Author, validate, preview, calibrate, diff, and export criterion-level AI evaluation rubrics.
-
Updated
Jul 20, 2026 - TypeScript
Author, validate, preview, calibrate, diff, and export criterion-level AI evaluation rubrics.
Probe judge behavior for position, verbosity, self-preference, paraphrase, anchoring, and calibration failures.
Validate and render a disclosure record for judge prompts, calibration, bias, limitations, and intended use.
Validate required Datasheet, Model Card, and Data Card sections and surface actionable review findings.
Measure inter-annotator agreement with modern metrics, uncertainty, ordinal support, and missing-data handling.
Import rubric JSON and response JSONL, validate them, run deterministic scoring, and inspect exact evidence in the browser.
Validate, lint, diff, and convert portable AuraOne Rubric Schema v1 files.
Convert OpenTelemetry or Phoenix GenAI trace exports into local evaluation regression cases and manifests.
Validate and summarize human intervention and recovery segments in robot episodes.
Run executable rubric-spec v1 compatibility checks and emit reproducible conformance evidence.
Score repository responses against an AuraOne rubric and publish one consistent GitHub evidence surface.
Normalize agent tool-call traces into deterministic replay artifacts and pytest cases.
Validate and render structured cards for robot morphology, sensors, control, environment, data, and known limitations.
Diff and lint rubric changes, annotate affected lines, and publish one merge decision for the exact commit.
Shared Proofline UI, runtime contracts, keychain, updater, and release infrastructure for AuraOne Open desktop Studios.
Add a description, image, and links to the auraone topic page so that developers can more easily learn about it.
To associate your repository with the auraone topic, visit your repo's landing page and select "manage topics."