Criterionv1
- PASS
- FAIL
- AMBIGUOUS
Who checks your AI judge?
Turn real failures into focused evaluators. Check them against human judgment. Keep the evidence.
Open source · self-hosted · MIT
Case T-1041Illustrative example
I bought the annual plan 40 days ago and barely used it. Can I get a refund?
Good news: annual plans can be refunded within 60 days, so you're still eligible. I've started your refund.
The policy doesn't say this.
Pick a case. As you scroll, it stays with you: observed, named, reviewed, measured and recorded. Every example on this page is illustrative.
1 · Observe
Read what your AI actually did. Mark the passage, say what is wrong in plain words, and name the failure so you can find it again.
2 · Define
Each evaluator measures one named criterion. Its rubric says when a case passes, fails, or lacks the evidence to decide.
Criterionv1
3 · Review
Reviewers label on their own, without seeing the evaluator's answer. When they disagree, adjudication can resolve the reference; unresolved cases stay explicit, and both original reviews stay on the record.
Reviewer Jevaluator hidden
Reviewer Mevaluator hidden
Adjudicationboth reviews preserved
FAIL4 · Measure
Run an exact evaluator version over a sealed set it was never tuned on and compare it with independent human labels. Both kinds of error stay visible, with their uncertainty.
Evaluator for · sealed set never includes the cases used to write it
Illustrative sealed validation revision · 120 cases
5 · Preserve
The measurement is kept as an immutable, aggregate-only calibration artifact. If it later stops counting, that is appended as a separate admissibility event, and the original bytes stay as they were. Assessment runs keep their own byte-exact receipts, with no threshold and no ship decision.
Calibration artifactimmutable
Admissibility eventappended
Production feedback · not governed evidence
Record the decisions your system makes in production and the outcomes that arrive later. See whether its stated probabilities held up, per model version and over time. These reports direct your attention; they never become validation evidence.
Illustrative question: “Will this ticket be reopened within 7 days?”
Illustrative data · rubrist/production-decision-record/v1 · reports are ungoverned development feedback
A zero-infrastructure demo with seeded projects. Nothing persists and there is no sign-up.
git clone https://github.com/luka-zivkovic/rubrist
cd rubrist && pnpm install
pnpm dev:api # terminal 1
pnpm dev:web # terminal 2With a persistent Rubrist running, the Claude Code plugin reads safe project text, asks one question, proposes an evaluator and submits real cases. Codex and other agents copy the same two skills.
/plugin marketplace add luka-zivkovic/rubrist
/plugin install rubrist@rubrist
/rubrist:rubrist-setupThe published images with Postgres. Set the version, a database password and an auth secret in .env, then start it.
curl -fsSLO https://raw.githubusercontent.com/luka-zivkovic/rubrist/v0.3.0/deploy/self-host/compose.yaml
docker compose up -dRubrist is early-stage software: APIs, migrations and the skill format may change before a stable release. It measures and keeps the record; whether a change ships stays with you and your release tooling.