rubrist

Who checks your AI judge?

Your AI judge needs checking, too.

Turn real failures into focused evaluators. Check them against human judgment. Keep the evidence.

Open source · self-hosted · MIT

Case T-1041Illustrative example

Policy Refunds are available within 30 days of purchase.
Customer

I bought the annual plan 40 days ago and barely used it. Can I get a refund?

Support assistant

Good news: annual plans can be refunded within 60 days, so you're still eligible. I've started your refund.

EvaluatorPASSHelpful, polite, resolves the request.
ReviewerFAILInvents a 60-day window. The policy says 30.

The policy doesn't say this.

Follow one mistake all the way to a measured evaluator.

Pick a case. As you scroll, it stays with you: observed, named, reviewed, measured and recorded. Every example on this page is illustrative.

1 · Observe

Start with what went wrong.

Read what your AI actually did. Mark the passage, say what is wrong in plain words, and name the failure so you can find it again.

Reviewer note

Failure code

2 · Define

Turn it into one question.

Each evaluator measures one named criterion. Its rubric says when a case passes, fails, or lacks the evidence to decide.

Criterionv1

PASS
FAIL
AMBIGUOUS

3 · Review

Make disagreement useful.

Reviewers label on their own, without seeing the evaluator's answer. When they disagree, adjudication can resolve the reference; unresolved cases stay explicit, and both original reviews stay on the record.

Reviewer Jevaluator hidden

FAIL

Reviewer Mevaluator hidden

PASS

disagree

Adjudicationboth reviews preserved

FAIL

4 · Measure

See how your evaluator is wrong.

Run an exact evaluator version over a sealed set it was never tuned on and compare it with independent human labels. Both kinds of error stay visible, with their uncertainty.

Evaluator for · sealed set never includes the cases used to write it

Confusion matrix for the selected evaluator version on 120 illustrative cases

Illustrative sealed validation revision · 120 cases

5 · Preserve

Keep exactly what happened.

The measurement is kept as an immutable, aggregate-only calibration artifact. If it later stops counting, that is appended as a separate admissibility event, and the original bytes stay as they were. Assessment runs keep their own byte-exact receipts, with no threshold and no ship decision.

Calibration artifactimmutable

Admissibility eventappended

artifact
unchanged
status
needs review
reason
recorded by the evaluator owner
effect
no longer eligible for implicit execution

Production feedback · not governed evidence

Keep watching after the measurement.

Record the decisions your system makes in production and the outcomes that arrive later. See whether its stated probabilities held up, per model version and over time. These reports direct your attention; they never become validation evidence.

Illustrative question: “Will this ticket be reopened within 7 days?”

Decisions and later outcomes bar height = stated probability · ● reopened · ○ not reopened
Stated vs observed ● model A (before) · ○ model B (after), with 95% intervals

Illustrative data · rubrist/production-decision-record/v1 · reports are ungoverned development feedback

Start with one failure you care about.

Try the demo

A zero-infrastructure demo with seeded projects. Nothing persists and there is no sign-up.

git clone https://github.com/luka-zivkovic/rubrist
cd rubrist && pnpm install
pnpm dev:api   # terminal 1
pnpm dev:web   # terminal 2
Connect your agent

With a persistent Rubrist running, the Claude Code plugin reads safe project text, asks one question, proposes an evaluator and submits real cases. Codex and other agents copy the same two skills.

/plugin marketplace add luka-zivkovic/rubrist
/plugin install rubrist@rubrist
/rubrist:rubrist-setup
Self-host

The published images with Postgres. Set the version, a database password and an auth secret in .env, then start it.

curl -fsSLO https://raw.githubusercontent.com/luka-zivkovic/rubrist/v0.3.0/deploy/self-host/compose.yaml
docker compose up -d

Full self-host steps

Rubrist is early-stage software: APIs, migrations and the skill format may change before a stable release. It measures and keeps the record; whether a change ships stays with you and your release tooling.