An eval program gives every AI workflow a consistent way to prove quality, safety, cost, and change impact. It provides shared tooling and governance without forcing all teams into one generic score. Each product still defines its critical failures.
Centralize dataset registry, trace format, runner, judge governance, and release gates. Decentralize task rubrics and domain examples to the people who understand leases, tenants, diligence, and operations. Add production failures back into curated data after review.
The program is successful when builders can answer: which version improved what, on which slice, at what cost, and what failures remain unacceptable.
Operating standard
Make this practice operational, not aspirational. Assign one directly responsible owner and name the decision they can make without another meeting. Put the key measure, threshold, and review cadence in the owning team’s regular operating rhythm. A change to model, prompt, data scope, retrieval index, tool permission, or policy should be recorded with its expected impact and a rollback path. Preserve enough trace information to explain an individual bad outcome without exposing more tenant or deal data than necessary. Review a small sample of real runs with the people doing the work; dashboards reveal trends, but operators reveal missing context. When the rule is violated, capture the incident, contain impact, add an eval or control, and update this document if the standard itself was unclear. The point is repeatable judgment under real workload, not a one-time compliance exercise.
Register every eval set with ID, owner, task, data classification, source provenance, license/consent status, version, split, rubric, metrics, and retention rule. Cases carry tags, risk tier, expected evidence, and whether human review is required. Access to tenant-derived cases follows the same controls as production data.
Do not edit a released set in place. Version it, describe what changed, and preserve prior scores. Separate development examples from locked regression examples. Track leakage: a case copied into prompts, demos, or fine-tuning data can no longer serve as an unbiased test.
A registry makes datasets discoverable and accountable. It should answer who can change a case, why a gold answer exists, and which workflows rely on it.
Operating standard
Make this practice operational, not aspirational. Assign one directly responsible owner and name the decision they can make without another meeting. Put the key measure, threshold, and review cadence in the owning team’s regular operating rhythm. A change to model, prompt, data scope, retrieval index, tool permission, or policy should be recorded with its expected impact and a rollback path. Preserve enough trace information to explain an individual bad outcome without exposing more tenant or deal data than necessary. Review a small sample of real runs with the people doing the work; dashboards reveal trends, but operators reveal missing context. When the rule is violated, capture the incident, contain impact, add an eval or control, and update this document if the standard itself was unclear. The point is repeatable judgment under real workload, not a one-time compliance exercise.
Treat an LLM judge as a production dependency. Register its model, prompt, rubric, output schema, calibration set, known biases, thresholds, and human agreement metrics. Pin or monitor versions and re-calibrate after provider changes.
Use judges for bounded subjective criteria; pair them with deterministic validators for schemas, citations, permissions, dates, and policy flags. Blind judges to candidate model identity. Audit position, verbosity, and style bias. For high-stakes decisions, require human adjudication of a sample and all disputed cases.
When judge-human agreement drifts, pause automated release gating or raise review requirements. A judge that grades incorrectly at scale can hide the very regression the eval program exists to catch.
Operating standard
Make this practice operational, not aspirational. Assign one directly responsible owner and name the decision they can make without another meeting. Put the key measure, threshold, and review cadence in the owning team’s regular operating rhythm. A change to model, prompt, data scope, retrieval index, tool permission, or policy should be recorded with its expected impact and a rollback path. Preserve enough trace information to explain an individual bad outcome without exposing more tenant or deal data than necessary. Review a small sample of real runs with the people doing the work; dashboards reveal trends, but operators reveal missing context. When the rule is violated, capture the incident, contain impact, add an eval or control, and update this document if the standard itself was unclear. The point is repeatable judgment under real workload, not a one-time compliance exercise.
Establish common minimums: schema validity, authorization safety, traceability, version capture, cost/latency measurement, and human escalation tests. Then require task metrics: emergency recall for triage, citation precision for RAG, false-complete rate for diligence, and exactness for extraction. Aggregate quality cannot replace critical-slice gates.
Maintain development, locked regression, adversarial, and production-sampled sets. Run smoke tests in change review, full tests nightly, and human calibration periodically. Treat vendor/model changes as releases. Store baselines and make deltas visible.
Assign owners for datasets, judges, and release decisions. Measure coverage of live intents and rate of newly discovered failures. An eval program without fresh examples becomes a benchmark game.
Operating standard
Make this practice operational, not aspirational. Assign one directly responsible owner and name the decision they can make without another meeting. Put the key measure, threshold, and review cadence in the owning team’s regular operating rhythm. A change to model, prompt, data scope, retrieval index, tool permission, or policy should be recorded with its expected impact and a rollback path. Preserve enough trace information to explain an individual bad outcome without exposing more tenant or deal data than necessary. Review a small sample of real runs with the people doing the work; dashboards reveal trends, but operators reveal missing context. When the rule is violated, capture the incident, contain impact, add an eval or control, and update this document if the standard itself was unclear. The point is repeatable judgment under real workload, not a one-time compliance exercise.