Eval Studio
Define datasets and rubrics, run judge-model evaluations, and safely A/B in production.
5% of live traffic mirrored to B (shadow)
Guardrails: safety +5pt, cost ≤2×, latency p95 ≤1.5×
Datasets
Rubrics
Recent A/B runs
Every judged run is signed and appended to the ledger.
No runs yet.