tlearthEval Studio
DK

Eval Studio

Define datasets and rubrics, run judge-model evaluations, and safely A/B in production.

judge: aios-judge-1
5% of live traffic mirrored to B (shadow)
Guardrails: safety +5pt, cost ≤2×, latency p95 ≤1.5×
Datasets
Rubrics

Recent A/B runs

Every judged run is signed and appended to the ledger.

No runs yet.