Info
Harness runs nightly against the shared benchmark set
Two product teams have contributed graded examples. We need more, and from more domains — if your team has a task with known-good answers, talk to us.
A shared way to measure whether a prompt or model change actually improved anything, before it reaches customers.
Next: rolling out on 7 December 2026, in 2 months.
A test runner for model-backed features: you define cases and graders, it runs them on every change and reports a comparison against the previous version.
Teams shipping features backed by a model. Nothing changes for you until rollout in December, but contributing evaluation cases now is what makes it useful then.
Two product teams have contributed graded examples. We need more, and from more domains — if your team has a task with known-good answers, talk to us.