Skip to content
Signpost
In developmentAI

Model evaluation harness

A shared way to measure whether a prompt or model change actually improved anything, before it reaches customers.

Lifecycle

Next: rolling out on 7 December 2026, in 2 months.

  1. Proposed4 months ago
  2. In development (current stage)2 months ago
  3. Rolling outplanned, in 2 months
  4. Generally availableplanned, in 5 months

Detail

What this is

A test runner for model-backed features: you define cases and graders, it runs them on every change and reports a comparison against the previous version.

Who this affects

Teams shipping features backed by a model. Nothing changes for you until rollout in December, but contributing evaluation cases now is what makes it useful then.

Updates

Info

Harness runs nightly against the shared benchmark set

Two product teams have contributed graded examples. We need more, and from more domains — if your team has a task with known-good answers, talk to us.