Build Templates Harness packs Evaluation runs
Evaluation · Support triage v2
Run the pack's evaluation set before publishing a new version.
What is a test run?
Each row here is one recorded test session: one frozen draft of one pack, run against its full test suite by a background worker. A draft can’t go live until a session for that exact draft finishes with every test passed — so this page is the release ledger for AI behavior changes.
Passed
37
of 40 cases
Failed
3
92.5% pass rate
Last run
Running
started 40s ago
Case results
Below 95% gate| Case | Expected | Got | Result |
|---|---|---|---|
| Refund request → route | Billing | Billing | Pass |
| Password reset → route | Support | Support | Pass |
| Angry churn → escalate | Escalate | Support | Fail |
| Feature ask → route | Product | Product | Pass |
Behind this page/templates/harnesses/:harnessPackKey/evaluations · CanManageTemplates
Retrieve
Sets/cases/runs read overHarnessEvalSets, HarnessEvalCases, HarnessEvalRuns; results poll to completion.Save
Queue →sp_QueueHarnessEvalRun; workers use sp_LeaseHarnessEvalRun / sp_CompleteHarnessEvalRun.Success
- Publish is blocked below the pass gate.
- Runs execute async with polling.
- Per-case expected vs got is shown.
CLI handoff
Implement this scaffold from the structured contract, then remove hard-coded preview rows. The source of truth is CLI Handoff and admin-cli-manifest.json.
Page IDharness-evaluation
Route/templates/harnesses/:harnessPackKey/evaluations
AccessCanManageTemplates
Statusworker-integration-required
Server-inject identity and scope values; never trust browser-supplied account, app, tenant, user, entitlement, price, or permission identifiers. Preserve the loading, empty, forbidden, failed, retrying, and completed states shown by the preview.