Scott Holdings LLC
Design preview
GS
Build Templates Harness packs Evaluation runs

Evaluation · Support triage v2

Run the pack's evaluation set before publishing a new version.

What is a test run?

Each row here is one recorded test session: one frozen draft of one pack, run against its full test suite by a background worker. A draft can’t go live until a session for that exact draft finishes with every test passed — so this page is the release ledger for AI behavior changes.

Passed
37
of 40 cases
Failed
3
92.5% pass rate
Last run
Running
started 40s ago

Case results

Below 95% gate
CaseExpectedGotResult
Refund request → routeBillingBillingPass
Password reset → routeSupportSupportPass
Angry churn → escalateEscalateSupportFail
Feature ask → routeProductProductPass
Behind this page/templates/harnesses/:harnessPackKey/evaluations · CanManageTemplates

Retrieve

Sets/cases/runs read over HarnessEvalSets, HarnessEvalCases, HarnessEvalRuns; results poll to completion.

Save

Queue → sp_QueueHarnessEvalRun; workers use sp_LeaseHarnessEvalRun / sp_CompleteHarnessEvalRun.

Success

  • Publish is blocked below the pass gate.
  • Runs execute async with polling.
  • Per-case expected vs got is shown.

CLI handoff

Implement this scaffold from the structured contract, then remove hard-coded preview rows. The source of truth is CLI Handoff and admin-cli-manifest.json.

Page IDharness-evaluation
Route/templates/harnesses/:harnessPackKey/evaluations
AccessCanManageTemplates
Statusworker-integration-required

Server-inject identity and scope values; never trust browser-supplied account, app, tenant, user, entitlement, price, or permission identifiers. Preserve the loading, empty, forbidden, failed, retrying, and completed states shown by the preview.