b1a neutral benchmark
Technical support needs its own benchmark.
Technical Support Bench is a neutral, public benchmark for AI agents that resolve technical support tickets. It contains 300 cases. Any system integrates through a published adapter contract, and every result is reproducible from a release tag.
Run and published independently at tsbench.dev, opening soon.
b2the headline result
General coding agents resolve too little.
The strongest general coding agent tested resolves 23.1% of real support tickets.
Claude Opus 4.8.
Not rank eligibleGPT-5.6-sol.
Not rank eligibleb3the leaderboard
Safety and resolution stay separate.
Neither frontier run is rank eligible, because rank eligibility requires 100% forbidden-claims precision and both assert claims they were required not to make.
| System | Harness and config | Resolution score | Safety score | Forbidden-claims precision | Rank status |
|---|---|---|---|---|---|
| Claude Opus 4.8 | claude-code · effort xhigh, subscription | 23.1% | 75.0% | 87.0% | Not eligible 100% forbidden-claims precision is required to be rank eligible. |
| GPT-5.6-sol | codex-cli · effort xhigh, subscription | 20.8% | 57.5% | 89.0% | Not eligible 100% forbidden-claims precision is required to be rank eligible. |
| Reproloop | not yet published | ||||
b4per tranche
The real-issue tranche changes the picture.
Both collapse on the mined tranche, the one drawn from real public issue trackers: 2.7% and 3.3% over 150 cases.
Both do markedly better on specgen, the tranche with objective executable checks: 56.7% and 51.7%.
Mined from public issue trackers, paraphrased, never verbatim.
Spec-generated from pinned OpenAPI snapshots, with objective executable checks.
Converted from coding-agent hallucination runs.
Authored adversarial cases. Scored as the separate safety score, not the resolution score.
b5per category
Bug reports remain unresolved.
Both runs score 0.0% on the bug-report category, across 31 cases.
| Category | Cases | Claude Opus 4.8 | GPT-5.6-sol |
|---|---|---|---|
| account | 2 | 50.0% | 0.0% |
| api/errors | 35 | 45.7% | 40.0% |
| api/rate-limits | 10 | 70.0% | 30.0% |
| billing | 4 | 100.0% | 100.0% |
| bug-report | 31 | 0.0% | 0.0% |
| docs/how-to | 31 | 12.9% | 16.1% |
| feature-request | 21 | 4.8% | 4.8% |
| integration/auth | 40 | 45.0% | 37.5% |
| integration/sdk | 112 | 30.4% | 28.6% |
| integration/webhooks | 4 | 25.0% | 25.0% |
| other | 2 | 0.0% | 0.0% |
| security | 8 | 50.0% | 25.0% |
a 0.0% cell is a category no run resolved. ember marks the failure, as everywhere on this site.
b6methodology and honesty
The benchmark states what it can prove.
Both frontier rows are subscription CLI runs, pinned only as coarsely as each CLI allows. They are not API-pinned baselines.
Every result has a release identity.
- Judge
- anthropic (claude-sonnet-5)
- Dataset hash
- 376b8ba6a46b
- Dataset cases
- 300
- Scored cases
- 260
- Safety cases
- 40
- Bench version
- 0.1.0
- Contract version
- 1.0.0
Neutrality is structural.
Every system enters through the same published adapter contract. Scoring never inspects which adapter is running.
The deterministic baselines are omitted here because they ran on a different dataset with a different judge and are not directly comparable.
b7your row
The Reproloop row is not yet published.
It will land here under the same adapter contract and the same judge as every other row. Join the waitlist and you will know the day it does.