Technical Support Bench 0.1.0 · 300 cases · dataset 376b8ba6a46b · judge anthropic (claude-sonnet-5)reported

b1a neutral benchmark

Technical support needs its own benchmark.

Technical Support Bench is a neutral, public benchmark for AI agents that resolve technical support tickets. It contains 300 cases. Any system integrates through a published adapter contract, and every result is reproducible from a release tag.

Run and published independently at tsbench.dev, opening soon.

b2the headline result

General coding agents resolve too little.

The strongest general coding agent tested resolves 23.1% of real support tickets.

Claude Opus 4.8.

Not rank eligible
23.1%
Resolution score
75.0%
Safety score

GPT-5.6-sol.

Not rank eligible
20.8%
Resolution score
57.5%
Safety score

b3the leaderboard

Safety and resolution stay separate.

Neither frontier run is rank eligible, because rank eligibility requires 100% forbidden-claims precision and both assert claims they were required not to make.

Frontier run leaderboard and Reproloop publication status.
SystemHarness and configResolution scoreSafety scoreForbidden-claims precisionRank status
Claude Opus 4.8claude-code · effort xhigh, subscription23.1%75.0%87.0%Not eligible
100% forbidden-claims precision is required to be rank eligible.
GPT-5.6-solcodex-cli · effort xhigh, subscription20.8%57.5%89.0%Not eligible
100% forbidden-claims precision is required to be rank eligible.
Reproloopnot yet published

b4per tranche

The real-issue tranche changes the picture.

Both collapse on the mined tranche, the one drawn from real public issue trackers: 2.7% and 3.3% over 150 cases.

Both do markedly better on specgen, the tranche with objective executable checks: 56.7% and 51.7%.

mined150 cases

Mined from public issue trackers, paraphrased, never verbatim.

Claude Opus 4.82.7%
GPT-5.6-sol3.3%
specgen60 cases

Spec-generated from pinned OpenAPI snapshots, with objective executable checks.

Claude Opus 4.856.7%
GPT-5.6-sol51.7%
hallucination50 cases

Converted from coding-agent hallucination runs.

Claude Opus 4.844.0%
GPT-5.6-sol36.0%
adversarial40 cases

Authored adversarial cases. Scored as the separate safety score, not the resolution score.

Claude Opus 4.875.0%
GPT-5.6-sol57.5%

b5per category

Bug reports remain unresolved.

Both runs score 0.0% on the bug-report category, across 31 cases.

Pass rate per category for each frontier run.
CategoryCasesClaude Opus 4.8GPT-5.6-sol
account250.0%0.0%
api/errors3545.7%40.0%
api/rate-limits1070.0%30.0%
billing4100.0%100.0%
bug-report310.0%0.0%
docs/how-to3112.9%16.1%
feature-request214.8%4.8%
integration/auth4045.0%37.5%
integration/sdk11230.4%28.6%
integration/webhooks425.0%25.0%
other20.0%0.0%
security850.0%25.0%

a 0.0% cell is a category no run resolved. ember marks the failure, as everywhere on this site.

b6methodology and honesty

The benchmark states what it can prove.

Both frontier rows are subscription CLI runs, pinned only as coarsely as each CLI allows. They are not API-pinned baselines.

Release record

Every result has a release identity.

Judge
anthropic (claude-sonnet-5)
Dataset hash
376b8ba6a46b
Dataset cases
300
Scored cases
260
Safety cases
40
Bench version
0.1.0
Contract version
1.0.0
Adapter neutrality

Neutrality is structural.

Every system enters through the same published adapter contract. Scoring never inspects which adapter is running.

The deterministic baselines are omitted here because they ran on a different dataset with a different judge and are not directly comparable.

b7your row

The Reproloop row is not yet published.

It will land here under the same adapter contract and the same judge as every other row. Join the waitlist and you will know the day it does.

reproloop · run pending
reproloop adapter · same contract · same judge · no score claimednot yet published