Benchmark / 03
The one-sided benchmark
TL;DRSol was far more reliable at making exact Unicode and ANSI edits without changing anything else: 22/24, versus 4/24 for Claude Opus 5 and 15/24 for DeepSeek V4 Pro.
Each model made one precise Unicode or ANSI edit. Any extra change made the answer wrong.
- openai/gpt-5.6-sol
- anthropic/claude-opus-5
- deepseek/deepseek-v4-pro-0813
Results
24 new cases
Main test96 related edits
Related cases30-case repeat
1 short of target24 new cases
- Only Sol right
- 18
- Only Opus right
- 0
- Same result
- 6
- p-value
- 0.00000763
96 related edits
| Operation | Sol | Opus | DeepSeek Pro |
|---|---|---|---|
| ANSI stripping | 21/24 | 7/24 | 13/24 |
| Marker removal | 24/24 | 12/24 | 23/24 |
| Scalar edit | 21/24 | 1/24 | 11/24 |
| Scalar span | 22/24 | 0/24 | 13/24 |
30-case repeat
| Operation | Sol | Opus | DeepSeek Pro |
|---|---|---|---|
| ESC string | 8/9 | 0/9 | 4/9 |
| C1 string | 7/9 | 0/9 | 2/9 |
| CSI ordinary | 7/8 | 0/8 | 1/8 |
| Benign compound | 3/4 | 0/4 | 1/4 |
- Target
- 26/30
- Sol
- 25/30
- Difference
- 1 short
Examples
Strip one OSC sequence
Remove the complete OSC sequence.
- Input
"left-fjorda\u0301 \u001b]titlebracket[31m\u0007bracket[31m right-bracket[31ma\u0301"- Expected
"left-fjorda\u0301 bracket[31m right-bracket[31ma\u0301"
"left-fjorda\u0301 bracket[31m right-bracket[31ma\u0301""left-fjord\u00e1 ]titlebracket[31mbracket[31m right-bracket[31m\u00e1""left-fjorda\u0301 bracket[31m right-bracket[31ma\u0301"Replace scalar index 4
Replace one Unicode scalar with λ.
- Input
"R1L-00:a\u0301\u0327-雪-🧑\u200d💻-ω"- Expected
"R1L-λ0:a\u0301\u0327-雪-🧑\u200d💻-ω"
"R1L-λ0:a\u0301\u0327-雪-🧑\u200d💻-ω""R1L-λ0:\u00e1\u0327-雪-🧑\u200d💻-ω""R1L-λ0:a\u0327\u0301-雪-🧑\u200d💻-ω"Remove markers; preserve NEL
Remove ⟦CUT⟧ and preserve U+0085.
- Input
"R1L-03⟦CUT⟧left\u0085mid⟦CUT⟧右\u0085tail-3"- Expected
"R1L-03left\u0085mid右\u0085tail-3"
"R1L-03left\u0085mid右\u0085tail-3""R1L-03leftmid右tail-3""R1L-03leftmid右tail-3"Setup
- Sol
- openai/gpt-5.6-sol
- Opus
- anthropic/claude-opus-5
- DeepSeek Pro
- deepseek/deepseek-v4-pro-0813
- Temperature
- 0
- Output ceiling
- 8,192 tokens
- Concurrency
- 24 total · 8 per model
- Scoring deadline
- None
- Failed requests
- 0/1,086 cells
- Provider-observed spend
- $10.981760539
Limits
- The 24-case test was created after the task family was chosen.
- In the 30-case repeat, Sol scored 25/30; the preset target was 26/30.
- The 96 cases include related variations of the same four tasks.
- Provider balance stopped the larger reservoir. No quota response was scored as a model failure.