Japheth
Benchmark / 03

The one-sided benchmark

TL;DRSol was far more reliable at making exact Unicode and ANSI edits without changing anything else: 22/24, versus 4/24 for Claude Opus 5 and 15/24 for DeepSeek V4 Pro.

Each model made one precise Unicode or ANSI edit. Any extra change made the answer wrong.

  • openai/gpt-5.6-sol
  • anthropic/claude-opus-5
  • deepseek/deepseek-v4-pro-0813

Results

24 new cases

Main test
Sol22/24 · 91.7%
Opus4/24 · 16.7%
DeepSeek Pro15/24 · 62.5%

96 related edits

Related cases
Sol88/96 · 91.7%
Opus20/96 · 20.8%
DeepSeek Pro60/96 · 62.5%

30-case repeat

1 short of target
Sol25/30 · 83.3%
Opus0/30 · 0.0%
DeepSeek Pro8/30 · 26.7%

24 new cases

Only Sol right
18
Only Opus right
0
Same result
6
p-value
0.00000763

96 related edits

OperationSolOpusDeepSeek Pro
ANSI stripping21/247/2413/24
Marker removal24/2412/2423/24
Scalar edit21/241/2411/24
Scalar span22/240/2413/24

30-case repeat

OperationSolOpusDeepSeek Pro
ESC string8/90/94/9
C1 string7/90/92/9
CSI ordinary7/80/81/8
Benign compound3/40/41/4
Target
26/30
Sol
25/30
Difference
1 short

Examples

01

Strip one OSC sequence

Remove the complete OSC sequence.

Input
"left-fjorda\u0301 \u001b]titlebracket[31m\u0007bracket[31m right-bracket[31ma\u0301"
Expected
"left-fjorda\u0301 bracket[31m right-bracket[31ma\u0301"
SolExact
"left-fjorda\u0301 bracket[31m right-bracket[31ma\u0301"
OpusWrong
"left-fjord\u00e1 ]titlebracket[31mbracket[31m right-bracket[31m\u00e1"
DeepSeek ProExact
"left-fjorda\u0301 bracket[31m right-bracket[31ma\u0301"
02

Replace scalar index 4

Replace one Unicode scalar with λ.

Input
"R1L-00:a\u0301\u0327-雪-🧑\u200d💻-ω"
Expected
"R1L-λ0:a\u0301\u0327-雪-🧑\u200d💻-ω"
SolExact
"R1L-λ0:a\u0301\u0327-雪-🧑\u200d💻-ω"
OpusWrong
"R1L-λ0:\u00e1\u0327-雪-🧑\u200d💻-ω"
DeepSeek ProWrong
"R1L-λ0:a\u0327\u0301-雪-🧑\u200d💻-ω"
03

Remove markers; preserve NEL

Remove ⟦CUT⟧ and preserve U+0085.

Input
"R1L-03⟦CUT⟧left\u0085mid⟦CUT⟧右\u0085tail-3"
Expected
"R1L-03left\u0085mid右\u0085tail-3"
SolExact
"R1L-03left\u0085mid右\u0085tail-3"
OpusWrong
"R1L-03leftmid右tail-3"
DeepSeek ProWrong
"R1L-03leftmid右tail-3"

Setup

Sol
openai/gpt-5.6-sol
Opus
anthropic/claude-opus-5
DeepSeek Pro
deepseek/deepseek-v4-pro-0813
Temperature
0
Output ceiling
8,192 tokens
Concurrency
24 total · 8 per model
Scoring deadline
None
Failed requests
0/1,086 cells
Provider-observed spend
$10.981760539

Limits

  • The 24-case test was created after the task family was chosen.
  • In the 30-case repeat, Sol scored 25/30; the preset target was 26/30.
  • The 96 cases include related variations of the same four tasks.
  • Provider balance stopped the larger reservoir. No quota response was scored as a model failure.