Ten beats: what changed, what moved, and what we had to stop claiming. Each result belongs to its own test, not one continuous performance curve.
12 seconds per beat. Pause to read, or choose a date.
01 · September 13Harness gain
Give the question one front door.
25tool choices
→
1call by question kind
The call dispatches internally. Naming the record on the surface removes a routing decision the agent had been getting wrong.
Arrival at the relevant tool
78.6% → 100%
Answer correctness
75.7% → 95.2%
Zero mis-routes across 1,631 dispatches. Still the largest single improvement in this record.
02 · September 14Instrument repair
The test could be passed without reading it.
1repetition
→
5repetitions by default
Separately, four question-blind rules each scored perfectly. The corpus was rebuilt around records that exist but lack the requested field.
Blind-rule abstention F1 · before rebuild
1.000Each of four rules
Best blind-rule F1 · after rebuild
0.889A baseline, not an agent improvement
Refusing everything had scored 0.667 against the attached arm’s 0.646. Repetitions and the corpus rebuild repaired different weaknesses in the instrument.
03 · September 15New measurement
Make useful work the unit of the test.
8real-work cases
→
4cumulative gates
Each requested item must be present, receipted, correct and sourced. Refusing everything cannot earn completeness.
Bare completeness floor
0.000Nothing is receiptable without a world
What this establishes
Traceable workNot general actuarial judgement
The cases and reference answers are tied to a named world build. This introduces a test; it is not a measured harness gain.
04 · September 16–17Completeness cost
The ladder spreads out. The critic costs.
0.750top completeness
↔
0.000bottom completeness
Models fail at different gates. A critic reading the progress ledger lowers completeness on both clean model pairs.
Gemini 2.5 Flash · without → with critic
0.423 → 0.180
DeepSeek V4 Flash · without → with critic
0.613 → 0.456
The ladder targeted eight models. Two model runs were excluded from comparison for faults; two additional models could not run because the provider rejected temperature zero.
05 · September 16–17Selection cost
Eight opinions lose to one fixed choice.
8critiques per answer
→
5/8selection threshold
A vote chooses between two models’ answers. It cannot reliably identify when to switch away from the stronger model.
Always Flash → selector · accuracy
0.935 → 0.908
Difference from always taking Flash
−0.027Interval [−0.041, −0.014]
1,400 observations, 11,200 verdicts. No threshold from one to eight reaches the stronger model alone.
06 · September 18No established routing gain
A sharper critic still cannot route.
8free-text critiques
→
1typed question per row
Check the figure against its payload deterministically, then ask one typed question. Discrimination improves at one eighth the calls and one fourteenth the tokens.
Typed-critic AUC · strict / registered
0.839 / 0.755Earlier critic: 0.654
Best selector · difference from Flash
−0.005Interval [−0.024, +0.014]; includes zero
A better critic score and lower cost did not establish better answer selection. This is not evidence that the critic learned nothing.
07 · September 18Delivery trade-off
A cleaner report can contain less work.
1.81unreceipted figures / turn
→
1.04with the delivery gate
A delivery tool blocks figures that no payload served. Sometimes the model deletes a correct item instead of fetching its evidence.
DeepSeek V4 Flash · completeness change
−0.083
Items retained · change out of 320
23 / 9 fewerReceipted / correct items
The block-fetch-redeliver loop closed 30 times in 31. That local success did not preserve completeness.
08 · September 18Comparison invalidated
The data stood still. The surface did not.
0.423one binary
→
0.230another binary
Same model, cases and world package; different runtime tool listing. The surface must be a registered variable.
Receipted items
228 → 115
Unsupported figures per turn
0.84 → 3.35
Do not carry a comparison across binaries. A shared data package does not establish a shared experimental configuration.
09 · September 23Calibration and scorer drift
An earlier effect falls inside the noise.
Bandestimated leakage noise
>
Meanmeasured leakage
Calibrate acceptance bands from saved runs. Separately, rescoring the same answers moves correctness about five points after a number-matcher change.
Completeness bands · by model
0.087 / 0.151
Abstention-F1 band
0.0088
Leakage differences in the record sit inside the estimated noise. These bands are provisional: resampling saved runs does not replace fresh, independent runs of an unchanged configuration.
10 · September 24Harness gain
Hide what the agent never used.
25declared data tools
→
15visible after pruning
Hide ten tools that no prior measured run called. The data stays the same. The gains sit in the two failing question families.
Arrival · 0.5386 → 0.6771
+0.138595% interval [+0.0786, +0.2000]
Correctness · 0.5286 → 0.6514
+0.122995% interval [+0.0657, +0.1814]
Gemini 2.5 Flash, 280 questions, five repetitions. The control lost 10% of rows to runtime failures; excluding lost rows left the arrival gain almost unchanged. This is a separate comparison from the front door.
The shape of the record
Both clear wins removed choices.
2 harness gains
The front door and deletion. Both simplify the route to evidence.
3 measured costs
The ledger critic and delivery gate cost completeness. The vote costs selection accuracy.
3 instrument repairs
A rebuilt corpus, five repetitions, and a calibrated noise band.
2 invalidating discoveries
The runtime surface moved with fixed data. The scorer moved with fixed answers.
A check has to improve the work, not just the quantity it checks.
These findings overlap the ten beats; they are not ten independent interventions. The sharper critic improved discrimination and efficiency, but no routing gain was established.