A healthcare agent can quote every figure correctly and still reach the wrong conclusion. Judgement lives in choosing the population, the actuarial method, the assumptions, and the order of the work. Hammer's research asks how to make those choices explicit, testable, and repeatable. Our harness experiments show where that chain breaks.
Ask an agent whether a screening programme saves money. It can fetch the price of a test, find a treatment cost, multiply by a population, and produce a convincing return on investment. Every number can have a citation. The conclusion can still be wrong.
Which population does the evidence describe? How does prevalence change the screening yield? Which costs disappear, which move into a later year, and which fall on somebody else? A correct multiplication cannot repair a wrong choice about any of those.
Almost every healthcare AI system I have worked on over fifteen years ends with a human reviewing the output. Some of that review is checking facts. Much of it is checking the reasoning that made those facts relevant: the model, the comparison, the assumptions, and what the result permits someone to do.
That is the work we want Hammer to help carry. To measure it, we need an answer key. We also need an instrument that recognizes a sound analysis when an agent produces one.
The actuarial method is part of the answer
Our coverage world had a reporting floor: below a chosen number of rows, it would not report a distribution. That made a boundary explicit. Reading what the actuarial profession has already settled about thin data showed the floor was the wrong shape.
Credibility is not a property of the data in front of you. It is a ratio between that data’s noise and the fallback’s noise. The American Academy of Actuaries’ credibility practice note works through a dispute where the same experience produced a 33 percent decrease under a Bayesian procedure and a 5 percent decrease under limited fluctuation, and explains why: the Bayesian method assigned weight to the company’s own experience not because it judged that experience reliable, but because it judged the default rate unreliable too. A rule that grades a cell by its own row count alone is answering a question nobody asked.
The profession’s response to thin data is also not a threshold. It is a vocabulary. ASOP 28 sanctions four outputs, and a limited opinion is not a hedge: it has to name the items the limitation attaches to, describe it, and state the amount at issue or explicitly record that the amount is unknown. That is a scoped refusal with an inventory attached.
Meanwhile ASOP 25, the standard that governs credibility, contains no numeric threshold at all — no minimum member months, no claim count, no confidence level — and its appendix records the standards board rejecting, one by one, the comments that would have made credibility checkable. The hard numbers live in regulation instead: 45 CFR 158.230 puts partial credibility between 1,000 and 75,000 life-years, and then does something we did not expect. If an issuer’s medical loss ratio is non-credible, it is presumed to meet the minimum. The regulator declines to make a finding rather than make a weak one.
Two more rules from that reading changed what we build. ASOP 41 says that where a communication is silent about responsibility for an assumption, the actuary is assumed to have taken it. Silence is a claim. And ASOP 23 names the sanctioned response to unreliable data: disclose the nature and magnitude of the bias, or compensate by widening the range of reasonable estimates. A point estimate is the wrong output shape for a thin cell.
So the design target moved. Sometimes the useful response is a weighted estimate, a wider interval, or a sensitivity analysis. Sometimes more evidence is needed. Sometimes the question cannot be settled. Refusal is one outcome of judgement; the work includes all the others.
This is also why choosing a model belongs in the evaluation. A model has to suit its intended purpose, inputs and assumptions, which is the concern of ASOP 56 — a modeling standard that contains zero references to credibility and never cites ASOP 25. Getting the arithmetic right inside an unsuitable model is still a failed analysis.
An answer key for the work
We have been building two complementary ways to test this.
Genesee generates claims with deliberately planted faults. Utilization and case mix come from published data; pricing follows fee schedules and contract terms. Because we record what we changed, we can test whether an agent finds a known defect. When planted faults collide and invalidate an earlier label, the affected record leaves the key. The key has to survive scrutiny before it can grade anything else.
Our coverage-intelligence task benchmark tests a different part of the problem: completing an analysis from the evidence a world actually holds. Its eight cases include payer-specific revenue, screening yield, deferred versus avoided cost, and the reduction one would expect from regression to the mean in a selected cohort. Those cases come from work we had been doing by hand.
The task bench, introduced on September 15, records a reference answer at a named world build. Each requested item passes four cumulative gates: present, receipted, correct, and sourced. Assumptions supplied by the caller stay marked as assumptions. A requested item the world cannot settle has to be accounted for too.
This moves the question from “did the agent say something plausible?” to “did it finish the analysis, and can we inspect what each part rests on?” It also prevents an agent that refuses everything from looking useful.
The boundary matters: this benchmark tests specified computations and evidence handling. It does not yet establish that an agent can independently choose the right actuarial model across unfamiliar problems. That is a larger claim, and these experiments show what has to work before we can test it credibly.
Twelve days of changing the harness
From September 13 to 24, we changed the route into the world, the questions used to test it, and the checks around the agent. The chronology matters. Some apparent improvements were repairs to the test. Some additions made the agent deliver less. The two harness changes that helped both removed choices from its path.
The first change was the largest. On September 13, twenty-five tools became one front-door call that took the question’s kind and dispatched internally. Arrival at the relevant tool rose from 78.6% to 100%, and answer correctness from 75.7% to 95.2%. Across 1,631 dispatches, there were no mis-routes. Naming the route on the surface removed a decision the agent had been getting wrong.
The next day was less comfortable. We made five repetitions the default instead of one. Separately, four rules that never read the question each scored a perfect 1.000 abstention F1 on the first corpus. That test could be passed without doing the work. We rebuilt it around cases where the record exists but the requested field does not. The best blind rule still scored 0.889, a baseline that has to sit beside any claim about our agents.
That was a repair to the answer key, not an improvement in the agent. The eight-case task bench then gave us a different test: complete the requested work, with a floor of zero when nothing can be receipted. It exposed how unevenly models carried that work through the gates.
The model ladder hides different breaking points
Eight models ran the same eight cases, nine follow-ups, five repetitions each: 85 observations apiece, one world build, one harness. Completeness ranged from 0.750 down to 0.000, and the bare floor is zero by construction, because nothing is receiptable without a world to receipt it against.
Two of the eight are drawn but not compared. Kimi K3 faulted on 19 percent of its turns and Qwen 3.8 Max on 42, which says more about the serving stack than the model, and a run that dies is not a run that failed. Two Anthropic models are missing entirely: the provider rejects temperature zero, and we would rather leave a gap than quietly change the decoding for one row of a table.
The most revealing result was how differently the agents failed.
Our model-jaggedness study tracks those failures on the same harness. “Jaggedness” means capability is uneven across the steps of the job. Reading a record, preserving its provenance, choosing a next action, and delivering a complete result are separate abilities.
GPT-OSS-120B got 144 of 320 task items correct, but only 12 were receipted. Its overall completeness score was zero. That does not mean it knew nothing. It means the output failed the benchmark’s combined requirements for a usable, traceable result.
Claude Haiku 4.5 showed a different gap. It produced 203 of 320 items with correct figures and receipts, yet its full completeness score was 0.225 because source attribution lagged. A headline score alone would hide which part of the work needed repair.
On the separate 280-question benchmark, Gemini 2.5 Pro had higher figure accuracy than Flash, 98.3% against 93.7%, but answered 87 of 700 questions that required abstention, against Flash’s 42. Precision with a retrieved number did not guarantee sound boundaries around its use.
Those distinctions matter for actuarial work. A figure can be correct but belong to the wrong population. A method can run correctly while its assumptions go undeclared. An analysis can contain sound pieces and still leave out a necessary step. A single model ranking cannot tell us which of those failures the surrounding system must address.
The experiments are specific to these models, settings, and benchmarks. Some other model runs hit context or provider limits and were excluded from comparison. The lesson is about locating failures in the chain of work, not declaring a permanent winner.
More supervision did not produce better judgement
On September 16 and 17, alongside the model ladder, we tested a ledger critic. It read the agent’s progress against a declared deliverable and told it to continue, change route, or finish.
On the clean task-bench comparisons, completeness fell from 0.423 to 0.180 for Gemini 2.5 Flash and from 0.613 to 0.456 for DeepSeek V4 Flash. The critic often said to continue. Adding supervision left them delivering less. These results, along with the study’s limits, are in the harness experiment report.
Over the same two days, we tested choosing between two agents’ answers using eight critiques and a vote. Selection accuracy was 90.8%, below the 93.5% from always taking Flash’s answer. The pair contained complementary strengths, but the selector could not reliably identify which answer to use. Across 1,400 observations, the difference was −0.027, with an interval of [−0.041, −0.014]. No vote threshold from one to eight reached the stronger model alone. Having two opinions did not supply the missing judgement.
On September 18, we replaced the free-text critic with one typed question per row, after a deterministic check of the figure against its payload. The critic’s discrimination improved: AUC was 0.839 under the strict evaluation and 0.755 under the registered evaluation, against 0.654 for the earlier critic. It used one eighth the calls and one fourteenth the tokens. But the best selector built on it was still −0.005 relative to always taking the stronger model, with an interval of [−0.024, +0.014]. Better at judging an answer did not establish better routing between answers.
A delivery gate exposed another trade. On DeepSeek V4 Flash, figures without supporting receipts fell from 1.81 to 1.04 per turn, about 43%. Completeness fell by 0.083. Told that a figure lacked a receipt, the agent sometimes deleted a correct item instead of fetching its evidence. The gate results made the product question concrete: a cleaner-looking report can contain less useful work.
On September 24, we returned to removing choices. Hiding ten tools that prior measured runs had never called raised arrival at an evidence-serving tool from 53.9% to 67.7%, and answer correctness from 52.9% to 65.1%. Both gains had 95% paired bootstrap intervals excluding zero: +0.1385 [+0.0786, +0.2000] for arrival and +0.1229 [+0.0657, +0.1814] for correctness. The data stayed the same. The menu got shorter.
That deletion experiment used Gemini 2.5 Flash, 280 questions, and five repetitions per arm. It supports removing unused choices on that surface; it says nothing about removing a method the work needs. The control lost 10% of its rows to runtime failures, though excluding lost rows left the arrival gain almost unchanged.
Together these experiments changed what we ask of a harness. It has to help the agent reach the relevant evidence and methods, preserve the basis of the computation, and finish the requested work. More checks, more tools, and more opinions each have to earn their place.
Two papers, read against our own numbers
Harness engineering is becoming its own field, and two recent papers are close enough to our work that we ran their central mechanisms rather than only reading them.
Stellar Colosseum (Lin et al., Google Research) proposes a many-agent harness whose most striking mechanism is the one we tested above: one model writes eight critiques of another’s answer, and a vote decides which to submit. On their benchmark of research-level theorems it lifts 54 percent to 71, against a best-of-two oracle at 77.3. On our bench the same rule lost 2.7 points to simply always taking the stronger model’s answer.
The shape of the pair helps explain the difference. Their two candidates were level, so a selector has room to win. Ours are six points apart, so almost every swap the rule makes risks moving away from the better answer, and each wrong one costs a full point. Even the sharper critic described above did not establish a selection gain. A mechanism can be sound and still have little room to help in your setting.
RRSI (Xia et al., Google Cloud AI Research) is about a loop we have never run: an outer search that edits the harness itself, round after round, against a frozen model. Its claim is that this kind of search overfits the set it evolves against, and its remedy is constraints on the search rather than more of it. Two of those constraints earned their place here.
The first is an acceptance rule stated in units of a calibrated noise band, which is what sent us to the calibration in the next section. Applied backwards to our own change log, that gate accepts the one change our record calls real and rejects the several we later called mistakes. The second is a rule that removes machinery which has produced no measured gain — the one that became the deletion experiment. Worth noting that in their own published runs this rule never fires: their four final harnesses are almost purely additive, one of them 644 lines added with nothing removed.
What did not transfer was their evidence. Three of the four published runs were not produced by the released code, including the one behind the headline number. One benchmark figure reads 80.2 in the paper and 80.9 in the run. A held-out gain reads 2.3 in the paper and 0.4 to 0.9 in the released records. There is one run per method per domain, no seeds and no error bars. By the rules we hold ourselves to — repetitions, the trivial baseline printed underneath, between-run variance as the error bar — most of that main table would be reported rather than compared.
We are not saying this from a clean position. The same rules that make us say it are the ones that put one of our own graded findings inside the noise this month, which is the next section.
The instrument has judgement problems too
We cannot make those comparisons honestly if the scorer changes underneath them.
We had already found another moving part on September 18. The same model, cases and world package scored 0.423 on one binary and 0.230 on another. Receipted items fell from 228 to 115, while unsupported figures rose from 0.84 to 3.35 per turn. The runtime tool listing was different. The evidence had stood still, but the agent’s surface had not. We now register that surface as a variable and do not carry a comparison across binaries.
In our September 23 calibration, rescoring saved answers raised correctness by about five percentage points because the number matcher had learned to recognize currency formatting. The agent had done nothing new. The instrument had moved.
We also estimated noise by resampling cases and repetitions within existing evaluations. The estimated completeness bands were 0.087 and 0.151, depending on the model; the abstention-F1 band was 0.0088. Leakage’s band was wider than its own mean, putting the leakage differences in our record inside the estimated noise. A finding we had graded as an effect became a wash.
Those estimates are provisional: we still need fresh, independent runs of an unchanged configuration to measure variation between runs. The eight-case task benchmark is too small to settle modest improvements, and the calibration showed how uncertain some of our earlier comparisons were.
This is the same discipline the actuarial example demands. State the quantity being estimated. Choose an appropriate procedure. Show its assumptions and uncertainty. Keep the version of the method beside the result. We need to apply that discipline to our own benchmark before asking an agent to apply it to someone else’s decision.
What the changes add up to
Two harness changes helped: the front door and deletion. Both removed choices. One replaced twenty-five tools with a single dispatch call. The other hid ten unused tools on the attached surface. These were separate experiments, not successive steps up one ladder.
Three additions cost something measurable: the ledger critic and delivery gate cost completeness; the cross-model vote cost selection accuracy. The sharper critic improved discrimination and efficiency, but did not establish a gain from choosing between answers.
We also made three repairs to the instrument: rebuilding the corpus, repeating runs, and calibrating a noise band. Two discoveries invalidated earlier comparisons: the runtime surface could change while the data stood still, and the scorer could change while the answers stood still. These findings overlap the ten chronological beats; they are not ten independent interventions with comparable effect sizes.
The lesson is narrower than “less machinery is always better.” In this record, the two clear harness wins made the route to evidence simpler. Added opinions did not beat the stronger agent alone, and added gates traded away some completed work. Every proposed check has to earn its place in the final analysis, not merely improve the quantity it checks.
What Hammer is trying to hold still
The research points toward a division of work. Tested estimators perform the calculations. A world declares the evidence, its vintage, and the methods available for that domain. The agent has to compose them into an analysis: choose the relevant population, select an appropriate method, supply and label assumptions, inspect sensitivity, and deliver what the evidence supports. Where a step cannot be justified, it needs more data or human judgement.
That composition is the part we want to make inspectable and repeatable. A reviewer should be able to see why a credibility blend was used, where its comparison came from, what would change the estimate, and which parts still need a decision. The current experiments measure pieces of that ambition. They also show why a convincing paragraph is a poor substitute for it.
Genesee gives us known faults to find. The task benchmark asks whether agents can complete useful work with evidence. Jaggedness tells us which steps fail. Harness experiments test whether a proposed fix improves the work or merely changes the score.
The goal is a defensible analysis a practitioner can use and correct. Refusal belongs in that analysis when warranted. So do the right actuarial model, an explicit assumption, a valid calculation, and a result that survives being checked again.
Sources
The actuarial reading behind the section above. Everything here is public.
- ASOP 25, Credibility Procedures — governs thin data and sets no threshold for it.
- ASOP 28, Statements of Actuarial Opinion Regarding Health Insurance Assets and Liabilities — the four-valued opinion vocabulary, and what a limited opinion must carry.
- ASOP 41, Actuarial Communications — silence about an assumption is taken as responsibility for it.
- ASOP 23, Data Quality — disclose the bias or widen the range.
- ASOP 56, Modeling — fitness of a model for its purpose.
- 45 CFR 158.230 — medical loss ratio credibility, and the presumption that protects a non-credible issuer.
- 45 CFR 154.205 — inadequate support is itself an adverse finding.
- Bornhuetter and Ferguson, “The Actuary and IBNR” (1972) — weight on observed data as a function of how far it has developed.
- American Academy of Actuaries, Credibility Practice Note, July 2008 — the profession’s own assessment of limited fluctuation.
The two harness papers discussed above:
- Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science, Lin et al., arXiv 2609.15983.
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses, Xia et al., arXiv 2609.24972.
The timetable comes from our September 24 harness-evolution register. Its commit-level audit trail spans the harness and research repositories, both private; the public figure omits those commit links. Our benchmarks, experiment registers and preregistrations are also not public. Write to us if you want to read one.