EVIDENCE REGISTER
What we sell, and what it actually rests on.
Every sentence in this register is one we publish somewhere — the homepage, the company page, the profile. Under each one: the study it rests on, and the point where that study stops. We opened every primary listed here. The two we could not open are named below rather than quoted second-hand.
Three of the seven entries argue against us, including the last one, which says the central claim of the offer has never been tested. A register that only supports the seller is a brochure with footnotes.
The register
"The model can do the work once. Not a thousand times in a row."
Homepage · company page
Measured
τ-bench (Yao et al., arXiv:2406.12045), verbatim: "even state-of-the-art function calling agents (like gpt-4o) succeed on < 50% of the tasks, and are quite inconsistent (pass^8 < 25% in retail)." CRMArena-Pro (Salesforce Research, arXiv:2505.18878): leading agents reach ~58% single-turn success and drop to 35% in multi-turn settings, across nineteen expert-validated tasks. METR (arXiv:2503.14499) finds the 80%-reliability time horizon is "roughly 5x shorter" than the 50% one — every "AI can now do X-hour tasks" headline quotes the coin-flip threshold.
Where it stopsτ-bench runs on simulated users and a synthetic database, with 2024 models. And the counterweight belongs here: on GDPval (arXiv:2510.04374, OpenAI's own benchmark), 47.6% of Claude Opus 4.1 deliverables beat or matched professionals averaging 14 years of experience — on the gold subset. One-shot quality is the demo. Repetition against state is production. Both are true.
Open the primary →"Nobody is replaced. The system applies their rules."
Profile · Experience entry
Measured
Goldberg, "Man versus model of man", Psychological Bulletin 73(6), 1970, pp. 422–432. Least-squares models were built of each of 29 clinical psychologists, on their own judgements of 861 MMPI profiles. Built on all 861 cases, 86% of the models predicted the actual diagnoses better than the clinician they were derived from (p < .001, sign test); for 28 of the 29 judges (97%) the model was at least as valid. The paper states the mechanism plainly: the clinician "has his days" — boredom, fatigue, distraction — and the model removes that noise, not the knowledge.
Where it stopsThe advantage shrinks with the sample the model is derived from: 86% on all cases, 79% on a seventh, 72% on a tenth, 62% on 33 cases. And the honest limit: modelling did NOT improve on the composite judgement of all 29 clinicians pooled. Against several experts pooled, the edge disappears. 1970s clinical judgement is not job pricing.
Open the primary →"We put the know-how that lives in people's heads into the process."
Homepage · profile · company page
Measured
Brynjolfsson, Li & Raymond, "Generative AI at Work", NBER w31161 (published in the Quarterly Journal of Economics 140(2), 2025). 5,179 customer-support agents, 3 million chats, staggered rollout, difference-in-differences on the firm's own administrative logs. Access to the assistant raised issues resolved per hour by 14% on average — 34% for novice and low-skilled workers, with minimal impact on the experienced and highly skilled. The tool was trained on the firm's own corpus, up-weighting the chats of top performers.
Where it stopsThe authors call the knowledge-diffusion mechanism "suggestive" themselves, and the design is quasi-experimental, not a randomised trial. We read the free NBER working paper (5,179 agents, 14%); the paywalled published version reports 5,172 and 15% — we cite what we opened. 89% of the agents worked outside the United States, mostly in the Philippines. And there is no generic-AI control arm: see the last entry.
Open the primary →"We first measure what the process costs today."
Homepage · profile · company page
Argues against us
METR (arXiv:2507.09089): a randomised controlled trial on 16 experienced open-source developers completing 246 real issues in repositories they had worked on for about five years. They forecast AI would cut their time by 24%. Afterwards they estimated it had cut it by 20%. Measured, it increased completion time by 19%. Economists had predicted −39%, ML experts −38%. Humlum & Vestergaard (NBER w33777): 25,000 workers across 7,000 workplaces in 11 exposed occupations in Denmark, survey responses linked to administrative registers — "precise null effects on earnings and recorded hours… ruling out effects larger than 2%" two years on. Adopters report saving about 3% of their work hours, and 85% of them reallocate that time to other tasks.
Where it stopsMETR is 16 developers, all expert maintainers, and its authors say explicitly this is not evidence that AI fails to speed up most developers. Generalise the perception gap, never the +19%. Humlum measures earnings and hours, not the quality of the work. What both establish is narrow and enough: self-report gets the sign wrong, so a number agreed before the work starts is not a formality.
Open the primary →"The system is yours: it evolves without me, with no subscription to my presence."
Company page · profile
Measured, once
Soloway, Bachant & Jensen (the last two at Digital Equipment Corporation), AAAI-87, pp. 824–829, published while XCON was in production. Verbatim: "Over 7 years, XCON has grown to about 6,200 rules, of which approximately 50% change every year. While the performance of XCON is satisfactory, it is increasingly becoming more difficult to change." It started at about 700 rules, and the paper notes that "the problems of continually updating such a large system do not grow linearly". The system worked. Keeping it alive was the problem.
Where it stopsOne industrial case, 1987, a rule-based system rather than a model. The mechanism transfers — custody and maintenance, not knowledge quality, is what kills these systems — the figures do not. The study we wanted beside it, Gill's 1995 survey of early expert systems in MIS Quarterly, is paywalled everywhere we looked. We did not read it, so it is not here.
Open the primary →"You leave with the value proven."
Homepage · company page
Argues against us
Kwan et al., BMJ 2020;370:m3216 — 108 studies, 122 trials, 1,203,053 patients and 10,790 providers. Computerised decision support raised the share of patients receiving the desired care by 5.8 percentage points (95% CI 4.0 to 7.6). But in the 30 trials that reported clinical endpoints, the median improvement was 0.3% (IQR −0.7% to +1.9%). Behaviour moves reliably. Outcomes much less so. Roshanov et al., BMJ 2013;346:f657 — meta-regression of 162 randomised trials: systems evaluated by their own developers report an odds ratio of 4.35 (95% CI 1.66 to 11.44), a result the authors describe as robust across methods and internal validation.
Where it stopsThat last number is about us. We build the system and we would be the ones reporting on it. There is no rhetorical answer to it — only a structural one: the baseline is agreed by you before we start, and measured against what the process costs today, which is the one arrangement a developer cannot quietly grade. Both papers are clinical; carrying them across to job pricing is reasoning, not measurement. And heterogeneity is high (I²=76%) — the top quartile of results ran from 10% to 62%.
Open the primary →"An assistant grounded in your know-how beats a generic model on your work."
The load-bearing claim of the whole offer
Untested
Nothing tests it directly. No randomised field experiment compares an internally-grounded assistant against a generic model on the same task; two research passes went looking and found none. What exists in that space is vendor marketing with no method and no control. The closest indirect evidence is a pair: Brynjolfsson's +34% for novices with a firm-tuned corpus, set against Otis et al. — 640 Kenyan entrepreneurs, randomised, given a generic GPT-4 business mentor over WhatsApp — where low performers did about 8% worse.
Where it stopsDifferent populations, different tasks, different designs: this is not a head-to-head, and we will not present it as one. Otis reports a null average effect under a pre-registered plan; the gain for high performers ("just over 15%") is not a significant average effect, only a quantile effect at the median. What Otis establishes solidly is the harm to low performers, not the benefit to anyone. And we read the working paper — the Management Science version is paywalled.
Open the primary →
What we could not read
The rule for this page is that nothing enters it whose primary we have not opened. Three sources we wanted are behind paywalls with no repository copy, so they support nothing here:
- Garg et al., JAMA 293(10), 2005 — the landmark review of decision support. Closed everywhere. Its findings are echoed by Kwan (2020) and Roshanov (2013), which we did read, so the entry above stands without it.
- Gill, MIS Quarterly 19(1), 1995 — the five-year follow-up on early expert systems. Closed. Its abstract matches what we argue, which is exactly why we will not cite it on an abstract.
- The published Quarterly Journal of Economics version of Brynjolfsson, Li & Raymond. We read the free NBER working paper instead, and cite its numbers — 5,179 agents and 14%, not the published 5,172 and 15%.
The gap, and what we do about it
The claim that carries the whole engagement — that grounding a system in your own know-how beats a generic model on your own work — has never been put to a controlled test by anyone. The mechanism is well evidenced on both sides. The head-to-head has not been run.
So we run it on your files. The engagement opens with a short slice on your own past jobs — your documents, not your people's time — and the people who do the work say whether what comes back is theirs or is generic. That is not a demo of what AI can do. It is the only place where this question actually gets settled, and the answer can be no.
This register is maintained from our internal evidence file, and corrected when a primary says otherwise — reading them for this page turned up three misquotations of our own. If you find another, we want to know.
Its companion asks the opposite question: where the AI statistics everyone quotes actually come from. Or bring us a process — twenty minutes, and a straight answer either way: the service.