We can't hire the experienced people we need — automate, train up, or outsource?
Ford rehired 300 senior engineers after AI failed at load-bearing tasks, and current AI agents score 20% or lower on complex multi-part tasks, risking expensive remediation and unreliable outputs.
The question
We have been trying for months to hire experienced people for judgment-heavy work — estimating, case handling, quality review — and the candidates are not there. Our options are to automate part of that work with AI, hire juniors and use AI to close the experience gap, outsource to a subcontractor, or keep searching and turn work away. Which of these holds up, what does each cost us in quality and risk, and what would have to be true for the AI option to be the right one rather than the fashionable one?
Counsel's position
Outsource judgment-heavy work to a subcontractor for immediate quality and risk mitigation, while piloting AI augmentation for junior staff on specific, less critical tasks.
Verdict
The verdict: Outsource judgment-heavy work to a subcontractor for immediate quality and risk mitigation, while piloting AI augmentation for junior staff on specific, less critical tasks.
How the criteria decide
2 of 3 criteria resolved on cited evidence. 1 had none either way.
| Criterion | Favours | Evidence |
|---|---|---|
| Quality and error risk on judgment work | Outsource to subcontractor | Ford rehired 300 senior engineers after AI failed at load-bearing tasks The car manufacturer Ford has rehired over 300 senior software engineers after attempts to replace them with AI failed. AI agents score below on complex multi-part tasks If we can get agents up to around 95% on these benchmarks then we are in a whole different world. This is like moving from driving assist features – which is where we are now with legal AI, to approaching Waymo Self-reported probabilities from LLMs are uncalibrated and unusable They had replaced ML models we built years earlier with AI agents. The results were unreliable, latency was much higher and outputs lacked a confidence score. |
| Time and cost to reach working capacity | Hire juniors with AI support | Agentic workflows consume hundreds of thousands of tokens per task Agentic workflows, where models call themselves recursively, can chew through hundreds of thousands of tokens to accomplish a task that a human could do in five minutes. AI touches a median of just 21% of adopting workers' tasks Yet, even within occupations that show usage, the median is just 21% of that occupation’s tasks. |
| Dependence created on a vendor, subcontractor or key person | No evidence either way |
Ford rehired 300 senior engineers after AI failed at load-bearing tasks
Generative AI has proven incapable of replacing critical human roles, leading companies to quietly rehire staff to clean up technical debt.
AI agents score below on complex multi-part tasks
Current AI agents require extensive human supervision and cannot yet be trusted with complex, multi-step professional work.
Agentic workflows consume hundreds of thousands of tokens per task
Automating complex workflows with recursive AI agents incurs prohibitive token costs compared to traditional software solutions.
AI touches a median of just 21% of adopting workers' tasks
AI functions overwhelmingly as an amplifier of human output rather than a substitute for it, and it tends to widen the gap between experienced and junior staff.
Self-reported probabilities from LLMs are uncalibrated and unusable
Replacing deterministic logic and traditional machine learning with generative AI agents often results in higher latency, prohibitive costs, and unreliable outputs.
Read another verdict
- Start with a small test, or take on the whole process at once?
- Our competitors advertise AI and we don't — match them, or hold the line?
- Our people already put client files into ChatGPT — ban it, frame it, or supply a tool?
- Our most experienced person retires in two years — how do we keep what they know?
- Slow our EU AI Act prep now the deadline's moved to 2027?
- Use AI to flatten middle management this year?
- Let an AI agent act on its own — or keep a human in the loop?
- Invest in pre-build costing or post-deployment ROI tracking?