We can't hire the experienced people we need — automate, train up, or outsource?

Ford rehired 300 senior engineers after AI failed at load-bearing tasks, and current AI agents score 20% or lower on complex multi-part tasks, risking expensive remediation and unreliable outputs.

· Counsel verdict · AIssential

The question

We have been trying for months to hire experienced people for judgment-heavy work — estimating, case handling, quality review — and the candidates are not there. Our options are to automate part of that work with AI, hire juniors and use AI to close the experience gap, outsource to a subcontractor, or keep searching and turn work away. Which of these holds up, what does each cost us in quality and risk, and what would have to be true for the AI option to be the right one rather than the fashionable one?

Counsel's position

Outsource judgment-heavy work to a subcontractor for immediate quality and risk mitigation, while piloting AI augmentation for junior staff on specific, less critical tasks.

Verdict

The verdict: Outsource judgment-heavy work to a subcontractor for immediate quality and risk mitigation, while piloting AI augmentation for junior staff on specific, less critical tasks.

How the criteria decide

2 of 3 criteria resolved on cited evidence. 1 had none either way.

CriterionFavoursEvidence
Quality and error risk on judgment workOutsource to subcontractor

Ford rehired 300 senior engineers after AI failed at load-bearing tasks

The car manufacturer Ford has rehired over 300 senior software engineers after attempts to replace them with AI failed.

Pivot to AI

AI agents score below on complex multi-part tasks

If we can get agents up to around 95% on these benchmarks then we are in a whole different world. This is like moving from driving assist features – which is where we are now with legal AI, to approaching Waymo

Artificial Lawyer

Self-reported probabilities from LLMs are uncalibrated and unusable

They had replaced ML models we built years earlier with AI agents. The results were unreliable, latency was much higher and outputs lacked a confidence score.

CIO

Time and cost to reach working capacityHire juniors with AI support

Agentic workflows consume hundreds of thousands of tokens per task

Agentic workflows, where models call themselves recursively, can chew through hundreds of thousands of tokens to accomplish a task that a human could do in five minutes.

LLM on Medium

AI touches a median of just 21% of adopting workers' tasks

Yet, even within occupations that show usage, the median is just 21% of that occupation’s tasks.

LeadDev

Dependence created on a vendor, subcontractor or key personNo evidence either way

Ford rehired 300 senior engineers after AI failed at load-bearing tasks

Generative AI has proven incapable of replacing critical human roles, leading companies to quietly rehire staff to clean up technical debt.

AI agents score below on complex multi-part tasks

Current AI agents require extensive human supervision and cannot yet be trusted with complex, multi-step professional work.

Agentic workflows consume hundreds of thousands of tokens per task

Automating complex workflows with recursive AI agents incurs prohibitive token costs compared to traditional software solutions.

AI touches a median of just 21% of adopting workers' tasks

AI functions overwhelmingly as an amplifier of human output rather than a substitute for it, and it tends to widen the gap between experienced and junior staff.

Self-reported probabilities from LLMs are uncalibrated and unusable

Replacing deterministic logic and traditional machine learning with generative AI agents often results in higher latency, prohibitive costs, and unreliable outputs.

Read another verdict

Get Counsel for your own decisions →