A scorecard for the AI age
Summary
A new "scorecard for the AI age" proposes "Useful Intelligence per Dollar" as the primary metric for CFOs to evaluate AI investments, moving beyond traditional software adoption metrics. This framework, introduced on July 17, 2026, addresses four key questions: the quantity of useful work completed, the cost per successful task, the dependability of AI results, and the value growth per AI dollar as usage scales. The article highlights that a lower cost per token does not always equate to a lower cost per outcome, emphasizing the full cost of a successful task including human review and retries. It cites GPT-5.6, released last week with Sol, Terra, and Luna tiers, noting GPT-5.6 Sol's 54% reduction in output tokens on the Artificial Analysis Coding Agent Index. The framework also stresses the importance of dependability, tracking "Ready to use," "Needs correction," and "Needs escalation" outcomes, and the compounding gains from improved compute and models.
Key takeaway
For CFOs and AI/ML Directors evaluating AI investments, shift your focus from cost per token to "Useful Intelligence per Dollar." This means assessing the full cost of successful outcomes, including human review and retries, against the value created. Implement metrics tracking useful work, cost per successful task, and dependability to ensure your AI spend genuinely reduces effort and scales value, rather than just optimizing raw compute.
Key insights
The true value of AI is measured by "Useful Intelligence per Dollar," focusing on successful outcomes over raw compute cost.
Principles
- AI value is determined by work accomplished, not just adoption or token cost.
- Dependability (accuracy, consistency) directly reduces rework and increases confidence.
- Economic gains compound as better models and infrastructure drive adoption and investment.
Method
The proposed method for calculating "Useful Intelligence per Dollar" involves adding the full cost of completing work, counting tasks meeting quality, and dividing total cost by successful tasks.
In practice
- Define "done" for specific workflows (e.g., customer issue resolved).
- Track outcomes: "Ready to use," "Needs correction," "Needs escalation."
- Establish clear boundaries for AI access, systems, and human review.
Topics
- AI Value Measurement
- Useful Intelligence per Dollar
- AI Cost Optimization
- AI Model Dependability
- Workflow Automation
- GPT-5.6
Best for: CTO, VP of Engineering/Data, AI Product Manager, Executive, Director of AI/ML, Consultant
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by OpenAI News.