Unlimited AI tokens aren't unlimited after all as US Army burns through supply
Summary
The US Army's Combat Capabilities Development Command (DEVCOM) faced an unexpected exhaustion of AI tokens for its Ask Sage platform, despite an Army CIO announcement in May 2026 of "unlimited tokens." By mid-June 2026, the token pool was depleted, forcing the re-establishment of usage limits, with future renewal uncertain after October 1. Ask Sage, accredited for Controlled Unclassified Information, integrates models like Alphabet's Gemini, Meta's Llama, and OpenAI's ChatGPT, and is used for tasks such as reclassifying personnel descriptions. The Army had an annual enterprise pack of 100,000,000 tokens, with employees initially allocated at least 200,000 tokens monthly. This rapid consumption mirrors similar challenges at Meta and Uber, where engineers quickly burned through token supplies. An anonymous Army employee reported finding the generative AI tools unreliable and not particularly useful for their work.
Key takeaway
For Directors of AI/ML deploying enterprise generative AI, your initial token budget estimates are likely insufficient. You must implement robust token monitoring and allocation policies from day one to prevent rapid exhaustion and service disruption. Thoroughly vet AI tool utility and reliability for specific tasks before encouraging widespread adoption, as uncritical rollout can lead to resource waste and user dissatisfaction. Proactive management of AI resource consumption is critical for sustainable deployment.
Key insights
Unmanaged generative AI adoption rapidly exhausts token supplies, forcing unexpected usage limits.
Principles
- Generative AI token consumption often exceeds initial estimates.
- Uncritical AI deployment risks inefficiency and unreliability.
- Enterprise AI platforms require robust resource management.
In practice
- Monitor token usage closely from initial deployment.
- Implement tiered access or hard caps early.
- Evaluate AI tool utility before widespread adoption.
Topics
- Generative AI
- Token Management
- Ask Sage Platform
- Large Language Models
- Department of Defense
- AI Resource Allocation
Best for: CTO, VP of Engineering/Data, Executive, Director of AI/ML, Policy Maker, Consultant
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by AI - Ars Technica.