AI inference is obviously profitable
Summary
Contrary to claims that AI inference is unprofitable and relies on investor subsidies, analysis suggests it is demonstrably profitable. Frontier AI providers report 70-80% gross margins on inference. Detailed estimates for running a dense 70B model on four Nvidia A100 GPUs show power and cooling costs around 13 cents per hour, plus GPU amortization of \$1.80 per hour over five years, leading to an approximate total inference cost of one dollar per million tokens. This contrasts with OpenAI's GPT-5.4-mini charging \$4.50 per million tokens, making high profit margins plausible. Further evidence comes from open-weights Chinese LLMs like DeepSeek-R1, which claim over 80% profit margins despite API pricing less than half of major providers, with market costs for DeepSeek-V4-Pro around 87 cents per million output tokens. AI labs like OpenAI and Anthropic maintain high inference prices to subsidize their substantial model training costs, not due to inherent unprofitability of inference itself.
Key takeaway
For AI/ML Directors evaluating LLM deployment strategies, recognize that AI inference is fundamentally profitable, often with significant margins. You should scrutinize API pricing from frontier model providers, as their high costs frequently subsidize their training efforts, not reflect actual inference expenses. Consider leveraging open-weight models or building your own inference infrastructure to achieve substantial cost savings, potentially reducing per-token costs to under one dollar. This approach can free up budget for other strategic investments.
Key insights
AI inference is inherently profitable, with high margins often used by labs to fund model training.
Principles
- Inference costs are significantly lower than current API pricing.
- Open-weight models drive down inference market prices.
- AI labs subsidize training with inference profits.
In practice
- Evaluate open-weight LLMs for cost-effective inference.
- Compare API token pricing against subscription models.
- Consider GPU amortization over five years for cost planning.
Topics
- AI Inference Costs
- LLM Profit Margins
- GPU Amortization
- Open-weight LLMs
- DeepSeek AI
- AI Lab Business Models
Code references
Best for: Entrepreneur, CTO, VP of Engineering/Data, Director of AI/ML, Investor, Consultant
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by All posts - seangoedecke.com RSS feed.