The AI landscape is shifting dramatically. In a single day, we witnessed the release of GPT-4.0 ImageGen, DeepSeek v3, and Gemini 2.5 Pro—with Google boldly claiming their new model is the best AI language model available. After extensive testing, I’ve found something more interesting than individual model capabilities: we’re witnessing the commoditization of AI intelligence.
Microsoft’s CEO recently claimed that AI models are becoming commodities with similar performance levels, suggesting that labs like OpenAI are merely “product companies selling an experience” rather than holding the secret to artificial general intelligence (AGI). This statement represents a massive shift from Microsoft’s previous celebration of its “special partnership” with OpenAI.
So does today’s news about Gemini 2.5 and DeepSeek v3 prove there’s no secret sauce to AI intelligence anymore? Let’s examine the evidence.
Benchmark Convergence: The New Reality
Looking at benchmark results across models reveals a striking pattern of convergence. On “humanity’s last exam”—a knowledge-intensive benchmark testing obscure trivia and complex translations—Gemini 2.5 appears to know the most. For incredibly difficult science questions, Gemini 2.5 Pro performs roughly on par with Claude 3.7 Sonnet and GPT-4o.
Direct comparisons are increasingly difficult (and perhaps deliberately so). Some companies use majority voting for benchmark scores, others don’t report benchmarks where they perform poorly, and some figures include tool use while others don’t. Despite these inconsistencies, performance is converging for models using similar computational resources.
This doesn’t mean progress has stopped—quite the opposite. Models are improving rapidly, but they’re improving together. The new Gemini 2.5 Pro excels at reading tables and charts, achieving near-human performance on the VISTA benchmark. It also handles an impressive million tokens of context, far beyond competitors.
The DeepSeek Challenge
Perhaps the most compelling evidence for AI convergence comes from DeepSeek v3, announced the same day. This Chinese company’s new base model performs comparably to GPT-4.5 from OpenAI—better in mathematics, slightly worse in science and general knowledge.
This is remarkable considering OpenAI was supposedly 6-12 months ahead of Chinese companies. The performance gap has essentially disappeared, challenging the notion that any single company has a significant technological moat in AI reasoning capabilities.
The primary differentiator now appears to be computational resources—how much money companies can spend on training larger models—rather than proprietary breakthroughs.
Microsoft’s Internal Efforts
Recent reports suggest Microsoft’s internal AI team, led by former Inflection AI head Mustafa Suleiman, has developed models performing nearly as well as those from OpenAI and Anthropic. This likely explains Satya Nadella’s confidence in declaring AI models commoditized.
An interesting anecdote: Suleiman reportedly became angry when OpenAI wouldn’t share how they programmed GPT-4o to “think” before answering queries. Microsoft claims they’ve since figured it out independently, suggesting these techniques are discoverable by multiple teams with sufficient resources.
The Reality Behind the Hype
While companies make bold claims about AI capabilities, reality often tells a different story. Anthropic’s CEO recently predicted AI would write “essentially all code” within 12 months. Yet they continue advertising high-salary software engineering positions with annual contracts. If their prediction were true, wouldn’t these roles become obsolete?
Similarly, for all the advanced capabilities these models demonstrate on benchmarks, they still struggle with relatively simple tasks like playing Pokémon—getting stuck and failing to progress despite their supposed reasoning abilities.
The current state of AI reveals three key insights:
- Performance is converging across model families using similar computational resources
- The primary differentiator is increasingly just raw compute power (money)
- Real-world capabilities still lag significantly behind benchmark performance and company hype
This doesn’t diminish the impressive progress we’re seeing. Gemini 2.5 Pro represents a significant advancement, particularly in visual understanding and context length. But it’s part of a broader trend where multiple companies achieve similar capabilities nearly simultaneously.
The era of any single company having a decisive lead in AI intelligence appears to be ending. Instead, we’re entering a phase where the real competition will be about product integration, user experience, and specialized applications—not raw intelligence.
For users, this is ultimately good news. Competition drives innovation and keeps prices in check. As AI capabilities become more standardized, the focus will shift to how these tools can best serve our needs rather than which model scores highest on obscure benchmarks.
Frequently Asked Questions
Q: What makes Gemini 2.5 Pro stand out from other AI models?
Gemini 2.5 Pro’s most distinctive feature is its exceptional context length capability, handling up to one million tokens (approximately 750,000 words). It also excels at visual understanding tasks, particularly interpreting tables, charts, and counting objects in images, where it approaches human-level performance.
Q: Are AI models from different companies becoming more similar in performance?
Yes, there’s a clear convergence in performance across models using similar computational resources. While each model may have specific strengths (like Claude’s writing style or DeepSeek’s cost efficiency), overall capabilities on standard benchmarks are increasingly comparable, suggesting the technical approaches are becoming standardized.
Q: What does the “commoditization of AI” mean for the industry?
As AI models converge in capabilities, the primary differentiator becomes computational resources rather than proprietary algorithms. This shifts competition toward user experience, specialized applications, and integration with existing products. For consumers, this likely means more choices at competitive prices as companies focus on differentiation beyond raw intelligence.
Q: How close are current AI models to artificial general intelligence (AGI)?
Despite impressive benchmark scores, current models still struggle with relatively simple real-world tasks and lack true understanding. The gap between benchmark performance and practical capabilities suggests we remain far from AGI. The convergence in model performance also indicates that no single company has discovered a secret path to general intelligence.
Q: Will AI replace programmers in the near future as some CEOs predict?
Despite bold predictions about AI writing “all code” within a year, the reality is more nuanced. Current models can generate code but struggle with complex programming tasks. Companies making these claims continue hiring software engineers at high salaries, suggesting they don’t fully believe their own predictions. Programming involves much more than code generation, including problem definition, architecture design, and system integration.








