The AI landscape just experienced a seismic shift. OpenAI has released two new models—O3 and O4 Mini—that are redefining what we expect from artificial intelligence. Having spent time testing these models since their release, I’m convinced we’re witnessing a significant leap forward in AI capabilities, particularly in how these models use tools and reason through problems.
What makes these models special isn’t just their raw intelligence but their ability to think through problems step by step, using tools when needed. It feels like having a research assistant who can search the web, run code, create visualizations, and analyze data—all within a single conversation.
The New Models: What Sets Them Apart
O3 is OpenAI’s new flagship model—larger, more capable, and designed to push boundaries in coding, math, science, and visual perception. Meanwhile, O4 Mini offers a glimpse into the future: it’s smaller and more cost-efficient than O3 but surprisingly powerful for its size.
The most striking feature of both models is their “agentic” quality. They don’t just respond to prompts; they actively work through problems:
- They analyze what you’re asking
- Develop a plan to solve it
- Use tools like web search, Python coding, or image manipulation
- Present findings with visualizations when helpful
This approach produces results that feel more thorough and trustworthy than previous AI generations. When I asked O3 to analyze whether the NBA should move its corner three-point line, it searched for data, ran calculations, created visualizations, and delivered a recommendation based on evidence—not just opinions.
Benchmark Performance: Impressive But Context Matters
On paper, these models are setting new records. O3 and O4 Mini achieve near-perfect scores on benchmarks like AIM 2024 and AIM 2025 when using tools. The improvement is dramatic—O4 Mini with Python access scores 99.5% on AIM 2024, compared to around 70% for the original GPT-4.
But here’s where things get interesting. When comparing these models to Google’s Gemini 2.5 Pro, the performance gap narrows significantly:
- On AIM benchmarks, OpenAI’s models lead by small margins
- On GPQA Diamond, Gemini 2.5 Pro actually outperforms both new models
- On Humanity’s Last Exam, O3 edges out Gemini 2.5 Pro, but O4 Mini falls behind
The real differentiator might be cost. O4 Mini is remarkably affordable at around $0.50 per million output tokens, compared to $10 for Gemini 2.5 Pro and a whopping $40 for O3. For many applications, O4 Mini offers the best balance of performance and cost.
The Multimodal Revolution
What truly sets these models apart is their ability to “think with images.” They don’t just describe what they see—they can zoom in on specific parts of an image, highlight relevant sections, and integrate visual information into their reasoning process.
In my testing, I uploaded a whiteboard photo with small text, and O3 automatically zoomed in to read it, then incorporated that information into its response. This feels like a genuine breakthrough in how AI interacts with visual content.
These models can integrate images directly into their chain of thought. They don’t just see an image—they actually think with it.
This capability opens up new possibilities for analyzing documents, diagrams, charts, and even handwritten notes that previous models struggled with.
The Coding Companion
OpenAI also released Codeex CLI, an open-source tool that lets these models interact directly with your computer for coding tasks. This appears to be their answer to Anthropic’s Claude Coder, but with the advantage of being open source.
While the models excel at generating shorter code snippets and solving specific problems, they still have limitations. My attempts to get O3 to create a complete Minecraft clone in three.js were unsuccessful—the model seemed restricted in how much code it could generate at once in the ChatGPT interface.
For smaller coding tasks, however, both models perform admirably, writing clean, efficient code and explaining their reasoning along the way.
The Future of AI Assistance
What we’re seeing with O3 and O4 Mini feels like a preview of truly agentic AI—systems that can take a high-level request and break it down into the necessary steps to solve it, using whatever tools are available.
The models still have limitations. Their context windows are limited to 200,000 tokens (far less than the million tokens some competitors offer), and they can’t generate images like GPT-4o can. But their reasoning capabilities and tool use represent a significant step forward.
As AI continues to evolve, the distinction between a chatbot and an assistant is becoming clearer. These new models don’t just answer questions—they solve problems. And that makes them valuable in ways their predecessors weren’t.
Whether these advances justify the cost depends on your needs. For many users, O4 Mini will provide the best value, offering much of O3’s capability at a fraction of the price. But for those working on cutting-edge research or complex analytical tasks, O3’s additional capabilities may well be worth the premium.
The AI race is accelerating, and OpenAI has just raised the bar again. The question now is how quickly competitors will respond—and what capabilities the full O4 model will bring when it eventually arrives.
Frequently Asked Questions
Q: How do O3 and O4 Mini compare to Google’s Gemini 2.5 Pro?
Performance-wise, they’re very close. O3 and O4 Mini slightly outperform Gemini 2.5 Pro on some benchmarks while falling slightly behind on others. The biggest difference is in pricing: O4 Mini is significantly cheaper than both Gemini 2.5 Pro and O3, making it potentially the best value option for many users.
Q: What makes these models “agentic”?
These models can break down complex problems into steps, decide which tools they need (like web search or code execution), and work through solutions methodically. They don’t just respond to prompts; they actively work to solve problems by choosing appropriate actions and tools based on the task at hand.
Q: Are these models available to everyone?
Currently, O3 and O4 Mini are available to ChatGPT Plus subscribers, with access rolling out to Pro, Teams, and Edu users shortly after. Free users will need to wait longer for access. However, some platforms like Windsurf are offering limited-time free access to O4 Mini for coding tasks.
Q: What’s the most impressive capability of these new models?
Their ability to work with images stands out as particularly impressive. They can analyze images as part of their reasoning process, zoom in on specific details, and extract information from visual content in ways previous models couldn’t. This makes them much more effective for tasks involving documents, diagrams, or other visual information.
Q: How do these models handle coding tasks?
Both models show significant improvements in coding ability, with O3 and O4 Mini achieving high scores on coding benchmarks. They’re particularly good at understanding coding problems, debugging issues, and writing efficient solutions. However, they still have limitations when generating very large amounts of code in the ChatGPT interface.








