The AI world is buzzing with excitement over OpenAI’s latest releases: GPT-4o Mini and GPT-3.5. While social media feeds fill with breathless proclamations of “AGI is here!” and “hallucination-free AI,” I’ve spent time testing both models extensively, and the reality deserves a more measured assessment.
These models represent significant progress—they’re undeniably better than their predecessors. But they’re not the superhuman intelligences some are claiming. Let me explain why.
The Reality Behind the Hype
OpenAI has a pattern of giving early access to people who will generate maximum excitement about their new models. This creates an echo chamber of hype that doesn’t always match reality. While GPT-4o Mini and GPT-3.5 are impressive improvements over GPT-4, claims that they represent Artificial General Intelligence (AGI) are premature.
For me, AGI means a model that can perform better than average humans across most tasks humans can do. These models excel at knowledge, coding, and mathematics compared to average humans, but they still make basic errors that even a moderately intelligent person wouldn’t.
For example, when I asked GPT-3.5 to count the number of intersection points between five lines, it confidently stated there were eight distinct points. This was simply wrong. In another test, when asked about a glove falling from a car trunk while crossing a bridge, it often concluded the glove would fall into the river below—completely forgetting about the bridge itself.
Would you consider someone “above genius level” if they couldn’t grasp that a glove falling from a car on a bridge would land on the bridge, not in the water below?
Benchmark Performance: Impressive But Costly
On benchmarks, these models show impressive capabilities:
- GPT-4o Mini scored 4/10 on my public benchmark questions—solid for a smaller model
- GPT-3.5 achieved 6/10 on the same test—the first OpenAI model to do so
- Both models excel at competitive mathematics and coding tasks
- For PhD-level science, GPT-3.5 scores 83.3% and GPT-4o Mini 81.4%
For comparison, Google’s Gemini 2.5 Pro scores 84% on PhD-level science benchmarks in a single attempt. This raises an interesting question: why wasn’t Gemini’s release met with the same “AGI is here” fanfare?
The pricing is also worth noting. Gemini 2.5 Pro is roughly 3-4 times cheaper than GPT-3.5. While OpenAI’s models might edge out competitors on some benchmarks, they don’t take the cost-effective lead.
Not “Hallucination-Free” as Claimed
OpenAI’s own release notes contradict some of the hype. They mention that “evaluations by external experts” found GPT-3.5 making “20% fewer major errors” than previous models. That’s great progress, but it directly contradicts claims of being “hallucination-free.”
If I were an average professional who saw Sam Altman retweet that a new model was “hallucination-free,” I’d be making dangerous assumptions about its reliability. The truth is these models still make significant errors, just fewer than before.
The Capabilities Gap
There are still notable limitations compared to competitors. For instance, Gemini 2.5 Pro can analyze YouTube videos and raw video content, while GPT-3.5 can only examine video metadata. This multimodal capability gap is significant for many use cases.
I was impressed with how GPT-3.5 analyzed my benchmark website, created a cover image, and provided nuanced advice about the benchmark’s limitations. But there’s another detail worth noting: OpenAI admitted that the GPT-3.5 version that crushed benchmarks months ago was “benchmark optimized”—meaning it had more compute time than the version being released to the public.
The AGI Question
My simplest definition of AGI is: Would I hire this AI over a smart human being? Could GPT-4o edit an entire video without random glitches? Could it do my Amazon shopping without putting me thousands in debt?
These models are incredible at drafting content that seems intelligent—and often is. They’re far smarter than me in many domains. I couldn’t score anywhere near them on coding or PhD-level exams.
But comparing them to human IQ is misleading. How many brilliant coders or PhDs do you know who can’t count intersections between lines or forget about a bridge when solving a simple physics problem?
Progress is happening, but we’re not at AGI yet. I’m not an AGI denialist—I believe it’s coming in the next few years. But today’s models, impressive as they are, don’t meet that threshold.
Frequently Asked Questions
Q: Are GPT-4o Mini and GPT-3.5 truly “hallucination-free” as some claim?
No, they are not hallucination-free. OpenAI’s own documentation states that external evaluations found GPT-3.5 makes “20% fewer major errors” than previous models, which directly acknowledges that errors still occur. The models have improved significantly but still make factual mistakes and logical errors.
Q: How do these models compare to Google’s Gemini 2.5 Pro in terms of cost and performance?
While OpenAI’s new models may slightly outperform Gemini 2.5 Pro on certain benchmarks, Gemini is approximately 3-4 times less expensive to use. For PhD-level science benchmarks, Gemini scores 84% on a single attempt compared to GPT-3.5’s 83.3%, making Gemini more cost-effective for many applications despite marginally lower performance in some areas.
Q: What does AGI (Artificial General Intelligence) actually mean, and are we there yet?
AGI refers to AI that can perform most tasks at or above average human level. While these models excel at knowledge, coding, and mathematics compared to average humans, they still make basic logical errors that most humans wouldn’t. A practical definition might be: would you hire this AI over a smart human for complex real-world tasks? Currently, the answer is still no for many situations, suggesting we haven’t reached true AGI yet.
Q: What are some key limitations of GPT-4o Mini and GPT-3.5 compared to competitors?
One notable limitation is multimodal capability. While Gemini 2.5 Pro can analyze YouTube videos and raw video content, GPT-3.5 can only examine video metadata. Additionally, OpenAI admitted that the version of GPT-3.5 that performed exceptionally well on benchmarks months ago was “benchmark optimized” with more compute time than the public release version.
Q: How do these models perform on your benchmark tests?
GPT-4o Mini scored 4/10 on my public benchmark questions, which is solid for a smaller model. GPT-3.5 achieved 6/10 on the same test, making it the first OpenAI model to reach that score. While impressive, both models still make fundamental errors on relatively simple reasoning tasks, such as the bridge-and-glove scenario where they often fail to consider that objects falling from a car on a bridge would land on the bridge itself.








