OpenAI’s GPT-4o Mini and GPT-3.5: Impressive, But Let’s Cut Through the Hype

REPURPOSE SOCIAL POSTS INTO CONTENT MARKETING

Create content 10x faster while staying authentic to your brand.

The AI world is buzzing with excitement over OpenAI’s latest releases: GPT-4o Mini and GPT-3.5. While social media feeds fill with breathless proclamations of “AGI is here!” and “hallucination-free AI,” I’ve spent time testing both models extensively, and the reality deserves a more measured assessment.

These models represent significant progress—they’re undeniably better than their predecessors. But they’re not the superhuman intelligences some are claiming. Let me explain why.

The Reality Behind the Hype

OpenAI has a pattern of giving early access to people who will generate maximum excitement about their new models. This creates an echo chamber of hype that doesn’t always match reality. While GPT-4o Mini and GPT-3.5 are impressive improvements over GPT-4, claims that they represent Artificial General Intelligence (AGI) are premature.

For me, AGI means a model that can perform better than average humans across most tasks humans can do. These models excel at knowledge, coding, and mathematics compared to average humans, but they still make basic errors that even a moderately intelligent person wouldn’t.

For example, when I asked GPT-3.5 to count the number of intersection points between five lines, it confidently stated there were eight distinct points. This was simply wrong. In another test, when asked about a glove falling from a car trunk while crossing a bridge, it often concluded the glove would fall into the river below—completely forgetting about the bridge itself.

Would you consider someone “above genius level” if they couldn’t grasp that a glove falling from a car on a bridge would land on the bridge, not in the water below?

Benchmark Performance: Impressive But Costly

On benchmarks, these models show impressive capabilities:

  • GPT-4o Mini scored 4/10 on my public benchmark questions—solid for a smaller model
  • GPT-3.5 achieved 6/10 on the same test—the first OpenAI model to do so
  • Both models excel at competitive mathematics and coding tasks
  • For PhD-level science, GPT-3.5 scores 83.3% and GPT-4o Mini 81.4%
See also  Microsoft 5.4 Release Review Shows Mixed Performance Against Leading AI Models

For comparison, Google’s Gemini 2.5 Pro scores 84% on PhD-level science benchmarks in a single attempt. This raises an interesting question: why wasn’t Gemini’s release met with the same “AGI is here” fanfare?

The pricing is also worth noting. Gemini 2.5 Pro is roughly 3-4 times cheaper than GPT-3.5. While OpenAI’s models might edge out competitors on some benchmarks, they don’t take the cost-effective lead.

Not “Hallucination-Free” as Claimed

OpenAI’s own release notes contradict some of the hype. They mention that “evaluations by external experts” found GPT-3.5 making “20% fewer major errors” than previous models. That’s great progress, but it directly contradicts claims of being “hallucination-free.”

If I were an average professional who saw Sam Altman retweet that a new model was “hallucination-free,” I’d be making dangerous assumptions about its reliability. The truth is these models still make significant errors, just fewer than before.

The Capabilities Gap

There are still notable limitations compared to competitors. For instance, Gemini 2.5 Pro can analyze YouTube videos and raw video content, while GPT-3.5 can only examine video metadata. This multimodal capability gap is significant for many use cases.

I was impressed with how GPT-3.5 analyzed my benchmark website, created a cover image, and provided nuanced advice about the benchmark’s limitations. But there’s another detail worth noting: OpenAI admitted that the GPT-3.5 version that crushed benchmarks months ago was “benchmark optimized”—meaning it had more compute time than the version being released to the public.

The AGI Question

My simplest definition of AGI is: Would I hire this AI over a smart human being? Could GPT-4o edit an entire video without random glitches? Could it do my Amazon shopping without putting me thousands in debt?

See also  Google's New AI Mode Blurs the Lines Between Search and Chatbots

These models are incredible at drafting content that seems intelligent—and often is. They’re far smarter than me in many domains. I couldn’t score anywhere near them on coding or PhD-level exams.

But comparing them to human IQ is misleading. How many brilliant coders or PhDs do you know who can’t count intersections between lines or forget about a bridge when solving a simple physics problem?

Progress is happening, but we’re not at AGI yet. I’m not an AGI denialist—I believe it’s coming in the next few years. But today’s models, impressive as they are, don’t meet that threshold.


Frequently Asked Questions

Q: Are GPT-4o Mini and GPT-3.5 truly “hallucination-free” as some claim?

No, they are not hallucination-free. OpenAI’s own documentation states that external evaluations found GPT-3.5 makes “20% fewer major errors” than previous models, which directly acknowledges that errors still occur. The models have improved significantly but still make factual mistakes and logical errors.

Q: How do these models compare to Google’s Gemini 2.5 Pro in terms of cost and performance?

While OpenAI’s new models may slightly outperform Gemini 2.5 Pro on certain benchmarks, Gemini is approximately 3-4 times less expensive to use. For PhD-level science benchmarks, Gemini scores 84% on a single attempt compared to GPT-3.5’s 83.3%, making Gemini more cost-effective for many applications despite marginally lower performance in some areas.

Q: What does AGI (Artificial General Intelligence) actually mean, and are we there yet?

AGI refers to AI that can perform most tasks at or above average human level. While these models excel at knowledge, coding, and mathematics compared to average humans, they still make basic logical errors that most humans wouldn’t. A practical definition might be: would you hire this AI over a smart human for complex real-world tasks? Currently, the answer is still no for many situations, suggesting we haven’t reached true AGI yet.

See also  OpenAI's 0101 Pro Mode Falls Short of Revolutionary Claims

Q: What are some key limitations of GPT-4o Mini and GPT-3.5 compared to competitors?

One notable limitation is multimodal capability. While Gemini 2.5 Pro can analyze YouTube videos and raw video content, GPT-3.5 can only examine video metadata. Additionally, OpenAI admitted that the version of GPT-3.5 that performed exceptionally well on benchmarks months ago was “benchmark optimized” with more compute time than the public release version.

Q: How do these models perform on your benchmark tests?

GPT-4o Mini scored 4/10 on my public benchmark questions, which is solid for a smaller model. GPT-3.5 achieved 6/10 on the same test, making it the first OpenAI model to reach that score. While impressive, both models still make fundamental errors on relatively simple reasoning tasks, such as the bridge-and-glove scenario where they often fail to consider that objects falling from a car on a bridge would land on the bridge itself.

 

About ArticleX

ArticleX is the leading content automation platform. Our expert staff writes about our tool, marketing automation, and the state of AI. The startup is dedicated to providing experts insights and useful guides to a larger audience.

If you have questions or concerns about an article, please contact [email protected]

Learn more.