Search

Current Challenges of Large AI Language Models

07 May 2025

Modern large AI language models, particularly OpenAI's o3 model, are found to be less accurate than their predecessors. This is reported by The New York Times, referencing multiple studies.

Accuracy issues are also observed in AI models from other major companies like Google and the Chinese startup DeepSeek. Despite improvements in mathematical capabilities, the number of errors in responses continues to rise.

One of the primary issues is the phenomenon of "hallucinations," where models fabricate information without proper sources. Amr Awaadalla, CEO of the startup Vectara, notes that these issues will persist despite developers' efforts.

For instance, a support bot for the Cursor tool incorrectly stated that the tool could only be used on one device, leading to user complaints. It was later revealed that the company had made no such changes — it was all a fabrication by the bot.

In tests of various models, the rate of fabricated facts reached 79%. Internal testing of the o3 model revealed a 33% hallucination rate in responses about celebrities, double that of the previous model o1. The new 04-mini model showed even worse results, with 48% errors.

When answering general questions, the hallucination rates for models o3 and o4-mini were even higher at 51% and 79% respectively, compared to the older model o1, which fabricated facts 44% of the time. OpenAI acknowledges the need for further studies to understand the causes of these errors.

Independent research confirms that hallucinations are also characteristic of reasoning-enabled models from Google and DeepSeek. Vectara's own study found that such models fabricate facts at least 3% of the time, with some instances reaching up to 27%. Despite companies' efforts, the hallucination rate has only decreased by 1-2% over the past year.