New Study Suggests ChatGPT Is Getting Dumber—But, Is It?

Saturday, July 22, 2023, 11:27:00 AM

A new study released on Tuesday by researchers from Stanford University and University of California, Berkeley has ignited debate within the AI community regarding the performance of OpenAI’s GPT-4 language model.

The paper, titled “How Is ChatGPT’s Behavior Changing over Time?” and published on arXiv by Lingjiao Chen, Matei Zaharia, and James Zou, investigates changes in GPT-4’s outputs over a span of a few months, suggesting a potential decline in coding and compositional task abilities.

The study utilized API access to test GPT-3.5 and GPT-4 versions from March and June 2023 on various tasks, including math problem-solving, answering sensitive questions, code generation, and visual reasoning.

Notably, the research found a significant drop in GPT-4’s ability to identify prime numbers, plummeting from an accuracy of 97.6 percent in March to just 2.4 percent in June. Surprisingly, GPT-3.5 displayed improved performance during the same period.

GPT-4 is getting worse over time, not better.

Many people have reported noticing a significant degradation in the quality of the model responses, but so far, it was all anecdotal.

But now we know.

At least one study shows how the June version of GPT-4 is objectively worse than… pic.twitter.com/whhELYY6M4
— Santiago (@svpino) July 19, 2023

This investigation comes amidst growing concerns expressed by users who have observed a subjective decline in GPT-4’s performance over the past few months. Speculations about the reasons behind this decline abound, including OpenAI’s possible distillation of models to enhance efficiency, fine-tuning to mitigate harmful outputs, and unfounded conspiracy theories suggesting a reduction in GPT-4’s coding capabilities to promote GitHub Copilot usage.

One Response

Wolf1 says:

July 24, 2023 10:40 AM at 10:40 am

This issue isn’t even dumber or smarter – the issue is the performance is changing radically even in a short time. Inconsistency directly correlates to unreliability, and this isn’t a GIGO situation so much as GO potentially happening any time for any reason.

Silver Is in a New Price Regime, and the Market Isn’t Used to It | Keith Neumeyer – First Majestic

Agnico Eagle Just Made a Massive Gold Land Grab

A Copper-Gold Deposit Caught the White House’s Attention | Rob McLeod – Cambria Gold

Recommended

Mercado Drills 256 g/t Silver Over 6.5 Metres In First Drill Hole of Inaugural Program

Antimony Resources Drills 4.38% Sb Over 7.05 Metres At Bald Hill In Final Hole Of 2025 Program

Trending

Are Humans Doomed to Destroy ChatGPT?

Humanity has always been afraid of artificial intelligence. We recognize its ability to fundamentally transform...

Monday, February 6, 2023, 02:17:00 PM

ChatGPT Ads Launch as OpenAI Burns Through Billions

OpenAI announced last week that it will begin testing advertisements in ChatGPT’s free tier and...

Tuesday, January 20, 2026, 12:07:00 PM

Are Hard-Working Content-Farmers being replaced by AI?

And what effect could the monoculture they create have on our information supply? Our last...

Wednesday, May 10, 2023, 03:03:00 PM

OpenAI’s Latest Chatbot Aces IQ Test, Now Smarter than 9 out of 10 People

OpenAI’s latest model o1 just passed the Norwegian Mensa IQ test, achieving a score that...

Tuesday, September 17, 2024, 12:58:26 PM

xAI Seeks Emergency Order to Block Ex-Employee From Starting at OpenAI

Elon Musk’s xAI has filed a federal lawsuit against a former engineer, alleging he stole...

Tuesday, September 2, 2025, 03:41:00 PM