ChatGPT can get better on benchmarks and worse at a task you rely on. Both can be true. A newer model might solve harder coding problems while producing writing you dislike, overlooking an instruction or taking more corrections to finish familiar work. The evidence does not support a simple, continuous decline. It does show that updates can introduce real regressions—and that stronger overall scores do not guarantee a better everyday experience.
Updated September 20, 2026. This article separates independent evaluations, OpenAI's own reports and user-experience complaints.
Quick answer: ChatGPT is improving on some measured capabilities, but it is not becoming uniformly better for every user. Judge an update by factual accuracy, instruction following and the amount of correction your actual work requires. There is no single reliable percentage describing how often all ChatGPT answers are wrong.
Table of Contents
- Is ChatGPT Getting Better or Worse?
- What the Data Actually Shows
- How Often Is ChatGPT Wrong?
- Why ChatGPT May Feel Worse
- Common Complaints and What to Check
- Writing, Coding and Research: Judge Them Separately
- How to Test Whether It Has Gotten Worse for You
- How to Get Better Answers
- When to Try Another Tool
- The Verdict
- Frequently Asked Questions
- Sources and Methodology
Is ChatGPT Getting Better or Worse?
The answer depends on what you measure. Solving a difficult problem, reliably following a brief and writing in a voice you enjoy are different abilities. An update can improve one without improving the others.
Imagine that your old workflow produced a usable newsletter after one edit. A newer model scores higher on a coding test but now needs five corrections to stop using stock phrases. For your newsletter, that is a practical decline. It does not establish that the model has become worse at coding or mathematics.
The reverse also matters: one bad answer does not prove a general regression. To establish a trend, compare repeated attempts on similar tasks under comparable conditions.
What the Data Actually Shows
September 2026: Measurable Gains, With Some Trade-offs
In its September 9 evaluation, Artificial Analysis reported that GPT-6 Astra gained six points over GPT-5.6 Sol on its Intelligence Index. But it also reported a roughly 45-point Elo decline on GDPval-AA v2, an evaluation of professional tasks, and lower presentation-quality results in AA-Briefcase.
These are results from the configurations and evaluation environments tested, not a score for every ChatGPT conversation. They support a more useful conclusion than either “everything improved” or “everything got worse”: improvements can coexist with weaknesses on particular tasks.
OpenAI's Factuality Tests Measure a Different Question
The GPT-5.6 system card reports slightly fewer factual errors for Sol than GPT-5.5, with a larger reduction in reproducing specific errors users had flagged. OpenAI explicitly says these examples were selected for being error-prone. They are not a representative sample of everyday conversations.
That distinction prevents an easy mistake: a result on difficult, previously failed questions cannot tell you the overall likelihood that an ordinary answer will be wrong.
A Bad Behavioral Update Has Been Documented
In April 2025, OpenAI rolled back a GPT-4o update after it made responses excessively agreeable and flattering. This behavior is called sycophancy. The episode establishes that an update can make the product less useful even when it was intended to improve the experience.
It also gives user reports an appropriate role: complaints can reveal a specific failure worth investigating. Their existence alone does not establish how common that failure is across all users.
The Famous Stanford/Berkeley Study Is Historical Evidence
A 2023 study by Lingjiao Chen, Matei Zaharia and James Zou compared March and June versions of GPT-3.5 and GPT-4. It found changes in instruction following, code formatting and other tasks, with improvements in some areas and declines in others.
This demonstrates why ongoing evaluation matters. It does not measure the quality of ChatGPT in September 2026. Presenting a three-year-old result as proof of a current collapse would be misleading.
How Often Is ChatGPT Wrong?
The sources reviewed here do not establish a single error rate for all ChatGPT use. A useful percentage must identify the model, task, tools, settings and grading method. It must also explain what counts as an error: one incorrect detail, a wholly wrong answer, an unsupported citation or failure to follow instructions.
Checking a supplied invoice is different from recalling an obscure historical fact. A sourced answer can still misread its source. A polished paragraph can contain one consequential mistake. Combining these into one accuracy number hides the distinctions a reader needs.
What Did the “86% Hallucination Rate” Actually Mean?
Artificial Analysis's April 2026 GPT-5.5 evaluation reported 57% accuracy and an 86% hallucination rate for its xhigh configuration on AA-Omniscience. Those percentages use different denominators.
The benchmark's published methodology defines accuracy across all questions. Its hallucination rate measures incorrect answers among responses that were not correct, including partial answers and questions the model did not attempt:
Hallucination rate = incorrect ÷ (incorrect + partial answers + not attempted).
For illustration, if a model answers 60 of 100 questions correctly, answers 30 incorrectly and declines 10, its accuracy is 60%, but this hallucination rate is 75%: 30 divided by 40. This is a hypothetical example, not a ChatGPT test result.
So the reported 86% did not mean that 86% of ordinary ChatGPT answers were false. It highlighted a tendency to give incorrect answers rather than withhold an answer when the model could not answer correctly.
In September, Artificial Analysis reported a decrease in this metric from 92% for GPT-5.6 Sol to 51% for GPT-6 Astra at max effort, alongside improved accuracy. That is encouraging within this test. It still does not supply an everyday ChatGPT error rate.
For more on fabricated claims and references, see What Is a Hallucination in AI?
Why ChatGPT May Feel Worse
The Model or Experience You Preferred Changed
ChatGPT is a changing service. OpenAI's retirement announcement set February 13, 2026 as the removal date for GPT-4o and several older models from ChatGPT. It acknowledged that some users preferred GPT-4o's conversational warmth and creative ideation.
A preference for that style is not a misunderstanding of benchmarks. Tone, pacing and creative choices are part of the product's usefulness.
More Capable Does Not Necessarily Mean Easier to Direct
OpenAI's GPT-6 Astra developer guidance describes behavior that can create friction: asking questions when users expect the model to proceed, sensitivity to instructions in supporting files, and recurring phrasing. This is model guidance for developers, not proof that every ChatGPT user experiences these problems.
It illustrates why “intelligence” is too broad a diagnosis. The failure might concern initiative, competing instructions or style rather than inability to understand the subject.
Your Conversation May Contain Competing Directions
Consider a chat where you first requested a short summary, later asked for a detailed report and then returned to an earlier draft. If the answer keeps reverting to the short version, a clean conversation containing only the current brief is a useful diagnostic test.
This is a possible explanation, not an excuse for ignoring a clear correction. If the failure persists with a short, unambiguous prompt, record it as an instruction-following problem.
Your Standard for a Useful Answer May Have Changed
You may now notice errors you missed when the tool was new, or ask it to do more demanding work. That does not invalidate frustration. It means a fair comparison should use the same task, not compare today's difficult project with a memorable easy answer from last year.
Common ChatGPT Complaints and What to Check
The following are troubleshooting categories, not a survey of how frequently users encounter each problem.
| Complaint | Useful check | What it tells you |
|---|---|---|
| It ignores my corrections. | Try the same task in a new chat with three explicit requirements. | A repeated failure is stronger evidence than one confused conversation. |
| It repeats itself or sounds generic. | Provide a short style example and identify the exact unwanted pattern. | You can separate a style mismatch from missing knowledge. |
| Answers are too short. | Specify required sections, examples and completeness. | Check whether it supplies the requested substance, not just more words. |
| It confidently invents facts. | Open the cited source and locate support for the claim. | A real link does not guarantee the claim is supported. |
| It agrees with a false premise. | Ask it to assess the premise against a supplied source. | Agreement is not independent verification. |
| It refuses a legitimate task. | Clarify the purpose and exact scope without changing the underlying task. | A refusal and an incorrect answer are different failure types. |
| It gets worse halfway through a project. | Restart from a concise record of current decisions. | This tests whether accumulated context contributes to the problem. |
Writing, Coding and Research: Judge Them Separately
Writing and Editing
Evaluate whether the response preserves your meaning, follows your voice and respects the brief. A fluent rewrite that removes a crucial qualification is not an improvement. Save examples of writing you liked so you can compare concrete outputs rather than impressions.
Coding
Use working behavior as the test: does the change solve the problem, fit the existing code and pass relevant checks? A confident explanation of untested code should not count as success. Compare similar tasks in the same environment before concluding that coding performance has broadly declined.
Research
Check whether each important source exists, supports the attached claim and is current enough for the question. Ask for facts and interpretations to be separated. Neither a long reference list nor another chatbot's agreement independently validates an answer.
Everyday Conversation
A conversational tool can become less enjoyable without becoming worse at formal reasoning. If warmth, brevity or imaginative discussion is your main use, include that preference in your evaluation rather than treating it as irrelevant.
How to Test Whether ChatGPT Has Gotten Worse for You
Try this small comparison before repeatedly changing prompts or abandoning a workflow:
- Choose five to ten representative tasks. Use work you actually repeat, such as editing a paragraph, extracting information or fixing a small bug.
- Define success first. Write down the facts, format and constraints each answer must satisfy.
- Keep conditions comparable. Record the date, visible model or mode, available tools, source files and relevant custom instructions.
- Repeat each task in fresh chats. Three attempts can expose variability, although this remains a small personal test.
- Track correction effort. Count factual errors, missed requirements, follow-up prompts and minutes spent repairing the output.
- Compare with saved outputs or another available model. If you cannot access the old model or its original inputs, acknowledge that limitation.
This will not establish a worldwide trend. It can establish whether the tool is saving you less time on your own work—the question that matters when deciding how to use it.
How to Get Better Answers From ChatGPT
Start with a brief that describes the finished result. Include the audience, source material, non-negotiable requirements and an example where style matters. Remove old directions that no longer apply.
Try this prompt:
Complete the task using the material below. Follow these requirements: [list]. Preserve these facts: [list]. If a claim needs outside verification, identify it rather than inventing support. Make reasonable assumptions for minor choices and state them briefly. Ask a question only if the missing information would materially change the result. Before finishing, check the answer against the requirements.
For long work, use manageable stages such as source review, draft and final checks. For current factual questions, request current sources and inspect the important ones yourself. If the conversation keeps circling an abandoned approach, move the current brief into a fresh chat.
A self-check may catch omissions, but it is not independent fact-checking. Asking the same model whether it is certain is also not proof. Better instructions help expose failures; they do not remove the model's responsibility to follow a clear request or eliminate its limitations.
When to Try Another AI Tool
Try a second tool when the same well-defined task repeatedly fails, when you need a different writing style or when the correction work outweighs the time saved. Give both tools the same materials and success criteria.
Choose the output that actually meets your requirements, then verify important claims against the underlying sources. For factual disagreements, the deciding evidence should be the documents, calculations or tests—not a vote between chatbots.
The Verdict: Capability and Reliability Are Different Questions
The evidence supports neither a universal collapse nor a promise that every update is an upgrade for every user. Documented regressions exist, and some newer evaluations show gains alongside declines.
For a reader asking “Is ChatGPT getting worse?”, the useful test is concrete: Does it complete your work correctly, follow your instructions and require less repair? If those results deteriorate consistently, your workflow has become worse even if a headline benchmark improves.
Frequently Asked Questions
Is ChatGPT getting worse in 2026?
Not uniformly. The evaluations reviewed here show improvements in some capabilities and declines in others. A particular workflow can become less reliable without demonstrating that every use of ChatGPT has deteriorated.
Is ChatGPT getting better or worse at answering questions?
Specify the kind of question. A model's score on difficult knowledge tests does not establish its accuracy when summarizing your file or researching today's news. Judge the relevant task and check the supporting evidence.
How often is ChatGPT wrong?
There is no single rate established by these sources for all ChatGPT conversations. Error rates depend on the model, task, tools and definition of an error. Always check what a benchmark percentage actually measures.
Why does ChatGPT ignore instructions and repeat itself?
Possible contributors include competing directions, accumulated conversation context and model behavior. Try a fresh chat with a short, specific brief. If the same failure persists, treat it as a reproducible problem rather than assuming you need an increasingly elaborate prompt.
Does a confident answer mean ChatGPT checked the facts?
No. Confident wording is not evidence that a claim was verified. Ask for the supporting source, open it and check whether it actually supports the answer.
Should I pay for ChatGPT if the answers feel worse?
Base the decision on your own repeated tasks. Compare usable results and correction time with the tools you already have. A subscription is valuable to you only if the access and features you use justify the cost.
Sources and Methodology
This is an evidence review, not an original benchmark or a representative survey of ChatGPT users. Numerical results are attributed to the organizations that published them. User-experience examples and troubleshooting suggestions are illustrative, not estimates of how often problems occur.
Evaluations can differ in model settings, tools, task selection and grading. Compare models within the same reported evaluation; do not combine scores from different versions of a benchmark into a single trend line. API and coding-agent results should not be treated as direct measurements of every ChatGPT mode.
- Artificial Analysis: Benchmarking GPT-6 Astra, September 9, 2026
- Artificial Analysis: GPT-5.5 evaluation, April 23, 2026
- Artificial Analysis: AA-Omniscience definitions and methodology
- OpenAI: GPT-5.6 system card, factuality evaluation
- OpenAI: Sycophancy in GPT-4o
- OpenAI: Retirement of GPT-4o and older ChatGPT models
- OpenAI developer guidance: GPT-6 Astra behavior
- Chen, Zaharia and Zou: How Is ChatGPT's Behavior Changing Over Time? Revised October 2023