5.1 is not good
Resolved 💬 11 comments Opened Nov 13, 2025 by perbinder Closed Nov 16, 2025
💡 Likely answer: A maintainer (etraut-openai, contributor)
responded on this thread — see the highlighted reply below.
What version of Codex is running?
0.58.0
What subscription do you have?
Teams
Which model were you using?
5.1
What platform is your computer?
Linux
What issue are you seeing?
019a7f29-4b96-7ef3-85c2-c218c93fb8a5
What steps can reproduce the bug?
Uploaded thread: 019a7f29-4b96-7ef3-85c2-
What is the expected behavior?
_No response_
Additional information
_No response_
Showing cached comments. Read the full discussion on GitHub ↗
10 Comments
Could you provide more details about what you're seeing? Also, it's very helpful if you can use the
/feedbacktool to upload your logs and optionally your conversation thread for reference (thanks for doing that, @perbinder).I am also seeing surprisingly bad results on a super small project - gave it multiple turns on the simplest of tasks and it continues to fail. Already filed a feedback. Will be interested to see if this is just a n of 2 or if others are experiencing the same.
https://github.com/openai/codex/issues/6629
I did use the feedback.
5.1 cannot even work out how to use logging in py and showing them in
systemd logs.
On Thu, 13 Nov 2025 at 23:05, Justin Kaufman @.***>
wrote:
Yea it lies, halucinates and doubles down on its believs even when proving wrong
yeah now with codex 5 quantized and codex 5.1 being this bad, there is no options left.
Please don’t deprecate gpt-5-codex. The gpt-5.1-codex model currently has serious regressions and fails on tasks that the previous version handled correctly.
Yeah, after trying the exact same tasks a few times with both GPT 5 and GPT 5.1, I unfortunately have to agree. 5.1 struggles with tasks and often tries to cheat or just skip doing any work, it feels very much like Sonnet. It seems like a massive regression from 5
We've shared the feedback with the team members who are responsible for model training. We're tracking this feedback and similar reports so we can make the model more robust in future iterations.
If you see other model behavior that you'd like to report, please use the recently-added /feedback command in the CLI. That will provide us with additional details including logs and conversation text.
It's memory is appalling. It doesn't even follow it's own, self-created plan from a working document. It does things that are so incredibly stupid... things gpt-5-codex would never have done. The way it talks is worse, often describing things in a way which is not accurate to the project, or in a way that could be understood as having two meanings and thus unclear is in what it is saying. It uses highly technical jargon mixed in, but not in a meaningful way, and even as a veteran developer I have to ask it what the hell it's trying to say. This should not have been released.
Exactly, it will assume I've made a mistake like not rebuilding the project or restarting the dev server. It will vehemently argue that it's my fault, and after 5 dialogues back and forth it will finally agree that it did, in fact, not know WTF was going on or that it had made a mistake.