GPT5.5 is too stupid to use it now!
Open 💬 8 comments Opened Jun 18, 2026 by withTheHidden
💡 Likely answer: A maintainer (github-actions[bot], contributor)
responded on this thread — see the highlighted reply below.
What issue are you seeing?
WTF!
I don’t wanna curse, but it’s been acting so damn stupid the last couple of days. Like, is the plan just to make GPT-5.5 so ridiculously dumb that when you guys drop any new model later, everyone’s gonna be like, 'Wow, so smart!'?
What steps can reproduce the bug?
use codex and chose the GPT 5.5 now!
What is the expected behavior?
_No response_
Additional information
_No response_
8 Comments
Potential duplicates detected. Please review them and close your issue if it is a duplicate.
Powered by Codex Action
honestly i am using gpt 5.4mini
no issues for me
tested open source models non compare even the 5.4mini (that being said 5,3 was kindda bad but the apporch was correct with caveman basic low amount of words logic)
but there might be a regression , due to huge spikes in CODEX usage , which is uncommon when they dont release a product
They targeted some accounts with limits but not others. I've got real evidence!
<img width="373" height="598" alt="Image" src="https://github.com/user-attachments/assets/69b78335-a4dd-4143-964a-7c23081a5bfd" />
can you share the eval code? my model is behaving strangely stupid with xhigh
can you share the eval.py script, so we can evaluate what it does and how (api calls - or does it called the
codexbinary ? )Look put yourself in openai shoes , what you choose is what you get
you mention some accounts are being limited (Free accounts? Subscribed accounts? Invited plus accounts? )
Openai codex is a very capiable product , (despite the fact I tend it to be a medium sized model )
you can get a good milage in the free-accounts version (when I subscribed to plus last month , it ended last week --- it was very good - im able to maintain about 20 codebases - and various tasks )
admitably ratelimits has increased , but I only got it once (and I use /goal heavily -- and the /goal is really good - its like by itself now finish tasks faster [which is very interesting despite using the same model -- and represent it as a fact the codex - goal software has increased] )
the reality is , openai has to deliever the models to a lot of people ,, and also maintain training on new models (and also know this, a lot of expirements fail too)
my points is , the service-status is really good -- they have to provide to a lot of customers -- and still also make training (which is consuming and not all gpu is reach more than 80% --- despite how much we can praise modern IC packagings -- )
my recomendation to you, do not gamble with prompts - but you can play with --- I recomend going old school --- learning how to use documentation -- pray on technicality -- heavily use docmentation -- heavily compare products (if you have issues , I recomend llm , and use the keywords FOSS COMPARE to X or Y ) if you want a good refrence --- with coding and software developing -- there is always "logic" you need to decide - so you dont find a 2 months later, what you dont is the bare minimum of basic
(some people in software development , demand high salaries, but barely can navigate GCP or AZURE or AWS -- when most of it is [declare to yourself what is the task] --> [find the term] --> [search on yandex/google/bing/yahoo/ documentation on how to make such task ]) , (and then a lot of things comes to cost -- which is them job - not the accountants job )
some use AI for [web development] and its fine , either you understand how design the end outlet/template you want (modify buttons size) -- either you learn to assist the documentation/toolusage/learn //// or you learn how to talk to the llm, to describe what you want - not neglecting other features you might want --- just hammering with I am not happy with the result not going to help (but there is truth to not being happy because its not perfect --- [it always looks like AI-ish] [somes its not look like ai - I guess you can ask it for follow basic html and no css - use w3c guidlnes ] and you tell if the features you want it to do )
you need to learn how to review, its a part of the job as a developer - if you need to know what the end result you want - (think vaguely in the back of your mind -- [whats the next best product -- if its something like math or medicine - apporch the litterature ]) -- some might want the job because its the joy of craft - and entering pathed landed on the earth
that being said, the reality is it - be surprised those ai models can be run as of now
Gotta disagreed hard here. I know exactly what the user that posted is talking about. I'm going through it right now. I would say that your assumption that the user is not following protocols or best practices properly and it's a user error sort of thing if he didn't say that it was acting incredibly dumb the last couple of days. That means, he did have it working and it suddenly went to crap. Which is exactly why i'm here right now because the past 2 days, i can barely get my agents to properly capture intent or follow basic semantic dialogue without drifting and showing other immediate signs of mode and collaboration drifts. I have strict guardrails and reanchors and relatches that i use. And mine are usually really good. Until the past 48 hours, i have been able to do nothing but try and get it to understand things that it never had a problem doing a week ago. Same model, same subscription, same project, same control docs. It just got retarded. So when a user like the original poster is posting, it's probably good to see that as a sign for the devs to be aware that the same model and same intelligence level is acting sub-par for users. You trying to tell them how to do certain things without knowing what they are doing is kind of a waste, especially if you failed to read that it was working fine until the past 2 days. That should be your clue.
Totally agree. GPT 5.5 xhigh has been worse than GLM 5.2 in my usage cases. Could be due to a reduced context window length.
Can you share exact prompts you used, at the launch
and prompts you gave it now ?
I agree , all of the models have a very tendacy to not respect the guildlines or system prompts if thats what you mean (I actually dont give it any ,,, ) for me what works is actually not to try to give it guildlines or too much
see https://learn.chatgpt.com/docs/agent-configuration/agents-md
(but a lot changes. and they also working on the agents features)
I agree to some level , but as for web developement , if you tell it not to do css or whatever it actually results in faster results for me (you know I think the answer is diffusion reasoning or whatever happening in the image models - like css can be long but you shouldnt spent to much time over it -- but you must treat all tokens equal)
I dont know to respond to that, but your point likely aims for a real place - look for what users prompts there should be a way to deal with the prompts they give (for example I want to see how people achieve what they want faster - but I dont have anywhere to see how they vibe code) I think the answer to this is actually start what is a bad prompt -- and its a big paradox --- we are only left with the user prompt --- and we cant really know what he really wanted or meant (but actually what I personally liking with language models --- is that they actually land you other ideas than you initially wanted , and you can get creative... )
but its task dependent -- I dont have any issues with codex --
i also have a subscribtion , for claude sonnet 5 , I used , and its not really something I enjoy to much , it thinks for too long it make up things --- but maybe it ends in the right place for you... but honestly I like gpt 5.4mini as its actually much smarter than what I tested it for... I even see users and hear off people - who hit limits very often - but do not on codex despite very heavy usage on codex allowing them a lot more prompts on codex.... I do like new writing styles - for me I done some work I wanted to do , I used Luna (the cheapest model) ($1.00 input / $6.00 output) which gets anything I wanted faster --- I tried rerunning tasks - I felt challenging with 5.4 and 5.5 - it got thinks very quickly.... but I tested today 5.4mini and it beaten claude sonnet 5 for me.... I had a lot of fun with the open source models today too... I actually feel open source models are actually surpierer to Anthropic Claude models -- but thats just my feeling (and ofcourse my personal opnion the codex models are on another level)
I personally think its either of two things
a. the model is too fast, you need to take a small break and think of your task.. I recomend trying to /goal feature or /plan instead of using "hey codex please launch two subagents" this is actually will break things for you.. I dont say dont use subagents. But I do say try other features like /goal which is fantastic and /plan which could be faster
a2. (you mention you use agents) , I personally tried using agents and its not a great expirence (there is in the documentation mentioning of automation / Scheduled tasks ,, see https://learn.chatgpt.com/docs/automations?surface=app )
b.I think your codebase might gotten big --- maybe you keep reusing the same chat (or even maybe you should reuse older chat after /compact)
c. maybe you are correct to some degree, I noticed some degraded
e. a lot of the things are being updated -- a lot of things in the api are getting broken - despite not for most users - (but again no user who pay money should have his service taken away --- or degraded more than generally understood to be believed in -- or generally offered to other users at the same price level without a degree of 20% difference between the same group of users) openai is a very honest company they dont try to make fun users who are heavy users -- or drain your usage limit too much
f. a lot of things work out of the box I actually recomend to if you like the
skillsthing , thats good! you can actually tell it "please make a skill based on our chats on how to restart our server, following the guildlines we end up wanting" I recomend try to summerize it and try to roast / grill it and try ti spec and specify what you want out of it to do (make sure you dont wear out and lose manual skills because without them you dont know how to dictate your language model) --- what I feel about it - doing it out of the box is actually not that great --- you should do it within an existing chat --- you shouldnt try use a different skill --- because I think you can run into some quick model-behavior limitation --- but you should see what works for you -- on topic I think the model didnt became stupid -- because I tried it , I tried prompts that were very complex and took the model more than 10 steps and it didnt failed , and it did another 10 steps without becoming dumbg. maybe the first time you get existed , and you do your prepartion and give better prompts because you are excited as a new brand new model has launch... but then you try it later and then its not the same thing (which look if this is the common users something should be about this type of users too)