Severe performance regression on the [$imagegen] skill in Codex App, GPT5.5 need investigation & fixes
What version of the Codex App are you using (From “About Codex” dialog)?
26.623.101652
What subscription do you have?
PRO
What platform is your computer?
MAC M4PRO Darwin 25.5.0 arm64 arm
What issue are you seeing?
<img width="720" height="420" alt="Image" src="https://github.com/user-attachments/assets/cd7d84ae-cb31-49ab-a14e-bb82824d0936" />
<img width="840" height="510" alt="Image" src="https://github.com/user-attachments/assets/dc7a2f80-497d-4321-9758-b257a41107f6" />
<img width="583" height="247" alt="Image" src="https://github.com/user-attachments/assets/dc0eadeb-0f0e-45cd-9bb4-ed303bc5a3b9" />
What steps can reproduce the bug?
The [$imagegen] skill in my Codex App has gotten way worse performance-wise lately, and I hope the team can look into this bug and push out fixes.
App version I'm running: 26.623.101652
I prompted it to make 3D-style icons, but the outputs constantly go against my prompts. All the icons it cranked out are plain, minimal flat designs instead. I've attached an icon I generated using the exact same [$imagegen] skill back in early June, and the gap in output quality is really night-and-day.
I've spotted another weird flaw with the [$imagegen] skill powered by GPT-5.5 on this 26.623.101652 build: it often deliberately veers away from prompt instructions and spits out totally unrelated imagery. I've attached two such irrelevant pictures as bug proof — the tool itself claimed those outputs were supposed to be icons. After it realized those results missed the mark, it only gave a short note saying the design had gone wrong. It felt like the underlying model got downgraded right after that; the icons it produced afterwards turned out even lower-quality and stuck to the unwanted flat aesthetic that broke my 3D requirements. Worst of all, the app mistakenly insisted those flat icons actually matched my requested 3D look.
I ran into all these issues when working with Codex App throughout late June to early July. For comparison, I tested Codex CLI using the same prompts, and it never showed any of these glitches, so the problem seems exclusive to the desktop app for now. I hold the Pro subscription, so I definitely didn't expect to hit such obvious drops in image-generation quality plus frequent prompt-drifting issues on GPT-5.5's [$imagegen] skill.
What is the expected behavior?
_No response_
Additional information
_No response_
This issue has 2 comments on GitHub. Read the full discussion on GitHub ↗