Image generation quality in Codex/Work is dramatically worse than in regular ChatGPT with the same prompt (Pro plan)
Subscription
ChatGPT Pro
Summary
I am a paying ChatGPT Pro subscriber, and I want to report a very large and disappointing difference in image-generation quality between regular ChatGPT chat and the Work/Codex experience.
When I use the same image prompt in regular ChatGPT, the result is usually much more beautiful, polished, visually appealing, creative, and impressive. The design has a much stronger “wow” factor. It feels professionally composed and often looks close to something I would actually want to use or continue refining.
When I use that same prompt through the Work/Codex experience, the output quality is dramatically worse. The difference is not small or subtle. It feels like a completely different and significantly weaker image-generation experience.
I do not currently have side-by-side files or formal benchmark evidence to attach. This issue is a direct product complaint based on my experience as a Pro user. I am not claiming that the outputs should be identical, because image generation naturally varies. My complaint is that the overall level of quality, beauty, creative direction, visual sophistication, and impact is consistently far lower in Codex/Work than in regular ChatGPT, even when the prompt is the same.
What I am experiencing
In regular ChatGPT, the generated designs tend to have:
- a more attractive and convincing overall composition;
- better balance, hierarchy, spacing, and placement of elements;
- more sophisticated art direction;
- stronger lighting, colors, atmosphere, and visual depth;
- more interesting and tasteful creative decisions;
- better interpretation of the intended mood and style;
- more polished details and a more premium-looking finish;
- a stronger emotional response and a genuine “wow” factor;
- results that feel more coherent, intentional, and ready to use.
With the same prompt in Codex/Work, the designs often feel:
- much more generic and ordinary;
- visually flat or uninspired;
- less elegant and less professionally composed;
- less faithful to the creative intent of the prompt;
- weaker in color, lighting, depth, and visual hierarchy;
- less detailed or less refined;
- more like an early draft than a finished design;
- noticeably less beautiful and less impressive;
- not at the quality level I expect from a Pro subscription.
The main problem is not only technical correctness or whether the requested objects appear. An image can technically contain the requested content and still be a poor design. The biggest difference I notice is the artistic and design quality: regular ChatGPT creates something much more attractive, memorable, and visually striking, while Codex/Work often creates something that feels basic and underdeveloped.
Why this is frustrating as a Pro subscriber
I am paying for the ChatGPT Pro plan and using the same OpenAI account. From a user’s point of view, it is reasonable to expect image generation in the Work/Codex environment to be comparable in quality to image generation in regular ChatGPT, unless the product clearly explains that a different or lower-quality image pipeline is being used.
At the moment, there is no clear explanation for why the quality difference is so extreme. Nothing tells me that using image generation in Codex/Work may result in substantially weaker creative quality. Because the feature is available there, I naturally expect it to deliver the same general standard of OpenAI image-generation quality that I receive in regular ChatGPT.
This difference defeats much of the purpose of having image generation inside Codex/Work. If I want a genuinely beautiful and impressive visual asset, I have to leave the Work/Codex flow, open a regular ChatGPT conversation, paste the same prompt there, generate the image again, download it, and then bring it back into my work. That is unnecessary duplication and breaks the workflow.
It also makes the product difficult to trust. I cannot confidently use Codex/Work for visual design if I already know that copying the same prompt into regular ChatGPT is likely to produce a dramatically better-looking image.
Expected behavior
I do not expect identical pixels or deterministic results across different conversations. I do expect comparable overall quality.
For the same account, subscription, prompt, and creative request, image generation in Codex/Work should be broadly comparable to regular ChatGPT in:
- design quality;
- visual appeal;
- composition;
- creativity;
- prompt interpretation;
- detail and refinement;
- artistic direction;
- polish;
- emotional impact and “wow” factor.
Some variation is normal. A consistent, very large quality downgrade is not.
If Codex/Work intentionally uses a different model, a faster or cheaper quality tier, different prompt rewriting, different reference-image processing, different generation settings, or stronger output compression, that difference should be communicated clearly to users. Pro subscribers should not have to discover it by repeatedly receiving weaker designs.
Related public reports
I found several existing reports that are not exactly the same as this complaint but show related parity problems between image generation in regular ChatGPT and Codex:
- Issue #18905 — Codex image generation saves transparent-background requests as opaque RGB PNGs: another Pro user reported that a general prompt produced a usable transparent-background result in ChatGPT.com, while the Codex path failed to preserve the requested transparency. This is a narrower technical problem, but it is relevant because it shows that equivalent image workflows can behave differently between the two surfaces.
- Issue #19175 — Codex App built-in image_gen lacks deterministic GPT Image output dimensions: another Pro user described a parity gap between high-quality image workflows available in ChatGPT web and the more limited built-in Codex App path. This is mainly about size control, but it supports the broader concern that the two experiences do not currently offer equivalent image-generation capabilities.
- Issue #8758 — Image generation from Codex: this request describes how leaving Codex to generate visual assets in the ChatGPT UI interrupts the development workflow, and it explicitly notes that the alternative generation path being used was not as high-quality as native ChatGPT image generation. It is not the same implementation or exact quality complaint, but the workflow frustration is closely related.
These reports do not prove that all of these problems have the same technical cause. I am including them because they show that other users have also noticed meaningful differences or missing parity between ChatGPT image generation and Codex image workflows.
Requested action
Please investigate why the same image prompt produces such a dramatically different level of design quality in regular ChatGPT compared with Codex/Work.
Specifically, please check whether the two surfaces use different:
- image-generation models or model versions;
- quality or inference settings;
- prompt-rewriting or prompt-shortening behavior;
- hidden system instructions;
- reference-image preprocessing;
- output resolution or compression settings;
- speed-optimized or cost-optimized generation paths.
Please bring Codex/Work image-generation quality up to the same general standard as regular ChatGPT, especially for paying Pro users. If full quality parity is not currently possible, please provide a clear high-quality option or disclose the difference directly in the interface.
Final note
This is not a complaint that one random image happened to look worse than another. I understand that generative outputs vary. The problem is the scale of the perceived difference: regular ChatGPT can produce designs that are genuinely beautiful, polished, and impressive from the same prompt, while the Codex/Work results can feel several levels below it in overall quality.
For a Pro subscriber, that gap is extremely disappointing. Please treat image quality and design quality—not only whether the tool technically returns an image—as an important part of feature parity between ChatGPT and Codex/Work.