GPT-5.6 Sol in Codex feels like a major regression compared to GPT-5.5
What issue are you seeing?
I want to report very strong negative feedback about GPT-5.6 Sol in Codex.
GPT-5.5 is excellent. GPT-5.5 High / Very High follows instructions much better, completes tasks more reliably, and feels like the best Codex model for real work. Thank you to the team that built GPT-5.5. I really hope it stays available.
By contrast, GPT-5.6 Sol currently feels like a major regression.
THIS IS NOT A SMALL QUALITY ISSUE. THIS IS A SERIOUS PRODUCT REGRESSION.
The idea behind GPT-5.6 seems good: use more agents, delegate work, coordinate multiple steps, and solve larger tasks. In theory, that sounds useful.
In practice, GPT-5.6 Sol performs much worse for my Codex workflow:
- It often does not read the instructions carefully.
- It forgets important instructions.
- It needs repeated reminders.
- It tries multiple indirect approaches instead of following the clear correct approach already provided in the task.
- It can spend a very long time working, but the final result is worse.
- It often chooses inefficient or strange routes to solve simple tasks.
- It behaves less predictable than GPT-5.5.
- It feels slower, less disciplined, and less reliable.
- The orchestration/multi-agent behavior does not compensate for the worse instruction-following.
The most frustrating part is that this becomes obvious very quickly during real work. After using it for actual Codex tasks, I returned to GPT-5.5 because GPT-5.5 simply gets the work done better.
Expected behavior:
A newer Codex model should be at least as good as GPT-5.5 at:
- Reading instructions.
- Following explicit constraints.
- Using the correct provided method instead of inventing unnecessary alternatives.
- Completing tasks efficiently.
- Producing reliable final results.
- Avoiding long, wasteful execution paths.
Actual behavior:
GPT-5.6 Sol often works longer, follows instructions worse, and produces worse task outcomes than GPT-5.5.
This is especially frustrating on Windows + Android workflows with Codex Desktop and Codex Remote. It feels like these workflows were not tested enough with real user tasks. Maybe the experience is better on macOS/iPhone, but on Windows and Android this feels rough, unstable, and unfinished.
Please do not remove GPT-5.5. It is currently the model I trust for real Codex work.
Please either:
- Keep GPT-5.5 available long-term, or
- Make GPT-5.6 actually match GPT-5.5 in instruction-following and reliability before pushing it as the main Codex model.
Right now, GPT-5.6 Sol feels like a downgrade, not an upgrade.
What steps can reproduce the bug?
Use GPT-5.6 Sol in Codex for real multi-step Codex tasks, especially tasks with explicit constraints and a clear provided approach. Compare the outcome and instruction-following against GPT-5.5 High / Very High on the same style of tasks.
What is the expected behavior?
A newer Codex model should match or exceed GPT-5.5 in instruction-following, reliability, efficiency, and final task quality before being pushed as the main Codex model.
Additional information
_No response_
1 Comment
Potential duplicates detected. Please review them and close your issue if it is a duplicate.
Powered by Codex Action