GPT-5.6 Sol in Codex feels like a major regression compared to GPT-5.5

Open 💬 1 comment Opened Aug 2, 2026 by Frie666
💡 Likely answer: A maintainer (github-actions[bot], contributor) responded on this thread — see the highlighted reply below.

What issue are you seeing?

I want to report very strong negative feedback about GPT-5.6 Sol in Codex.

GPT-5.5 is excellent. GPT-5.5 High / Very High follows instructions much better, completes tasks more reliably, and feels like the best Codex model for real work. Thank you to the team that built GPT-5.5. I really hope it stays available.

By contrast, GPT-5.6 Sol currently feels like a major regression.

THIS IS NOT A SMALL QUALITY ISSUE. THIS IS A SERIOUS PRODUCT REGRESSION.

The idea behind GPT-5.6 seems good: use more agents, delegate work, coordinate multiple steps, and solve larger tasks. In theory, that sounds useful.

In practice, GPT-5.6 Sol performs much worse for my Codex workflow:

  • It often does not read the instructions carefully.
  • It forgets important instructions.
  • It needs repeated reminders.
  • It tries multiple indirect approaches instead of following the clear correct approach already provided in the task.
  • It can spend a very long time working, but the final result is worse.
  • It often chooses inefficient or strange routes to solve simple tasks.
  • It behaves less predictable than GPT-5.5.
  • It feels slower, less disciplined, and less reliable.
  • The orchestration/multi-agent behavior does not compensate for the worse instruction-following.

The most frustrating part is that this becomes obvious very quickly during real work. After using it for actual Codex tasks, I returned to GPT-5.5 because GPT-5.5 simply gets the work done better.

Expected behavior:
A newer Codex model should be at least as good as GPT-5.5 at:

  • Reading instructions.
  • Following explicit constraints.
  • Using the correct provided method instead of inventing unnecessary alternatives.
  • Completing tasks efficiently.
  • Producing reliable final results.
  • Avoiding long, wasteful execution paths.

Actual behavior:
GPT-5.6 Sol often works longer, follows instructions worse, and produces worse task outcomes than GPT-5.5.

This is especially frustrating on Windows + Android workflows with Codex Desktop and Codex Remote. It feels like these workflows were not tested enough with real user tasks. Maybe the experience is better on macOS/iPhone, but on Windows and Android this feels rough, unstable, and unfinished.

Please do not remove GPT-5.5. It is currently the model I trust for real Codex work.

Please either:

  1. Keep GPT-5.5 available long-term, or
  2. Make GPT-5.6 actually match GPT-5.5 in instruction-following and reliability before pushing it as the main Codex model.

Right now, GPT-5.6 Sol feels like a downgrade, not an upgrade.

What steps can reproduce the bug?

Use GPT-5.6 Sol in Codex for real multi-step Codex tasks, especially tasks with explicit constraints and a clear provided approach. Compare the outcome and instruction-following against GPT-5.5 High / Very High on the same style of tasks.

What is the expected behavior?

A newer Codex model should match or exceed GPT-5.5 in instruction-following, reliability, efficiency, and final task quality before being pushed as the main Codex model.

Additional information

_No response_

View original on GitHub ↗

1 Comment

github-actions[bot] contributor · 26 days ago

Potential duplicates detected. Please review them and close your issue if it is a duplicate.

  • #36229
  • #36086

Powered by Codex Action