Significant reasoning degradation
What version of the Codex App are you using (From “About Codex” dialog)?
GPT-5.6 Sol High Reasoning behaves significantly differently between Mac desktop app and iOS app
What subscription do you have?
Premium
What platform is your computer?
Mac
What issue are you seeing?
Issue: GPT-5.6 Sol High Reasoning behaves significantly differently between Mac desktop app and iOS app
I am reporting a reproducible issue affecting GPT-5.6 Sol reasoning quality when using the same ChatGPT account, same conversation, same model and same reasoning setting across different clients.
I am a ChatGPT Plus user using GPT extensively for complex analytical and professional workflows.
Recently I have observed a severe degradation in reasoning quality when using the ChatGPT desktop application on Mac. The issue does not appear to be related to the GPT model capability itself, because the same prompts, in the same conversations, continue to produce high-quality results when submitted through the ChatGPT iOS application.
The problem appears to be specific to the Mac desktop application execution path, model routing, reasoning configuration, or another client-specific factor.
Reproduction steps
The issue can be reproduced as follows:
- Open an existing complex ChatGPT conversation.
- Ensure the selected model is GPT-5.6 Sol.
- Select High reasoning mode.
- Submit a complex analytical prompt.
When submitted through the Mac desktop application:
- The response quality is dramatically lower than expected.
- Reasoning depth appears significantly reduced despite High reasoning being selected.
- The model reaches conclusions prematurely.
- Important constraints, dependencies and contradictions are frequently missed.
- Responses often resemble a lower reasoning-effort mode rather than High reasoning.
If I take the exact same conversation and exact same prompt and submit it through the ChatGPT iOS application:
- The expected High reasoning behaviour returns.
- The model performs deep analysis.
- It considers alternatives and trade-offs.
- It maintains context and constraints significantly better.
- Output quality is comparable to the behaviour I previously experienced before this issue appeared.
Consistency of the issue
This is not based on one isolated example.
I have tested this repeatedly across:
- multiple different projects;
- multiple unrelated conversations;
- different complex analytical tasks;
- many repeated prompt comparisons.
The pattern is highly consistent:
- Mac desktop application → degraded reasoning quality
- iOS ChatGPT application → expected high-quality reasoning
The difference is reproducible enough that normal response randomness does not appear to be a likely explanation.
A particularly strong reproduction case is:
- Submit a prompt from the Mac desktop application and receive a poor-quality response.
- Edit or resend the same prompt in the same conversation from the iPhone application.
- Receive a substantially better response with deeper reasoning and better adherence to constraints.
The conversation history, account, model selection, reasoning setting and prompt are effectively unchanged. The main variable appears to be the client used to submit the request.
Potential areas to investigate
Could you please investigate whether there are differences between Mac desktop and iOS clients affecting:
- Whether the High reasoning setting is correctly transmitted from the Mac desktop application.
- Whether GPT-5.6 Sol requests from the Mac desktop application are routed identically to iOS requests.
- Whether there are client-specific feature flags, rollout configurations or backend settings affecting reasoning behaviour.
- Whether the Mac desktop application may incorrectly display High reasoning while executing with a different reasoning configuration.
- Whether there are known regressions affecting GPT-5.6 Sol reasoning quality in the current Mac desktop application.
Environment
- Plan: ChatGPT Plus
- Model: GPT-5.6 Sol
- Reasoning mode: High
- Client affected: ChatGPT Mac desktop application
- Client not affected: ChatGPT iOS application
- macOS version: [insert]
- Mac ChatGPT application version: [insert]
- iOS version: [insert]
- ChatGPT iOS application version: [insert]
- Approximate date issue started: [insert]
Additional information
This issue is materially affecting my workflow because I use ChatGPT for high-complexity analytical work where reasoning depth, maintaining constraints, evaluating alternatives and avoiding premature conclusions are critical.
I would appreciate this being investigated as a potential desktop client/model-routing/reasoning-configuration issue rather than treated only as subjective feedback about response quality.
I can provide paired examples from the same conversations demonstrating the difference between Mac desktop and iOS outputs if required.
What steps can reproduce the bug?
Reproduction steps
The issue can be reproduced as follows:
- Open an existing complex ChatGPT conversation.
- Ensure the selected model is GPT-5.6 Sol.
- Select High reasoning mode.
- Submit a complex analytical prompt.
When submitted through the Mac desktop application:
- The response quality is dramatically lower than expected.
- Reasoning depth appears significantly reduced despite High reasoning being selected.
- The model reaches conclusions prematurely.
- Important constraints, dependencies and contradictions are frequently missed.
- Responses often resemble a lower reasoning-effort mode rather than High reasoning.
If I take the exact same conversation and exact same prompt and submit it through the ChatGPT iOS application:
- The expected High reasoning behaviour returns.
- The model performs deep analysis.
- It considers alternatives and trade-offs.
- It maintains context and constraints significantly better.
- Output quality is comparable to the behaviour I previously experienced before this issue appeared.
What is the expected behavior?
Model picker control will be respected - otherwise the quality of output is degrading significantly and GPT becomes usless.
Additional information
_No response_
1 Comment
I have now confirmed this is not a single-chat or prompt-specific issue. The same pattern reproduces across multiple unrelated projects. The same conversation and exact same prompt produce significantly different outputs depending on whether the request is submitted from the Mac desktop application or iOS application.
Importantly, resubmitting the same prompt from iOS restores the expected High-reasoning behaviour, which suggests the model capability itself is available and the issue may be related to desktop-specific routing, configuration, rollout or reasoning-effort application.