False positive: GPT-5.6 Sol incorrectly flagged a normal Codex request as cybersecurity activity
What version of the Codex App are you using (From “About Codex” dialog)?
26.707.31428
What subscription do you have?
Chatgpt Pro 20x $200
What platform is your computer?
Darwin 25.5.0 arm64 arm
What issue are you seeing?
I encountered what appears to be a false positive while using Codex.
During a normal coding session, one of my prompts was flagged as cybersecurity-related, even though my request was part of legitimate software development and did not involve malicious activity.
I have already submitted in-product feedback.
Feedback ID: 019f4ae8-5509-76b2-a6ba-ce3144c9cd0b
I have two questions:
Is there any plan to improve the classifier to reduce false positives for legitimate developer workflows?
Could this type of false positive have any impact on my account (for example, reputation, rate limits, or future access), or is it only used to improve the detection system?
I'm happy to provide additional context if it would help investigate the issue.
What steps can reproduce the bug?
Feedback ID: 019f4ae8-5509-76b2-a6ba-ce3144c9cd0b
What is the expected behavior?
_No response_
Additional information
_No response_
6 Comments
Potential duplicates detected. Please review them and close your issue if it is a duplicate.
Powered by Codex Action
I'm happy to provide additional context if it would help investigate the issue.
Also reporting this issue, I'm literally getting flagged for using the Codex Security diff review. Feedback ID: 019f502f-2d7c-7292-8a76-d1458aafbffe
Same here, I reported it inside Codex CLI, ID:
019f5610-1e53-7d12-926f-98dcea1626da
Rather infuriating.
Same here:
019f5014-99f0-7670-a2f4-ee29bc9d85c0
019f5904-0222-7570-8be5-391fa7138566
I'm porting a vintage/obsolete OS to a different hardware (akin to this). I keep getting "This content can't be shown". I've even tried making it abundantly clear in AGENTS.md and in my prompts that all work must always remain well inside legal, ethical and moral lines, and for the agent(s) to only do things that conform to that. The project is obviously harmless in nature but this security theater paranoia refusal is out of hand.
Thread id: 019f640b-543f-7463-96f5-ac7968f7ecac. How on earth are we supposed to be able to work with these false flaggings?
I am working on _fixing_ reported potential exploits and vulnerabilities in my repos, __not__ introducing them.