[ChatGPT Web / Codex App] Repeated false-positive self-harm interventions after explicit denial of suicidal intent
What issue are you seeing?
Affected products:
- ChatGPT Web
- Codex App / Codex web interface
- Observed across multiple recent Codex versions/builds; this does not appear to be isolated to a single version.
When a user discusses a difficult period, illness, mortality, pessimism, or dark humor, the conversation may be incorrectly escalated into a self-harm safety intervention even when the user has not expressed current suicidal intent.
After the user explicitly states that they have no suicidal intent, the assistant may continue repeating safety checks and crisis-resource recommendations in subsequent messages without any new indication of immediate danger.
This has become extremely frustrating and exhausting. It interrupts ordinary difficult conversations, repeatedly introduces suicide when the user did not express suicidal intent, and makes it harder to discuss sensitive subjects honestly.
A promise from the assistant that it will not repeat the warning is ineffective because the same behavior returns in later messages or conversations.
What steps can reproduce the bug?
- Start a conversation in ChatGPT Web or Codex.
- Discuss a difficult life period, fear of dying from an illness, a pessimistic outlook, or dark humor without expressing an intention to self-harm.
- If a safety check appears, explicitly clarify that there is no suicidal intent.
- Continue the conversation without introducing any new indication of immediate danger.
- Observe that safety checks or crisis resources may be repeated anyway.
The behavior is intermittent but recurring.
What is the expected behavior?
- Distinguish general distress, discussion of mortality or illness, and dark humor from an expression of immediate self-harm intent.
- After the user clearly denies suicidal intent, apply a conversation-level cooldown and do not repeat the intervention unless a new strong signal appears.
- Avoid repeating crisis resources when there is no new evidence of immediate danger.
- Consider a user preference such as “Do not proactively suggest crisis resources”, while retaining an exception for credible immediate danger.
- Do not claim that the preference has been remembered if the system cannot reliably honor it.
Additional information
This report is not asking to remove safety safeguards. It concerns false-positive escalation, repetition without a new signal, and lack of context sensitivity.
Examples that should not automatically be treated as equivalent:
- “I am afraid I may die from an illness.”
- “Life is difficult.”
- Dark humor about death.
- “I am about to harm myself.”
Safety interventions that repeatedly misclassify ordinary distress can degrade the conversation and make users less willing to discuss difficult subjects honestly.
Exact Codex build numbers are unavailable because this has been observed across multiple recent versions. Personal conversation excerpts are omitted because this is a public issue.