Agent safeguard repeatedly flags ordinary dev automation as "possible cybersecurity risk," escalating to a full goal block on a long-running unattended task

Open 💬 2 comments Opened Aug 17, 2026 by ThatDragonOverThere
💡 Likely answer: A maintainer (github-actions[bot], contributor) responded on this thread — see the highlighted reply below.

Environment

  • Codex mobile app (iOS)
  • An agent session running an unattended, multi-day automated software task (build/test/refactor work on a personal project), controlled remotely via the mobile app
  • No security research, exploit development, red-team activity, or credential-related work occurred in any of the flagged steps

What happened

Over a roughly 20-minute window in one continuous session, the same false-positive safety flag fired four separate times, each time immediately after entirely ordinary agent actions:

  • Running shell/PowerShell commands (routine build and test tooling)
  • Editing and creating project script and test files
  • Sending a status message to a sub-agent handling a separate bounded task
  • Agent reasoning text describing pipeline steps, internal gating/validation logic, and a fix to an orchestration bug — no security-relevant content, no credentials, no exploit code, no mention of any external target

Every one of the four occurrences produced the identical UI message:

"This content was flagged for possible cybersecurity risk. If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program: https://chatgpt.com/cyber"

Impact — this is not just a warning

This is not a single false-positive notice that the user can shrug off. Across the four occurrences the session's own status indicator visibly degraded:

  • Pursuing goal (first two occurrences)
  • Queued (third occurrence, mid-response)
  • Goal blocked (fourth occurrence)

The safeguard did not just mis-flag benign automation — it accumulated across repeated ordinary tool calls until it halted a multi-day unattended background task. There was no in-session way to provide clarifying context, confirm the work was benign, or resume; the task simply stopped making progress.

What was NOT present in any flagged content

  • No security tooling, exploit code, or offensive-technique references
  • No production/live-system access — the agent's own stated task explicitly scoped itself to test/paper-only work and named the exact gates blocking any live/production activation
  • No credentials, tokens, or secrets in any visible text
  • No content resembling an actual security-research request of any kind

Suspected trigger (unconfirmed)

The flagged steps have one thing in common: they're the routine shape of any autonomous coding agent's work — repeated shell command invocations, file edits, and inter-agent messaging inside one long-running, unattended session. It looks more like a volume/pattern-based heuristic (many tool calls of a certain shape in sequence) than a content classifier reacting to anything actually said. If that's correct, the "possible cybersecurity risk" framing is actively misleading — a rate-limit or throttling message would be an honest description of what's happening; "you might be doing security research" is not.

Ask

  1. Investigate why an unattended agent's routine shell-command / file-edit / sub-agent-messaging automation is tripping a "possible cybersecurity risk" classifier with no security-relevant content anywhere in the flagged turns.
  2. Separate the failure modes: a soft "this got flagged, want to rephrase?" notice is recoverable; silently degrading a long-running unattended task's status to Goal blocked with no way to self-correct in-session is a much more severe failure and needs its own fix independent of the classifier accuracy issue.
  3. If this is intentional throttling rather than a security determination, surface it as that — the current "Trusted Access for Cyber" framing tells the user they're doing something they're not.

Evidence available

Four mobile-app screenshots from one continuous session, timestamped roughly 5-6 minutes apart across a 20-minute span, each showing the flag box immediately following a benign, visible tool-call sequence (shell command runs, a file edit, a test-file creation, a sub-agent message). Happy to attach on request.

View original on GitHub ↗

2 Comments

github-actions[bot] contributor · 10 days ago

Potential duplicates detected. Please review them and close your issue if it is a duplicate.

  • #37854
  • #38516
  • #37702
  • #38464

Powered by Codex Action

nuliknol · 6 days ago

this security filter is total madness, I can't work. If I ask "what are current bugs in in my app", it blocks conversation. Instead of helping us to write better code, they are helping hackers to attack us by not allowing us to write bug free code. This is like not allowing home owners to replace locks on their doors. Stupid AI, Singularity is not near, it is not even on the horizon