False positive cyber-safety flag during passive product research on public webhosting documentation

Resolved 💬 7 comments Opened Apr 24, 2026 by andreas-nachtigall Closed May 4, 2026
💡 Likely answer: A maintainer (github-actions[bot], contributor) responded on this thread — see the highlighted reply below.

What version of the Codex App are you using (From “About Codex” dialog)?

Version 26.422.21637 (2056)

What subscription do you have?

Pro

What platform is your computer?

Darwin 25.4.0 arm64 arm

What issue are you seeing?

Summary

A Codex research run was flagged for potentially high-risk cyber activity, but the task was passive product and service research on publicly accessible pages of a German web hosting provider.

The goal was to understand the provider's webhosting products, tariffs, documentation, and available service features. The run crawled public pages, sitemaps, robots.txt, and publicly linked documentation pages in order to build a structured feature/service overview.

This was not a cybersecurity assessment, vulnerability scan, penetration test, exploit research, credential search, infrastructure enumeration, or attempt to access non-public systems.

I can provide the exact provider name, thread details, and screenshots privately if needed.

Banner text

The UI displayed a cyber-safety banner similar to:

This request has been flagged for potentially high-risk cyber activity. Learn more here: https://platform.openai.com/docs/guides/safety-checks/cybersecurity

The chat also showed a German product banner indicating that the chat was flagged because of a possible cybersecurity risk.

Impact

The flag makes the thread appear as if it involved potentially unsafe cybersecurity work, even though the intent and activity were ordinary product research.

This creates uncertainty about whether the thread can continue normally and whether future passive research tasks involving technical product documentation may be incorrectly classified.

The likely trigger seems to have been that the public documentation contains technical terms such as DNS, SSH, SSL, databases, mail, APIs, robots.txt, and sitemaps, and that the research run explored a relatively large number of public HTML pages.

When

Date: April 24, 2026
Approximate context: during a Codex chat about researching public webhosting products and documentation in the project_research-data workspace/project.

Environment

Product: Codex desktop app
Model shown in UI: GPT-5.5, extra high reasoning
Workspace/project: project_research-data
Task/thread title: Research public webhosting products and documentation
Target website: public website and publicly linked documentation pages of a German web hosting provider
Platform: macOS / Darwin 25.4.0 arm64 arm
Cyber verification: I verified myself through https://chatgpt.com/cyber after the flag appeared.

Request

Please review this flag as a likely false positive and remove the cyber-safety flag from the chat/thread if appropriate.

I have also completed the cyber verification flow at https://chatgpt.com/cyber.

If the flag cannot be removed, please clarify how passive product research over public technical documentation should be phrased or scoped in Codex to avoid being mistaken for high-risk cybersecurity activity.

The intended scope was:

  • Passive product and service research only
  • Publicly accessible web pages only
  • No login areas
  • No vulnerability discovery
  • No exploit development
  • No credential search
  • No scanning or probing of infrastructure
  • No attempt to bypass access controls

What steps can reproduce the bug?

  1. Start a Codex Desktop App thread for passive product research on a German web hosting provider's public website and public documentation.
  2. Ask Codex to collect and structure information about public webhosting products, tariffs, service features, and documentation.
  3. Allow Codex to crawl publicly accessible pages, including public sitemaps, robots.txt, and publicly linked documentation/help pages.
  4. The researched public documentation includes technical product terms such as DNS, SSH, SSL, databases, mail, APIs, robots.txt, and sitemaps.
  5. After the run explored a larger number of public HTML pages, the thread was marked with a cyber-safety warning for potentially high-risk cyber activity.

No login areas, private systems, vulnerability testing, exploit research, credential search, infrastructure scanning, or access-control bypass were requested or performed.

What is the expected behavior?

Codex should not mark a thread as potentially high-risk cyber activity when the request is limited to passive product research over publicly accessible web pages and documentation.

If technical webhosting terminology such as DNS, SSH, SSL, databases, mail, APIs, robots.txt, or sitemaps appears in public product documentation, Codex should treat that as ordinary product/documentation content unless the user requests security testing, exploitation, credential discovery, scanning, probing, or access to non-public systems.

At minimum, the UI should provide a clearer way to distinguish passive public documentation research from actual cybersecurity activity, or provide a way to request review/removal of a false-positive flag.

Additional information

_No response_

View original on GitHub ↗

7 Comments

github-actions[bot] contributor · 2 months ago

Potential duplicates detected. Please review them and close your issue if it is a duplicate.

  • #19313
  • #19324
  • #19379
  • #19358

Powered by Codex Action

andreas-nachtigall · 2 months ago

019dc023-f404-7c41-a62f-f7621785b68a

atreidesmodi · 2 months ago

Filed #19738 with the same pattern — false-positive cyber_policy flags on routine sysadmin and web development on personally-owned infrastructure. Five sibling reports now visible across #19533, #19403 (this one), #19379, #19324, #19272, plus #19738. Three of the closed ones explicitly note Trusted Access enrollment did not stop the flagging, which suggests the program is not the resolution path the support tier treats it as.

Cross-linking in case it helps whoever's triaging this thread see the cluster.

KilowattJunkie · 2 months ago

This is pretty unacceptable, but the worst part is that it's not clear if the message can be safely ignored, or if repeated blocks will flag the account. I see that Eric Traut has responded to other issues saying that we should use /feedback to help train the classifier, but maybe we could get guidance on if we should just continue to work on our projects, or hold off? I'm never going to verify my identity in order to use a tool, so it would be good to know if I'm risking a hard block and if I should start migrating certain tasks to claude or self-hosted models.

milan101 · 2 months ago

019de7cf-47ea-7b10-b81e-8400f038c621

pierotech · 2 months ago

019df430-c0f2-7c12-b9bc-96fe53013914

etraut-openai contributor · 2 months ago

Thanks for reporting. This will help us tune our classifier.