CLI worker silent death after parallel-command burst (announce → mkdir → probe → exit) on Linux/WSL2 — `Turn completed` event never fires
Summary
The Codex CLI app-server worker process exits with code 0 mid-turn on Linux/WSL2 immediately after a specific tool-call sequence pattern, leaving the broker stuck and the wrapper layer hanging on state.completion. No error event is propagated. Reproducible 4 times across 24 hours on codex-cli 0.130.0 and 0.129.0.
Environment
- Codex CLI: 0.129.0 → 0.130.0 (both reproduce)
- Companion / wrapper:
@openai/codexplugin runtime viacodex-companion.mjs - Host: Windows 11 + WSL2 (Ubuntu 22.04, kernel 6.6.114.1-microsoft-standard-WSL2)
- Model: gpt-5.5 with xhigh reasoning effort (default for codex)
- Sandbox:
danger-full-access - Trigger: dispatched as background job from a Claude Code orchestrator
Reproducer (high reliability — 4/4 in 24h)
The worker dies after this exact sequence in a single turn:
- Multiple
psql/sedexploration commands (parallel batch, all complete with exit 0) - Model emits an "announcement" assistant message ("I'm going to write
<file>...") mkdir -p <output_dir>— succeeds- ONE more
psqlor shell probe (e.g.,psql -Atc "SELECT exchange, contract_type FROM symbols") — succeeds - Worker silently exits. No
Turn completedevent, no error, no SIGTERM in logs. Companionstate.completionPromise hangs indefinitely.
Observed across 4 distinct Codex jobs today:
task-moxeya7t-248j70(v4 analysis executor, 21:20 UTC)task-moxpj6m0-zr9uvr(PR creation task, 02:13 UTC next day)task-moxq4ixm-aln8jh(REST probe investigation, 02:29 UTC)bt0dkgkg6/b77sx7ne9(analytical re-run, 16:00 UTC)
What I see in logs
/home/user/.claude/plugins/data/codex-openai-codex/state/<workspace>/jobs/<task-id>.log ends with:
[ts] Command completed: /bin/bash -lc 'psql ...' (exit 0)
and zero lines after that. No turn/completed notification, no client.exit, no error event.
broker.log is empty (0 bytes) on the workspace's /tmp/cxc-*/ directory.
Expected behavior
Either:
- Turn completes normally with subsequent file-write tool calls
- OR worker emits an error event before exiting
Neither happens — exit is silent.
Workaround applied locally (wrapper-layer)
Patched lib/codex.mjs captureTurn() to use Promise.race([state.completion, exitRace]) where exitRace watches client.exitPromise and throws if turn isn't yet complete. This converts indefinite hang into clean exception, but doesn't fix the underlying app-server crash.
Related (non-duplicate)
- #21813 — Detached task-worker exits without writing failed status on broker socket disconnect (related: same silent-death symptom, but that issue's trigger is a broker socket reset; here the broker is alive and broker.log is 0 bytes — no socket event is involved)
- #14731 — Turn completes prematurely when unified_exec background processes are still running (different failure mode)
- #16271 — Any Single Engine Crash Leaves Orphaned Threads (Windows app, different surface)
Reproduction setup (if upstream wants to bisect)
I can provide:
- All 4 task log files (~10-90 lines each, JSON-ish single-line structure)
- Companion broker.log captures
- Process tree at moment of failure (broker alive, app-server dead)
- The wrapper-layer Promise.race patch as before/after diff
Request: enable structured logging in app-server's exit path, OR add a "worker-died-mid-turn" event to surface this state instead of silent exit.
This issue has 5 comments on GitHub. Read the full discussion on GitHub ↗