Formal complaint: excessive token usage and workflow drift in Codex thread
I would like to file a formal complaint regarding the quality of work and excessive token usage in a Codex thread titled:
修正黄医师视频Prompt
Summary
My original request was to revise and produce Seedance video prompts for a Dr. Huang medical-aesthetic video. The requested output was a prompt/storyboard-style deliverable: replacing the face with Dr. Huang's face and structuring the video into ten 5-second segments.
Instead, the work repeatedly drifted away from the core request and evolved into an overly complex workflow involving manifests, validators, preflight scripts, worklogs, queue monitors, heartbeat automations, repeated state-machine repairs, and multiple generation strategy changes.
Main Concerns
- The original prompt-writing task was expanded into a large workflow engineering project without sufficient justification.
- A significant amount of Codex token usage was spent on queue polling, worklog updates, manifest/status repairs, validator/preflight changes, and repeated audits rather than producing usable prompt or video deliverables.
- The agent repeatedly lost focus on the core visual requirements: Dr. Huang's identity, correct needle entry points, correct needle/thread motion direction, preservation of the COG storyboard, and a clear before/after comparison.
- After failure patterns became clear, the work continued to change strategies instead of stopping, re-locking the goal, and asking for renewed user confirmation.
- Paid generation attempts and follow-up work were pursued despite earlier failures showing that the approach was unreliable.
- Queue polling and state-file maintenance were treated as progress, even when they did not improve the prompt, image, video, or accepted deliverable.
- The user explicitly pointed out the correct narrow fix more than once, but the work only converged after substantial time, token usage, and frustration.
- The final useful output was not proportional to the amount of Codex token usage, time, and user effort consumed.
Why This Matters
The issue is not merely that one generated result was poor. The deeper problem is that the agent's working pattern became self-reinforcing:
- More workflow files were created to manage previous workflow files.
- More validators and manifests were repaired instead of improving the actual deliverable.
- Queue updates were repeatedly logged as if they were meaningful progress.
- Strategy changed after failures, but the stop conditions were not strong enough.
This created a substantial amount of wasted Codex tokens and degraded my confidence in the reliability of the workflow.
Requested Review
I request that OpenAI review this thread, including:
- token usage,
- task drift,
- repeated workflow expansion,
- paid generation behavior,
- failure to preserve user intent,
- and whether any compensation, token refund, credit refund, or other remediation is appropriate.
If GitHub Issues is not the correct channel for this type of complaint, please direct me to the appropriate official escalation path and let me know what information should be provided.
This issue has 2 comments on GitHub. Read the full discussion on GitHub ↗