Codex 5-hour usage meter appears to consume much faster than comparable historical usage
Summary
I am seeing a major discrepancy between local Codex/Hermes usage telemetry and the Codex 5-hour usage meter. Recent usage appears to consume a much larger percentage of the 5-hour allowance than historically comparable or much larger sessions.
This may indicate a regression in Codex quota accounting, prompt-cache accounting, context/tool-schema weighting, or effective allowance size.
Current observed Codex usage
From the authenticated Codex usage endpoint (/backend-api/codex/usage), redacted and summarized:
- Plan type:
prolite - Primary window: 5 hours / 18,000 seconds
- Primary
used_percent: 33% - Secondary weekly
used_percent: 5% limit_reached: false
A tiny live Codex probe succeeded, so the account was not hard rate-limited at the time of measurement.
Local usage comparison
Across local Hermes/Codex usage ledgers for the same environment/accounts:
Current combined last 5-hour usage:
- ~5.53M total tokens, including cached tokens
Historical rolling 5-hour windows in the last 30 days include:
- ~91.27M tokens
- ~91.21M tokens
- ~87.59M tokens
- ~80.78M tokens
- ~78.01M tokens
- ~75.13M tokens
The current 5-hour period is only around 6.1% of the largest recent 5-hour local usage window, yet the Codex usage endpoint reports 33% of the current 5-hour allowance consumed.
Why this seems anomalous
This account has had the subscription for months and has previously run much larger sessions with much larger local token payloads without the 5-hour allowance moving this aggressively.
The current behavior feels like one of:
- the effective 5-hour allowance was reduced,
- Codex/GPT-5.5 usage weighting changed,
- cached/tool/context/schema tokens are being counted differently,
- prompt caching is failing or no longer credited correctly,
- or the usage endpoint/UI is reporting a different allowance/accounting model than before.
Related public reports
Similar Codex usage-limit complaints appear in:
- #18422 — usage exhausted too quickly / feels more restrictive than before
- #19607 — single prompts heavily draining 5h/weekly usage
- #24560 — severe quota drain, suspected prompt-cache/context issue
- #24431 — GPT-5.5 reliability degradation wasting Pro quota
- #22650 — usage limit reached while status/web showed remaining quota
- #27083 — extension reports usage limit despite available quota
- #27850 — blocking with 100% left
- #28246 — quota windows anchoring/reset behavior concerns
Expected behavior
Comparable local token usage should consume a broadly comparable fraction of the Codex 5-hour allowance over time, or the product should expose clear accounting showing what changed.
Requested clarification
Can OpenAI confirm whether any of the following changed recently?
- Codex 5-hour allowance size
- GPT-5.5 quota weighting
- cached token accounting
- tool/schema/context token accounting
- primary/secondary window behavior
proliteplan allowance behavior
19 Comments
Potential duplicates detected. Please review them and close your issue if it is a duplicate.
Powered by Codex Action
Tokens consumption were not revealed after finishing a task (
Additional datapoint with local token telemetry, sanitized as combined local agent usage only:
OpenAI Codex usage endpoint:
limit_reached: falseLocal combined agent telemetry for the current rolling 5-hour window:
Historical local rolling 5-hour windows from the same aggregate ledgers include:
The current local combined 5-hour token volume is far below those historical 5-hour windows, yet the Codex usage endpoint is showing 79% of the primary 5-hour window used. This reinforces the concern that the discrepancy is not simply large local token volume, unless the effective allowance or token weighting/accounting changed substantially.
Clarification: I understand Codex subscription limits may not be raw-token based. The anomaly is that the effective drain rate changed materially for comparable local workflows, while the usage endpoint exposes only percentages and no accounting breakdown. Please expose or confirm the weighting inputs used for the 5-hour meter: request count, model multiplier, context length, cached-token handling, tool/schema tokens, speed mode, and weekly-window interaction.
确实是,10分钟干完了5小时限制
I'm experiencing the identical discrepancy on a plus-tier account.
From my local Hermes/Codex telemetry for the same weekly window:
This feels like either:
+1 to this. I am encountering the exact same regression on the Pro plan. My quota and top-up budget were completely wiped out in under 5 hours today, which has never happened before. It seems the token consumption or rate-limit calculation is heavily bugged right now, just as described in this issue. It is literally impossible to maintain a normal development workflow when a 5-hour budget can't even get you through the morning.
Additional fresh datapoint from the same account, with a short-sequence endpoint trace and local ledger delta.
Authenticated endpoint:
https://chatgpt.com/backend-api/codex/usagePlan reported:
proliteLocal timezone: AEST (+1000)
Endpoint sequence during one short troubleshooting/chat session:
During the first part of that sequence, local session counters changed by approximately:
Over that same evidence window, the endpoint primary 5-hour meter moved from 6% to 18% (+12 points), and the weekly meter moved from 51% to 53% (+2 points). That works out to roughly 101k local counted tokens per 1 percentage point of the 5-hour meter in this short sample, which is difficult to reconcile without knowing the hidden weighting/accounting model.
Also, the human-visible UI was observed as showing a different weekly percentage around the same time (47%) while the authenticated endpoint reported 53%. If those are different quota pools, the UI should label them differently; if they are intended to represent the same Codex weekly subscription meter, the UI and API appear to disagree materially.
I understand the subscription meters may not be raw-token based. The ask is for an explanation or exposed breakdown of the accounting inputs, especially:
prolitequota/weighting/backend policy changedWithout a per-model/per-feature breakdown, users cannot reconcile local telemetry with the account meter.
I’ve also escalated this privately to OpenAI Support for account-side inspection, referencing this issue and the timestamped endpoint sequence above.
The support request asks them to inspect the 2026-06-19 16:14–16:28 AEST window specifically, where the authenticated endpoint moved from 6% → 19% on the primary 5-hour meter and 51% → 53% on the weekly meter, while local telemetry only increased by ~1.22M counted tokens / 6 model calls.
The main requested clarification is whether cached tokens, tool schemas/results, hidden initialization/system context, retries, background/session-management calls, title/compaction calls, or model/feature multipliers are contributing to the subscription meter in a way users cannot see from local logs.
If Support provides a ticket/reference ID or explanation, I’ll add the sanitized outcome here.
Update: the drain has continued, and there is now an additional misleading-product-flow concern.
At 2026-06-19 19:06 AEST, the authenticated backend endpoint reported:
This is less than one day into the weekly window. The human-visible product flow also sent a message effectively saying “upgrade now because you are nearing limits.” That is not an adequate answer when the user is already on a paid Pro/5x-style plan and the dispute is that the allowance appears to be draining far faster than OpenAI’s own public wording would reasonably imply.
OpenAI’s public Codex pricing wording says:
A Pro/5x user seeing 52% of the 5-hour window and 59% of the weekly window consumed from very small recent interaction counts needs a transparent per-turn ledger, not an upsell prompt.
Additional ecosystem-impact point: this is not only a standalone Codex UI problem. Hermes Agent documents OpenAI Codex as a first-class provider path via ChatGPT OAuth / Codex models, including
openai-codex (ChatGPT Plus/Pro)and Codex-specific runtime support. If the underlying Codex quota is silently reweighted or mis-accounted, it breaks downstream paid agent workflows built around that documented provider path.Requested again:
plan_type: proliteand how that maps to paid Pro/5x entitlement.Same here on a ChatGPT Pro / Pro 5x account.
I primarily use Codex through the CLI. Over the last ~2 days, the 5-hour usage window has started draining much faster than normal, despite my workflow being basically the same as before.
Same here for 10x account
This needs a reconciliation trace per quota window. For each run, record model, effort, input uncached, input cached, output, reasoning, tool schema bytes, tool result bytes, retries, background calls, and backend window id. Then compare local ledgers against backend used_percent deltas by timestamp. Without origin buckets, every unexplained drain looks like ordinary token use.
---
_Generated with ax._
I'm using the pro20x plan, and I think the current 5-hour allowance is only 30% of what it was before. I'm using my own modified superpowers. Previously, with two threads running continuously for 5 hours, it would only use about 50% of the allowance, but now it only lasts for 3 hours.
Facing same issue with plus plan.
Why don’t they make an official announcement instead of doing this?
Additional datapoint from a Pro 5x user:
I observed what looks like an abnormal Codex 5-hour quota drain. After sending a single instruction, around 8% of my 5-hour usage balance was consumed within roughly 1 minute and 30 seconds. At that point, only about five commands had been executed, and the task had not involved a long-running process or an unusually heavy workload.
This drain rate seems much higher than expected for the amount of visible activity. It makes me suspect either:
Expected behavior: the 5-hour meter should decrease in a way that is understandable and broadly proportional to the actual amount of work performed, or the product should expose enough usage breakdown to explain why a short interaction with only a few commands consumed such a large percentage of the quota.
Please investigate whether Pro 5x quota accounting changed recently or whether this is a metering regression.
Adding a Pro account data point specifically for the historical-usage mismatch described in this issue.
I inspected local Codex
token_count.last_token_usageevents across prior months and compared them with the current quota behavior. The main thing that stands out is that June is elevated, but it is not my highest usage month, while the 5-hour / weekly limit pain is much worse than before.Monthly local telemetry summary:
| Month | Raw
last_totalsum | Approx uncached input + output | Active days ||---|---:|---:|---:|
| 2026-03 | 28.0B | 1.53B | 24 |
| 2026-04 | 9.47B | 706.8M | 23 |
| 2026-02 | 6.31B | 243.6M | 16 |
| 2026-06 | 5.34B | 288.1M | 22 |
| 2026-05 | 3.78B | 151.6M | 20 |
So June is not light usage, but it is also not an outlier compared with earlier months where I did not experience this kind of quota drain.
A daily comparison shows the same pattern:
last_totallast_totalThose are comparable local-log usage days, but June 22 was when I had to use the new reset/banked reset function because the 5-hour and weekly limits became effectively unusable.
One additional anomaly that may be relevant to this thread: after the manual reset, local
rate_limitsstate did not look cleanly synchronized across active sessions.Observed reset sequence:
primary.used_percent=39,secondary.used_percent=96primary.used_percent=0,secondary.used_percent=0token_countevents reported the new reset counters, while other active events still showed old or exhausted weekly state around96-100%.This may be a separate issue from raw usage weighting, but it fits the broader concern here: local telemetry and the quota meters are difficult to reconcile, and comparable or heavier historical usage did not behave this way.
What would make this diagnosable is a quota ledger per turn/window showing:
Without that, it is hard to tell whether the current behavior is expected reweighting, stale reset state, duplicated accounting, or a regression in the 5-hour meter.
Same here, my usage is abnormally high.
According to Codex I used almost 30% of my weekly Pro(x5) quota within a single day. This is in stark contrast to my usage the weeks before. I used 99.5M Tokens on week of Jun 28 (this week) AND I was using a free reset on Sunday (28th) AND I was granted a free reset, BUT still I am already down to 71%.
Last week I used 633M Tokens on week of Jun 21st, the week before I was using 405.1M Tokens on week of Jun 14th, and both weeks DID NOT reach the weekly Limit at all.
Please look into this!