Codex 5-hour usage meter appears to consume much faster than comparable historical usage

Open 💬 19 comments Opened Jun 18, 2026 by Truck0ff
💡 Likely answer: A maintainer (github-actions[bot], contributor) responded on this thread — see the highlighted reply below.

Summary

I am seeing a major discrepancy between local Codex/Hermes usage telemetry and the Codex 5-hour usage meter. Recent usage appears to consume a much larger percentage of the 5-hour allowance than historically comparable or much larger sessions.

This may indicate a regression in Codex quota accounting, prompt-cache accounting, context/tool-schema weighting, or effective allowance size.

Current observed Codex usage

From the authenticated Codex usage endpoint (/backend-api/codex/usage), redacted and summarized:

  • Plan type: prolite
  • Primary window: 5 hours / 18,000 seconds
  • Primary used_percent: 33%
  • Secondary weekly used_percent: 5%
  • limit_reached: false

A tiny live Codex probe succeeded, so the account was not hard rate-limited at the time of measurement.

Local usage comparison

Across local Hermes/Codex usage ledgers for the same environment/accounts:

Current combined last 5-hour usage:

  • ~5.53M total tokens, including cached tokens

Historical rolling 5-hour windows in the last 30 days include:

  • ~91.27M tokens
  • ~91.21M tokens
  • ~87.59M tokens
  • ~80.78M tokens
  • ~78.01M tokens
  • ~75.13M tokens

The current 5-hour period is only around 6.1% of the largest recent 5-hour local usage window, yet the Codex usage endpoint reports 33% of the current 5-hour allowance consumed.

Why this seems anomalous

This account has had the subscription for months and has previously run much larger sessions with much larger local token payloads without the 5-hour allowance moving this aggressively.

The current behavior feels like one of:

  1. the effective 5-hour allowance was reduced,
  2. Codex/GPT-5.5 usage weighting changed,
  3. cached/tool/context/schema tokens are being counted differently,
  4. prompt caching is failing or no longer credited correctly,
  5. or the usage endpoint/UI is reporting a different allowance/accounting model than before.

Related public reports

Similar Codex usage-limit complaints appear in:

  • #18422 — usage exhausted too quickly / feels more restrictive than before
  • #19607 — single prompts heavily draining 5h/weekly usage
  • #24560 — severe quota drain, suspected prompt-cache/context issue
  • #24431 — GPT-5.5 reliability degradation wasting Pro quota
  • #22650 — usage limit reached while status/web showed remaining quota
  • #27083 — extension reports usage limit despite available quota
  • #27850 — blocking with 100% left
  • #28246 — quota windows anchoring/reset behavior concerns

Expected behavior

Comparable local token usage should consume a broadly comparable fraction of the Codex 5-hour allowance over time, or the product should expose clear accounting showing what changed.

Requested clarification

Can OpenAI confirm whether any of the following changed recently?

  • Codex 5-hour allowance size
  • GPT-5.5 quota weighting
  • cached token accounting
  • tool/schema/context token accounting
  • primary/secondary window behavior
  • prolite plan allowance behavior

View original on GitHub ↗

19 Comments

github-actions[bot] contributor · 1 month ago

Potential duplicates detected. Please review them and close your issue if it is a duplicate.

  • #28498
  • #28727
  • #28687
  • #27908
  • #27242

Powered by Codex Action

dragon-zhang-woo · 1 month ago

Tokens consumption were not revealed after finishing a task (

Truck0ff · 1 month ago

Additional datapoint with local token telemetry, sanitized as combined local agent usage only:

OpenAI Codex usage endpoint:

  • Current check: 2026-06-18 11:20:11 AEST
  • Primary 5-hour window: 79% used; reset in 2:26:41; reset at 2026-06-18 13:46:51 AEST
  • Secondary weekly window: 12% used; reset in 6 days, 21:26:41; reset at 2026-06-25 08:46:51 AEST
  • limit_reached: false

Local combined agent telemetry for the current rolling 5-hour window:

  • Total tokens, including cached tokens: ~13,833,620
  • Input tokens: ~10,812,673
  • Cached input tokens: ~2,927,616
  • Output tokens: ~71,502
  • Reasoning tokens: ~21,829
  • API calls: 149
  • Tool calls: 105

Historical local rolling 5-hour windows from the same aggregate ledgers include:

  • ~91.27M tokens
  • ~91.21M tokens
  • ~87.59M tokens
  • ~80.78M tokens
  • ~78.01M tokens
  • ~75.13M tokens

The current local combined 5-hour token volume is far below those historical 5-hour windows, yet the Codex usage endpoint is showing 79% of the primary 5-hour window used. This reinforces the concern that the discrepancy is not simply large local token volume, unless the effective allowance or token weighting/accounting changed substantially.

Truck0ff · 1 month ago

Clarification: I understand Codex subscription limits may not be raw-token based. The anomaly is that the effective drain rate changed materially for comparable local workflows, while the usage endpoint exposes only percentages and no accounting breakdown. Please expose or confirm the weighting inputs used for the 5-hour meter: request count, model multiplier, context length, cached-token handling, tool/schema tokens, speed mode, and weekly-window interaction.

stealth-Lee · 1 month ago

确实是,10分钟干完了5小时限制

MDGChamomile · 1 month ago

I'm experiencing the identical discrepancy on a plus-tier account.

  • Plan type: plus
  • Weekly used_percent: 34%
  • 5-hour used_percent: 100%

From my local Hermes/Codex telemetry for the same weekly window:

  • input tokens: ~7.2M
  • output tokens: ~26K
  • reasoning tokens: ~6K
  • cache_read tokens: ~2.7M
  • Estimated billed (cache@ 0.5x): ~8.6M tokens

This feels like either:

  1. the effective weekly allowance has been silently reduced,
  2. the weighting on cached/tool-schema/context tokens has changed,
  3. or there is a regression in the usage endpoint's accounting model.
eva763057345-lab · 1 month ago

+1 to this. I am encountering the exact same regression on the Pro plan. My quota and top-up budget were completely wiped out in under 5 hours today, which has never happened before. It seems the token consumption or rate-limit calculation is heavily bugged right now, just as described in this issue. It is literally impossible to maintain a normal development workflow when a 5-hour budget can't even get you through the morning.

Truck0ff · 1 month ago

Additional fresh datapoint from the same account, with a short-sequence endpoint trace and local ledger delta.

Authenticated endpoint: https://chatgpt.com/backend-api/codex/usage
Plan reported: prolite
Local timezone: AEST (+1000)

Endpoint sequence during one short troubleshooting/chat session:

2026-06-19 16:14 AEST
  primary 5-hour: 6%
  secondary 7-day: 51%
  GPT-5.3-Codex-Spark secondary: 16%

2026-06-19 16:19:52 AEST
  primary 5-hour: 13%
  secondary 7-day: 52%
  GPT-5.3-Codex-Spark secondary: 16%

2026-06-19 16:23:27 AEST
  primary 5-hour: 18%
  secondary 7-day: 53%
  GPT-5.3-Codex-Spark secondary: 16%

2026-06-19 16:27:26 AEST
  primary 5-hour: 19%
  secondary 7-day: 53%
  GPT-5.3-Codex-Spark secondary: 16%

During the first part of that sequence, local session counters changed by approximately:

model API calls: +6
tool-result messages: +16
input tokens: +1,014,026
cache-read tokens: +29,696
output tokens: +6,071
reasoning tokens: +1,149
total local counted tokens: +1,217,742

Over that same evidence window, the endpoint primary 5-hour meter moved from 6% to 18% (+12 points), and the weekly meter moved from 51% to 53% (+2 points). That works out to roughly 101k local counted tokens per 1 percentage point of the 5-hour meter in this short sample, which is difficult to reconcile without knowing the hidden weighting/accounting model.

Also, the human-visible UI was observed as showing a different weekly percentage around the same time (47%) while the authenticated endpoint reported 53%. If those are different quota pools, the UI should label them differently; if they are intended to represent the same Codex weekly subscription meter, the UI and API appear to disagree materially.

I understand the subscription meters may not be raw-token based. The ask is for an explanation or exposed breakdown of the accounting inputs, especially:

  • cached-token handling
  • tool schema / tool result handling
  • system prompt and hidden initialization context
  • failed/retried/background/session-management calls
  • model/feature multipliers
  • whether the 5-hour and weekly windows use different accounting rules
  • whether a recent prolite quota/weighting/backend policy changed

Without a per-model/per-feature breakdown, users cannot reconcile local telemetry with the account meter.

Truck0ff · 1 month ago

I’ve also escalated this privately to OpenAI Support for account-side inspection, referencing this issue and the timestamped endpoint sequence above.

The support request asks them to inspect the 2026-06-19 16:14–16:28 AEST window specifically, where the authenticated endpoint moved from 6% → 19% on the primary 5-hour meter and 51% → 53% on the weekly meter, while local telemetry only increased by ~1.22M counted tokens / 6 model calls.

The main requested clarification is whether cached tokens, tool schemas/results, hidden initialization/system context, retries, background/session-management calls, title/compaction calls, or model/feature multipliers are contributing to the subscription meter in a way users cannot see from local logs.

If Support provides a ticket/reference ID or explanation, I’ll add the sanitized outcome here.

Truck0ff · 1 month ago

Update: the drain has continued, and there is now an additional misleading-product-flow concern.

At 2026-06-19 19:06 AEST, the authenticated backend endpoint reported:

plan_type: prolite
primary 5-hour used: 52%
secondary weekly used: 59%
weekly reset: 2026-06-25 08:46 AEST
GPT-5.3-Codex-Spark separate bucket: 16% weekly used, 0% 5-hour used
rate_limit_reset_credits.available_count: 2
credits.balance: 0

This is less than one day into the weekly window. The human-visible product flow also sent a message effectively saying “upgrade now because you are nearing limits.” That is not an adequate answer when the user is already on a paid Pro/5x-style plan and the dispute is that the allowance appears to be draining far faster than OpenAI’s own public wording would reasonably imply.

OpenAI’s public Codex pricing wording says:

  • Pro: “5x or 20x higher rate limits than Plus.”
  • Plus GPT-5.5 local messages: 15–80 / 5h.
  • GPT-5.5 usage averages 5–45 credits per message.

A Pro/5x user seeing 52% of the 5-hour window and 59% of the weekly window consumed from very small recent interaction counts needs a transparent per-turn ledger, not an upsell prompt.

Additional ecosystem-impact point: this is not only a standalone Codex UI problem. Hermes Agent documents OpenAI Codex as a first-class provider path via ChatGPT OAuth / Codex models, including openai-codex (ChatGPT Plus/Pro) and Codex-specific runtime support. If the underlying Codex quota is silently reweighted or mis-accounted, it breaks downstream paid agent workflows built around that documented provider path.

Requested again:

  1. Confirm whether Pro/5x Codex quota weighting changed around June 16–19.
  2. Explain why the backend reports plan_type: prolite and how that maps to paid Pro/5x entitlement.
  3. Provide per-turn/per-feature usage receipts: model, client/surface, bucket charged, input/cache/output/reasoning tokens, tool/MCP/image/retry/startup/compaction/background charges, and any multipliers.
  4. Restore wrongly consumed allowance/credits/resets if this is mis-accounted.
  5. Clarify why the product is pushing “upgrade because you are nearing limits” instead of explaining the sudden drain on an already-paid plan.
caioarotolo · 1 month ago

Same here on a ChatGPT Pro / Pro 5x account.

I primarily use Codex through the CLI. Over the last ~2 days, the 5-hour usage window has started draining much faster than normal, despite my workflow being basically the same as before.

hzhua · 1 month ago

Same here for 10x account

Necmttn · 29 days ago

This needs a reconciliation trace per quota window. For each run, record model, effort, input uncached, input cached, output, reasoning, tool schema bytes, tool result bytes, retries, background calls, and backend window id. Then compare local ledgers against backend used_percent deltas by timestamp. Without origin buckets, every unexplained drain looks like ordinary token use.

---

_Generated with ax._

FYZAFH · 28 days ago

I'm using the pro20x plan, and I think the current 5-hour allowance is only 30% of what it was before. I'm using my own modified superpowers. Previously, with two threads running continuously for 5 hours, it would only use about 50% of the allowance, but now it only lasts for 3 hours.

mUsmanGoally · 28 days ago

Facing same issue with plus plan.

duynguyenbui · 28 days ago

Why don’t they make an official announcement instead of doing this?

zixuanjiang332 · 27 days ago

Additional datapoint from a Pro 5x user:

I observed what looks like an abnormal Codex 5-hour quota drain. After sending a single instruction, around 8% of my 5-hour usage balance was consumed within roughly 1 minute and 30 seconds. At that point, only about five commands had been executed, and the task had not involved a long-running process or an unusually heavy workload.

This drain rate seems much higher than expected for the amount of visible activity. It makes me suspect either:

  1. a bug in quota/usage accounting,
  2. a recent change in how tool calls, hidden context, cached tokens, or command execution are weighted,
  3. or a reduction/change in the effective allowance for the Pro 5x tier that is not clearly surfaced in the UI.

Expected behavior: the 5-hour meter should decrease in a way that is understandable and broadly proportional to the actual amount of work performed, or the product should expose enough usage breakdown to explain why a short interaction with only a few commands consumed such a large percentage of the quota.

Please investigate whether Pro 5x quota accounting changed recently or whether this is a metering regression.

loredan1994 · 25 days ago

Adding a Pro account data point specifically for the historical-usage mismatch described in this issue.

I inspected local Codex token_count.last_token_usage events across prior months and compared them with the current quota behavior. The main thing that stands out is that June is elevated, but it is not my highest usage month, while the 5-hour / weekly limit pain is much worse than before.

Monthly local telemetry summary:

| Month | Raw last_total sum | Approx uncached input + output | Active days |
|---|---:|---:|---:|
| 2026-03 | 28.0B | 1.53B | 24 |
| 2026-04 | 9.47B | 706.8M | 23 |
| 2026-02 | 6.31B | 243.6M | 16 |
| 2026-06 | 5.34B | 288.1M | 22 |
| 2026-05 | 3.78B | 151.6M | 20 |

So June is not light usage, but it is also not an outlier compared with earlier months where I did not experience this kind of quota drain.

A daily comparison shows the same pattern:

  • 2026-06-03: ~597.8M raw last_total
  • 2026-06-22: ~582.9M raw last_total

Those are comparable local-log usage days, but June 22 was when I had to use the new reset/banked reset function because the 5-hour and weekly limits became effectively unusable.

One additional anomaly that may be relevant to this thread: after the manual reset, local rate_limits state did not look cleanly synchronized across active sessions.

Observed reset sequence:

  • 2026-06-22 17:48 EEST, before reset: primary.used_percent=39, secondary.used_percent=96
  • 2026-06-22 17:49 EEST, immediately after reset: primary.used_percent=0, secondary.used_percent=0
  • After that, some active token_count events reported the new reset counters, while other active events still showed old or exhausted weekly state around 96-100%.

This may be a separate issue from raw usage weighting, but it fits the broader concern here: local telemetry and the quota meters are difficult to reconcile, and comparable or heavier historical usage did not behave this way.

What would make this diagnosable is a quota ledger per turn/window showing:

  • model and client surface
  • bucket charged
  • input, cached input, output, and reasoning tokens
  • tool/schema/result/background/retry/compaction charges
  • model or feature multipliers
  • backend quota window id / reset id

Without that, it is hard to tell whether the current behavior is expected reweighting, stale reset state, duplicated accounting, or a regression in the 5-hour meter.

midnighter90 · 19 days ago

Same here, my usage is abnormally high.

According to Codex I used almost 30% of my weekly Pro(x5) quota within a single day. This is in stark contrast to my usage the weeks before. I used 99.5M Tokens on week of Jun 28 (this week) AND I was using a free reset on Sunday (28th) AND I was granted a free reset, BUT still I am already down to 71%.

Last week I used 633M Tokens on week of Jun 21st, the week before I was using 405.1M Tokens on week of Jun 14th, and both weeks DID NOT reach the weekly Limit at all.

Please look into this!