Burning tokens very fast
Open 💬 627 comments Opened Mar 13, 2026 by cy-ooi88
💡 Likely answer: A maintainer (github-actions[bot], contributor)
responded on this thread — see the highlighted reply below.
What version of the IDE extension are you using?
26.311.21342
What subscription do you have?
Business
Which IDE are you using?
VS Code
What platform is your computer?
Microsoft Windows NT 10.0.22631.0 x64
What issue are you seeing?
Am I the only one still seeing my tokens burning very fast following today's extension update? e.g. just by writing 1 or 2 prompts, usage drops by 1%, within 2 hours of working I managed to burn through ~20% of my tokens. Tried using GPT5.3 and 5.4, High reasoning effort. I know High reasoning effort will consume more quickly, but is it considered normal to consume one fifth of the tokens within 2 hours even for business accounts? Doesn't make sense to me, but please clarify
What steps can reproduce the bug?
just by writing reasonably simple prompts. GPT5.3 and 5.4, High reasoning effort.
What is the expected behavior?
Last week token usage was like 25% slower.
Additional information
_No response_
Showing cached comments. Read the full discussion on GitHub ↗
626 Comments
Potential duplicates detected. Please review them and close your issue if it is a duplicate.
Powered by Codex Action
Not a duplicate. I am aware about the separate reports that are now closed. But with my VS Code extension updated today, it seems the issue has not been resolved, at least for my account.
Please refer to my post here for reasons why you may be seeing increased usage.
Not a duplicate. I've reviewed etraut's post on the closed thread and none of the listed factors apply to my case.
I'm on a Pro account. Over the past couple of days, my weekly limit has been draining dramatically faster despite the same workload, regardless of model (GPT-5.3 or 5.4), reasoning effort, or channel (Codex App, CLI, VS Code). No sub-agents, no experimental features enabled. I've burned through ~70% of my weekly limit in a single day on what I'd consider normal, moderate usage. This is a clear regression from last week's consumption rates.
The regional sync fix mentioned in the closed thread doesn't appear to explain this. The drain is consistent and directional, not oscillating.
Thanks @nishikawa7863, exactly, this is what I am experiencing. I will do a little more evidence / data gathering & post here.
Same here never ever burned this much usage on CLI 5.3 codex med.... Lost arround 25% weekly within few prompt on only one project without using fast or anything else.. My 5hours usage goes under 20% which never hapened either.
If this is considered 2 times the normal usage it's straight unusable for any coding activites at least for plus member.
Edit : Just reached 5hours limit for the first time ever.
Token usage: total=742,555 input=697,188 (+ 9,077,504 cached) output=45,367 (reasoning 11,450)
Almost used half of my weekly session on this one.
I don't really think this has been solved, with all due respect.
First the issue is being attributed to "network", when deflected by actual users the thread gets closed and now "its fixed and thread closed" but we still are experiencing a definitively more than 30% usage attributed to the new model.
That's our opinion based on usage. Thanks.
I'm experiencing a similar issue on my PRO account and the usage spiked in the last few days despite no changes in my usual workflow
<img width="2408" height="572" alt="Image" src="https://github.com/user-attachments/assets/cb20cb3d-4f7d-4455-8bf2-fdadd2d83558" />
I'm using codex mainly via VSCode extension, current version is 26.311.21342
Yeah I was the original poster on that thread @etraut-openai and it's still not fixed. Still using 5.3 codex like I've always been, not particularly high usage compared to previously, but usage limits are still burning very fast.
It's insane how OpenAI insists everything is normal... there are people with Pro subscriptions that are hitting 70% of the weekly limit in 1 day... how is that even normal, specially considering the x2 promo?
Either...
a) the x2 promo is not being applied to everybody
b) there is a bug somewhere
c) the limits were significantly nerfed
And at this rate, Codex will be useless in April considering a lot of people will hit their weekly limits in 2/3 days.
On February/March, I've been often hitting my weekly limits and yet, there were weeks in October when I spent more or the same number of tokens.
I completely agree something’s not right. I’ve been using Codex for a few months, and I’ve had enabled subagents since they were released. I usually was ending my week with a 40% remaining usage and using it without worries.
For the past two weeks, I’ve been struggling on my 3rd day and using much less subagents.
I noticed that my quota decrease quite fast when I encounter compaction loop bug.
The compaction launch but the context window stay quite high (>75%), it compacts again and so on.
It doesn’t matter; they’ve already made it very clear that they’re not going to do anything. Apparently, it’s our fault because we’re idiots and we’ve been using 2x or 5.4 without even knowing it. We all changed our usage pattern with Codex at the same time, and there’s nothing they can do because, supposedly, it’s our fault
My experience is the same. Whilst previously, my Plus subscription lasted the whole week, then now I basically drain one Plus subscription per day.
Same experience. Relatively light usage, almost never hit 5 hour or weekly limits (sometimes I did via heavy weekend usage). Suddenly burned the weekly limit in 1.5 days (light usage, not weekend). I did use 5.4 but I don't see how that explains how rapid it depleted. I use codex cli only (Linux).
Based on my total token and cost data from ccusage (which may not be 100% accurate) it used to be very rare for me to hit the weekly limit. Now it’s happening much more often, even with a 2× promo. And yet, neither my total tokens nor my cost are anywhere near 2× my previous peak. I really don’t understand how this can be called 2× when neither my experience nor my data support that...
<img width="1056" height="460" alt="Image" src="https://github.com/user-attachments/assets/ae582cf0-9d58-4f72-877c-4129b20da4dc" />
<img width="1057" height="472" alt="Image" src="https://github.com/user-attachments/assets/d514e265-ab24-4789-a7f9-adf0ab97fdcb" />
I use my ChatGPT account on two computers, that's why there are 2 colors, it's combining ccusage data from both.
Here is my extremely simple experiment, I hope it helps OpenAI team clarifies what the expected behavior is.
EDIT: this experiment consumed 22 credits. Is it considered "normal" ?
<img width="1377" height="309" alt="Image" src="https://github.com/user-attachments/assets/e9d6a4b4-2914-4659-aa65-4e20b74f035e" />
<img width="1222" height="390" alt="Image" src="https://github.com/user-attachments/assets/c5719a89-e0fa-471b-bbe9-ccea43431cd9" />
I'm experiencing the exact same issue, and there's no way this is acceptable — even before factoring in the 2x multiplier. We should never be having this conversation at 2x to begin with, yet consumption is so extreme that usage is completely drained within a single day.
If this is considered normal, developers will be cancelling their subscriptions en masse come April 11th. Claude Code gives significantly more tokens with Opus than Codex does at this rate — and we never had these problems with Codex ever before, Codex 5.4 was supposed to be more token-efficient, and OpenAI is racing to match Claude Code on speed — so how is it possible that this many developers are reporting the exact same problem while OpenAI insists everything is fine?
I really hope they will put more effort into finding the issue. otherwise we will need 3 subscriptions per week.
It's not fine.
Hey everyone, it is the weekend and we got to cut the OpenAI folks some slack here. In the coming week, hope that @etraut-openai acknowledges that it is an issue & look into it. OpenAI are doing great things so please continue with that, and do not let one issue spoil the whole pot of soup.
This is a cop out for not resolving the issue, sorry. The drain is drastic and still on going, I'm using 5.2-codex.
Same issue here, it is just burning tokens like crazy, I have switched to 5.3 codex model for now, and it works and uses token normally. In 5.4, each prompt takes few % of weekly usage, with 5.3 doing even more complex tasks and it does not drain that fast. Something is wrong.
hey everyone, while waiting for OpenAI team to look into this, can I suggest we all help out each other by sharing what our current "workarounds" / band aid solutions are:
For me, I had to "downgrade" to GPT-5.2 Codex (Normal speed, Medium reasoning effort) to avoid the super quick burn. GPT-5.3 Codex was still quite fast in burning.
@etraut-openai Not possible. I'm experiencing increased usage on 5.3 Codex and 5.4 both. The limits are not being respected.
Not using multiagents/subagent/fast mode.
https://github.com/openai/codex/issues/14349
Folks. I have three accounts, I blow through each in a day's worth of coding. It was OK when "weekly" limits were reset for 3 days, but now I have three accounts and all of them are out of quota for the next 4 days????
Time to try out the glorified chinese models, I guess.
And let me be clear: I've had two Google subscriptions prior to subscribing to ChatGPT. Which were plenty, and my usage never drained with both. In fact, I had more and more coming (those models weren't as good at planning/debugging though). But nevertheless!
Exactly what's happening with me, it's become a joke and almost as bad as claude. ChatGPT's selling point was the token usage, now it's pretty much nil.
I still think something is wrong with usage given I've had multiple instances of 5.2 or 5.4 counting to spark usage, see https://github.com/openai/codex/issues/13854
That's easy to spot because spark is a seperate usage bucket, but it makes me wonder if something that's not visible is happening hitting the main usage (double counting to two models, or counting cheaper models to a more expensive model, etc)
My usuage dropped from 65% remaining (used all week at work) to less than 20% remaining from one quick job on a Saturday when I'm not even working properly. 5.3 medium only, have hard disabled subagents since last week's usage drama. Something is not right here.
Since 5.4 came out, usage is dramatically faster. Been using 5.3-codex hoping for better usage. Entire team on business plan started seeing issues after the 5.4 release.
I'm using two Pro subscriptions, and after some testing, I found that using up a full 5-hour session still consumes about 30% of the weekly limit. After an x2 rate limit event, does that mean the exact same 5-hour usage now becomes 60% of the weekly limit?
If that's by design, then ChatGPT Pro is not sustainable for heavy users at all. I'm already on two Pro subscriptions, and after the promotion, I might need four. At that point, I should just buy some 6000 Blackwell GPUs and run local models instead.
la pacchia e' finita
I am having a problem with a business / Team account - it reset unexpectedly on Tuesday morning and since then has been consuming very fast. Today I used my personal account on personal projects which is going down much slower and is as I would expect. I'm not using any of the other features (fast mode, agents, extra high). The contrast between the two is very clear, and the Team account would normally have more usage for a week.
Using 5.3-codex medium is burning through my usage much faster than it was for the same workload, and it feels like a regression from the issues last week. What’s most frustrating is the inconsistency, from one day to the next, I have no sense at all of how much work I can actually do because the usage drain changes so dramatically. As a user this unpredictability feels really bad.
Yeah, I burned through my quota in just 2 days. It changed either yesterday or the day before. The weekly limit reset, and within 48 hours, I was already down to less than 10%. Definitely a recent change, and it basically killed the limits.
The token burn is real after 5.4. It started with 5.3-codex, but this is just crazy. Imagine what will happen after April 1st. It's not just the VS Code extension, it's the CLI as well.
Same issue. 5 business accounts all burned through weekly limits within a couple of days.
Another +1 on token burn using 5.3-codex medium. Blew weekly allowance rapidly in a couple 5 hour sessions, never hitting the 5hr window limit. Even compared to claude with similar usage, seem weekly measuring is off. Only noticed after 5.4 release.
I have new data for you guys. My subscription expired yesterday, until yesterday my token usage was normal, after resubscribing, my usage has increased significantly. One request to write a 300 letter doc used 3%, where before this it used 1% or the meter did not move at all.
A recap of what we have already investigated:
A recap of known factors under user control that can result in higher usage than in previous weeks:
multi_agentfeature, subagents will typically consume tokens at a higher rate than if you are not using subagents. As we continue to refine this experimental feature, some changes may affect token consumption. If you are concerned about usage consumption, leave this feature in its default (disabled) state for now.Based on the reports in the above thread, it sounds like some are still experiencing higher-than-expected (or higher-than-previously-experienced) usage consumption. We don't think this problem is affecting the majority of Codex users. Many users have reported that they are not seeing higher-than-usual usage. That leads me to believe that there's some common thread among those of you who are observing this. Let's see if we can figure out what it is.
There are a couple of theories in the thread above:
multi_agentexperimental feature.Here are some other ideas:
/reviewcommand for local code reviews?Any other theories?
The only one of those relevant to me is a lot of /review
But also the issue #13854 I submitted might be related
No to all of your questions except "Do your sessions tend to be really long and require many compactions?". I tend to have very very long sessions with tens or hundreds of compactions, on a large-ish codebase (200k+ loc).
Note: that was already the case before March 5 when I first reported usage dropping too quickly, I didn't change my codex habits at all.
hi @etraut-openai , thanks for looking into this. Not sure how the below info helps, but here you go.
Do you frequently use the /review command for local code reviews?
Do your sessions tend to be really long and require many compactions?
Do you use long AGENTS.md?
Do you have a large number of MCPs or skills?
I don't know. I'm on the PRO plan, and between March 11th and 13th, I used up 50% of my weekly limits on 5.4 High in /fast mode. On March 14th, I realized my limits were burning up way too fast, so I turned off /fast mode. It's still really frustrating, though, and I'll probably run out of limits completely by tomorrow. And that's even with the 2x promo limits!
<img width="427" height="238" alt="Image" src="https://github.com/user-attachments/assets/c90e470b-85d8-40e7-ae2e-5cdf7fe26866" />
They're passing the buck and blaming users, its not a "small" subset, reddit is flooded with complaints.
As above, multi agent is disabled in my config and I'm only using 5.3 medium. I use /review, but no more than usual. My biggest usage day was a Saturday that i wasn't working and only checked a couple of things through ssh on my phone (on the same repo that i worked literally all day the two previous). There is no possible way that my increased usage is explained by any factors in your list.
Hmm. I only use long sessions here, sometimes jumping between topics.
But @etraut-openai - why is it more expensive? compression is done completely from cache, is it not?
I have been mainly working with 5.3 Codex, maintaining the same usage pattern as before. However, starting March 12, my weekly usage consumption increased dramatically.
I was initially using it through the VSCode extension, but because the token usage seemed too high, I switched to the Codex Windows app. Unfortunately, that did not improve the situation. In fact, the token consumption still remained very fast.
Because of this, I am now mainly using the “high” setting instead of “xhigh” to reduce usage.
I am not using the API, but rather accessing Codex through a GPT Pro subscription. I am also not using multi-agent setups or similar advanced configurations.
Given that my usage pattern has not changed, something clearly seems wrong.
If this is caused by a feature that was enabled due to a recent update, then it would mean users are unknowingly consuming significantly more usage without being aware that such a feature has been activated, which would clearly be problematic.
Could this issue be related to updates released around the time the Codex Windows app was launched?
<img width="622" height="215" alt="Image" src="https://github.com/user-attachments/assets/4370bd33-ff6e-41bb-861a-707f41cbea7a" />
This is just my hypothesis, but it seems like the issue may be related to "turn"
When I used 5.2 before, it usually worked by completing the task in one go and then sending a final message. Now, however, it reports progress intermittently during the process, and it feels like this has caused the token usage to increase significantly.
I’m currently using 5.2 through the codex-cli version before the 5.4 release (0.9x.0), and it seems to be a bit better.
I appreciate you engaging on a weekend, @etraut-openai.
Of your questions, @etraut-openai, only long sessions for me.
I think the most important point is that it was like a switch flipped somewhere in late Feb or early March. This is the earliest report I quickly find: https://github.com/openai/codex/issues/12728
My particular issue didn't show until around Mar 1.
My only other thought is that at some point, let me estimate November, my quota went way up. Before that, I could get a lot done, but busy weeks I'd use all the quota. Between November and this issue, I couldn't use it up if I tried. So I hope this isn't the case, but maybe the November-ish shift to the generous quota for a subset of users is the anomaly and that generosity was recently and inadvertently "fixed" by some other change.
None to all of these questions. I do not have a single MCP installed -- they hugely pollute the context. My AGENTS and SKILLS havent changed since the last 2-3 weeks before this started happening, and gpt-5.4 usage is also not the main culprit as 5.3-codex also is dropping usage really fast.
I do not think that this is related to the codex-cli app either -- this seems to be an usage accounting related issue/cache tokens not being properly counted as cache tokens -- because I use opencode. So the problem has to related to usage reporting.
<img width="1263" height="762" alt="Image" src="https://github.com/user-attachments/assets/3fe550e0-e4d8-4643-86dd-84b86c38c075" />
Look at the usage graph. Given that I did not change my usage pattern that I've been using for > 30 days -- this spike at 11th Mar is unmistakable.
<img width="1257" height="448" alt="Image" src="https://github.com/user-attachments/assets/8949ba18-1a46-4bd3-a106-ed94c6f9bccf" />
The usage data makes the problem unmistakable.
The highest graph corresponds to March 12, which is the usage recorded after my weekly limit was reset.
In other words, the three graphs on the far right represent March 12, 13, and 14, in that order.
March 13 still shows abnormally high usage compared to my normal baseline.
The critical point is that my workload on March 12 and 13 was not materially different from my usual workload.
Despite that, usage surged dramatically. Based on the graph, it appears to be roughly double my normal level.
As for March 14, token consumption was so clearly abnormal that I deliberately reduced my workload to about one-third of my usual average.
Only after cutting my workload down that far did token usage return to something close to normal.
That is not a minor fluctuation.
It strongly suggests that something is wrong with how usage is being measured or charged.
This is clearly a serious issue, not normal variation.
Additional information:
I asked codex which operations were burning the most tokens:
I then asked "Can you identify where token usage optimizations could occur without reducing fidelity and how much was possibly wasted without without the optimization?"
I then asked "Can you identify in this repo if a prompt directly caused 'Heavy log/session introspection' or was codex deciding to do this on its own?"
Yes.
So perhaps this is codex internal tooling causing spikes in usage?
I can also corroborate significantly increased usage.
I used to run 5.4 xhigh pretty much all the time and never hit 5 hour or weekly limits, and now I'm hitting my 5 hour limit within 1-2 hours and hitting weekly limit in a couple days. Haven't really changed what I'm doing or using Codex for, and spending roughly the same amount of time using Codex on a day-to-day (within the limits I'm hitting now). I have also disabled agents (I was using them before), have no MCPs and dropped down to high instead of xhigh.
I am on the Plus subscription.
Specifically:
Do you frequently use the /review command for local code reviews?
No, I never use /review
Do your sessions tend to be really long and require many compactions?
Typically yes. But even in new chats, usage is still significantly higher than it was previously with longer chats.
Do you use long AGENTS.md?
I do not have an AGENTS.md
Do you have a large number of MCPs or skills?
No
You can see here on the 13th March:
<img width="179" height="343" alt="Image" src="https://github.com/user-attachments/assets/96cb8f0c-3517-485e-9eb3-b2b85ae1d11f" />
How quickly my 5 hour limit was used up. From my refresh, it took 1 hour 30 (2:30pm to 4pm) to hit 87% on my 5 hour limit.
Are they going to do something about it???
If increased usage is related to a client-side change, then downgrading to an earlier version of the CLI should result in reduced usage. If someone is able to demonstrate that there was a specific version of the CLI where usage rates increased, that would help us isolate the cause.
Don't think this is a client side change - rather probably somehow usage not being computed/accounted for properly - because I'm on OpenCode on the OpenAI codex sub (and not a codex-cli related change I think) and seeing increased usage (and no - this is not due to opencode tool-calling more frequently - have been on this for > 3 months and it was fine) - this started happening on 10th/11th March for me.
Subscription: Business. happening on both accounts on my team
@eternal-openai Ty for referencing my post here. Plus plan, Gpt 5.4, Extra high, Standard speed,n o MCP, 80 lines AGENTS tops, WSL with Codex App installed on Windows.
My user ID:
user-Ifae4RD8fJVLSlCOhx3wjeL3Also, I have this statistic from Tokscale. 🤔 Maybe it's the new reality — massive token usage. Earlier, even when I tried hard, I couldn't hit my weekly limits, but now even on PRO I need to think about my token usage.
<img width="638" height="534" alt="Image" src="https://github.com/user-attachments/assets/def34e8c-3fd5-4603-b8b2-12251582a5c1" />
I was using gpt 5.4, but then switched back to 5.3 codex due to the token issue, yet it persists. No multi agent use, never use /review, I dont use fast, I dont use extended context window, no change in session lengths, agents.md under 100 lines, 8 skills, context7 is only MCP. As others have said, I had to physically try to hit the usage limits even before the 2x usage. Now even with the 2x usage, 5 different business accounts all burn through 5 hours windows within 1-2 hours and weekly usage within a couple of days.
How can we do that when we're all out of usage lol
If I remember correctly it was on 0.110.0. I'm fairly sure I upgraded from 107 to 110 and then this started. But it could be that this started, and then I upgraded
Just chiming in, I've got 4 pro accounts which just about last me a week perfectly under normal use (sometimes I have to ration the last day!). This weekend I burned through 2 plans in 36 hours which is absolutely wild even for me (all four plans are now used up which considering they reset on Wednesday is crazy ahead of schedule for me).
I am using a range of 0.110-0.114, across 2 Macs, 2 Windows and a Linux instance. Ive been constantly hitting 5 hour limits all of last week (i'd only hit it twice before over the last year). Using 5.2 across the board, no sub agents, no MCPs on one project and Bitbucket/Jira on the other, no /fast, its as bog standard as it comes I think. Unfortunately I dont use any of the token tracking tools so I cant give much more specifics than that.
My gut feeling is an accounting thing, I cant see otherwise what would have changed. My workflow and work items have been extremely consistent the last couple of weeks so I really struggle to see what might have changed my end. Just my 2c
Same here, the speed at which the limits are being used up are insane.
Here is an example using VibePulse which counts local token usage.
20th Feb and 10th March have about the same token use shown in VibePulse, but 2-3x the codex consumption (red and orange).
I had a usage reset at 7:42 UTC on 10th March and since then consumption has been 2-3x faster than expected. This was only 2 hours before my expected scheduled reset on this account (just mentioning it as a data point since it is not affecting all users, perhaps the ad-hoc reset being same day or so close to the scheduled weekly reset could be an edge case). I am not using any of the extra features or fast mode, working only in codex.app and have switched to 5.3-codex medium but still see accelerated consumption. I have a personal account I use on the weekend which does not have this problem.
The account showing increased consumption has account id
b31d47aa-e5c3-476a-89d2-4b5601d4c652Feb 19th and March 11th show almost the same usage in the codex usage plot, but vibepulse shows $17.43 vs $7.83 of tokens. This ~2.3x factor matches my subjective impression of the increased usage (I would usually finish a week with around 50% usage left, now it is running out after 3-4 days).
<img width="1166" height="683" alt="Image" src="https://github.com/user-attachments/assets/ed549cb3-c7a5-44d4-a1e0-661f7058696c" />
So do you think we'll get our credits back, for those of us where the 2x clearly isn't applying? It feels like we're paying for extra credits if we want to have usage in line with what we should expect given the (soon-ending) promotion
For what it's worth, after taking a step back from codex a week or so ago after the issue appeared, I'm back to coding today. My burn rate is much more to my expectation. Hopefully it's a permanent change.
If I did something on my end to fix it, unfortunately, I have no idea what it was.
user-vCJq0VW2uSEmtAVWM6wTz0Hf
Frequent compaction may have something to do with it. 5.3 and 5.4 are good enough that my discipline around single purpose chats have grown lax. Still, something doesn't add up. Since the last issue thread I've disabled fast inference and multi-agent yet still burned though my weekly limit (again). While I am using 5.4 I find it hard to believe that going from 5.3 codex to 5.4 is the difference between "not knowing rate limits were a thing" to "burning through my weekly allotment by midweek".
This is ridiculous, In the past two days on a single thread work my PRO limit went from 72% to 25%. Just by using standard context, no mcps, no extra settings than default and half of this time on standard speed because fast was burning literally 1-2% per prompt. I just tried spark, I though I have a few questions about few files in repo, so I spent in total under 20minutes with spark my weekly spark usage on Spark went from 100% to 89% on PRO plan, Weekly! And I only use MacOS and Windows native updated to newest apps with standard settings. By the time I am writing this reply it went down by 1% again.
I honestly think it’s an absolute disgrace how OpenAI is treating us here. We’re paying $200 or more, and the support really leaves a lot to be desired. The issue has been ongoing for several days now, and nothing is being done about it, nor is there any real information being provided. Instead, we’re being asked to figure out the problem ourselves, which is neither our responsibility nor within our capabilities as customers. That really makes me wonder what’s going on. This is not customer-friendly at all. I’ve also already contacted OpenAI support and was basically brushed off, being told to monitor my usage and switch to lighter models. What kind of response is that?
Apparently it's a very rare issue (we're not that many commenting here, maybe a few hundred, even if ~10k users globally are affected that's still less than 1% of the codex user base. And given that there is no real alternative (let's not pretend claude code is good), and that we just pay for credits, they have no real incentive to solve an issue which may have dozens of possible causes (and it looks like many of us are having slightly different issues, likely caused by different things; and we can't exclude that for a slice of us this may be pebcak).
Yes, that’s true, I’m afraid you’ve hit the nail on the head there. I still hope something will be done and that we won’t be overlooked, even if we’re not many...
+1, I have similar observations with the token usage
i have the same workflow and got vacuum-cleaned on 2 different instances. something like this has never happened to me before. one medium task took 50% of my weekly limits. i was using my Plus account and my jobs account
on top of that, it almost felt like i was using completely different models. they all got significantly dumber right after Codex started consuming enormous amounts of tokens
More informations about my case :
I tend to have long conversations with lot of compactions.
The first compactions are quite efficient but the compaction quality degrade so much that it force me to create new thread (new feature idea : thread should be deletable not just archivable).
I don't use VPN but I have apple private relay activated
Considering we should have 2 / 1.3 > 1.53 rate limits which is more than a 50% increment over old limits at 5.3 codex era, there is no way so many users are burning at 2x or more their usual rate. It's basically an effectively 3-4x reduction in limits.
Since I reset my
.codexprofile, the issue seems resolved on my side for now. I'm still posting my setup here in case it helps compare patterns with other reports.Setup
Usage Pattern
Additional Context
I'm not sure whether this actually fixed the issue, but I:
.codexfolder completely.codexdirectorySince then, my usage limits seem to be decreasing at a more normal rate.
I noticed this issue around the time I installed a much of skills to try using the VS Code extension, but never got the chance to try them before I quickly ran out of usage with each prompt on existing chats consuming ~7% and up to 15% of weekly usage on Plus plan. I tried uninstalling them, but even though it said succeeded nothing seemed uninstalled. So I just deleted my .codex folder suggested above to reset everything, but waiting on weekly reset in 24 hours to see if it actually helps.
I think Codex is the issue; the CLI does not appear to drop as fast. Go into Codex, and my limit rapidly drops.
Opencode is also dropping usage really fast -- not sure if the issue is just the app
I don't have too many skills + anytime skills are being used - it's shown in the UI that I use - it definitely hasn't changed in the last ~1 week since this issue started. Dont think too many tokens are the issue.
I made a small reusable Codex skill to help collect the setup/context for this issue in a consistent way:
$codex-setup-report
It doesn’t measure quota usage, but it gathers the local environment/config details that seem relevant here and can generate a report comment.
For some user-side fields, I kept the workflow simple and just ask the user directly instead of trying to infer everything from session history, for example /review usage, plan mode usage, and whether the behavior is new vs the usual baseline.
Skill folder:
https://github.com/J3m5/skills/tree/main/skills/codex-setup-report
Going to try now.
Just chiming in to say I've somehow gone through 2% of my weekly limit with 1 5.1-codex-max prompt in an empty repo 👎 .
Using Plus btw.
fast mode: disabled1M context window: disabledsub-agents: disabledIf yes, which ones?
/reviewoften? yesAGENTS.md? nogpt-5.3-codex,gpt-5.4, or both? It affects all modelsSince I reset my .codex profile, the issue seems resolved on my side for now. I'm still posting my setup here in case it helps compare patterns with other reports.
Backed up ~/.codex
Ran codex logout
Deleted the .codex folder completely
Logged back in with codex login, which recreated a fresh empty .codex directory
Since then, my usage limits seem to be decreasing at a more normal rate.
------
I tried this too, and it seems like it may have helped.
Previously, every 2–4% of my 5-hour limit would consume 1% of my weekly limit. But after resetting .codex, the 5-hour limit seems to be depleting at a similar rate, while the weekly limit appears to be draining more slowly.
I haven’t tested it extensively yet, but it does seem like something meaningful might be going on.
Just tested it and I've been at it for the past hour with 5.3 with subagents. Only dropped 5% on 5h and 1% on weekly
absolutely ridiculous, but this helped. i fully wiped the codex folder except for auth.json and skills, and it solved my problem. the models got smarter instantly, and usage went down. the folder was initially 500 mb, then dropped to 800 kb. it seems like this whole mess polluted codex's context, which hurt answer quality and caused enormous token usage
Same config.toml? Otherwise config could explain it. If it works fine even with the old config.toml, then it would be really bizarre. Like why would old conversations and the size of the DB affect its intelligence and token usage?
not the same config since it got wiped. i only preserved the skills and auth.json files
I would be happy to wipe the config but reluctant to lose my historical sessions. I guess I will try this though.
You can back up ~/.codex first so you do not lose your old sessions. Then try the reset on a clean profile. If that fixes the issue, you still have the option to restore or selectively copy session data back later.
Is this working? I'm not sure how this would work with opencode -- I believe the data to be wiped might be different/interact differently but this is very weird.
In any case -- old conversations should not affect new ever? The size of the codex folder should never matter
ok, I ran the ccusage tool and it showed me for today this usage:
<img width="798" height="106" alt="Image" src="https://github.com/user-attachments/assets/17b3c95c-1ea7-4d72-80b1-04a35e6dfa7e" />
It was exclusively GPT 5.4 medium and used about 10% of my weekly usage. No 2x, no subagents, no 1M context.
<img width="1181" height="429" alt="Image" src="https://github.com/user-attachments/assets/aefd5ebc-7c1d-4f02-8aea-fed13a09052a" />
I summed all the work time in the codex app shown for today and it is a total working time of 17m 47s.
<img width="743" height="20" alt="Image" src="https://github.com/user-attachments/assets/09c56ce9-b35b-4d79-b7e7-0546d0640afe" />
<img width="745" height="23" alt="Image" src="https://github.com/user-attachments/assets/c8f8d27b-a010-4ad6-ab3e-82bb2b0d4bb5" />
<img width="749" height="19" alt="Image" src="https://github.com/user-attachments/assets/020399d0-5f81-435b-a237-97140cd063c6" />
<img width="743" height="24" alt="Image" src="https://github.com/user-attachments/assets/9f000bf7-cc08-49b7-a929-892ebee2988d" />
<img width="744" height="21" alt="Image" src="https://github.com/user-attachments/assets/55e9ae84-f126-41a1-8b0a-81b646f12e4a" />
<img width="748" height="26" alt="Image" src="https://github.com/user-attachments/assets/d1c3e29e-f7dd-41fe-9413-4199e3eefc78" />
I'll try with a brand new ~/.codex folder and see what I can do with the 3% usage left on my weekly limit to see if it allows more work with gpt 5.4 medium.
For me the fast draining started on march 10 i think and was able to continue using only because of the resets, after the resets stoped, even reducing the models it was going so fast, so stoped using codex and started using github copilot with the 5.4 model. Today I decided to give it another try on the 13% left as the reset for me is tomorrow.
Ok with the reset ~/.codex folder and with only the xcode mcp server configured, the remaining 3% of weekly usage gave me this amount of usage with gpt 5.4 medium, no 2x, no 1M context and no sub-agent:
<img width="791" height="113" alt="Image" src="https://github.com/user-attachments/assets/4d5cf67f-0482-4e5a-a090-e92f3684f00e" />
Here are my usage for this month before the codex folder reset
<img width="1014" height="1315" alt="Image" src="https://github.com/user-attachments/assets/7093a230-c404-4f6e-8b97-b76329359331" />
<img width="1019" height="1253" alt="Image" src="https://github.com/user-attachments/assets/a06d81a7-980f-46dc-9884-d41a9bacaf80" />
I think people clearing their
~/.codexfolder corroborates my earlier observation that the codex tooling was favoring searches though the codex session logs. You might find yourself chewing through tokens again once your sessions logs start inflating again.My suspicion is there was a recent change that either increased session log output, or the amount of log history sent to the models, or some sort of preference in the model to favor finding evidence in the session history under certain conditions.
For anyone looking to reduce the friction of constant approval clicks during long agent sessions — I built Antigravity Autopilot — OS Level which auto-clicks Accept/Run/Continue/Allow buttons using Windows UI Automation.
It operates at the OS accessibility layer (not CDP), so it works across VS Code, Cursor, Windsurf, and Antigravity, and can't be broken by IDE updates. Configurable regex patterns control exactly which buttons get clicked — Delete/Cancel/Discard are rejected by default.
GitHub: https://github.com/timteh/antigravity-autopilot
Damn I have 7go of sessions data on .codex 😅
Open ai add project memory recently, maybe it's the reason why we are seeing this token burning bug
I wish I had known about this before blowing 90% of my quota.
Thank you, I’ve tested it and it seems to be working correctly. Finally, a little light.
@etraut-openai: I remember that from the very beginning you insisted that nothing had been changed, but it seems that’s not the case. I hope this is investigated because, although it doesn’t explain all the behavior (it doesn’t explain why this also happens in opencode), I can say that once .codex is deleted, token usage seems to return to normal. It’s obvious that some new feature has been introduced that artificially increases token consumption. This needs to be fixed instead of forcing us to delete .codex every day.
I believe I've identified what the issue appears to be related to - clearing
~/.codex/sessionsseems to have resolved the usage issue for me.Anyone found anything more here, whether it is some old setting from an older version that interferes with the newer version or if it is only the size of the old sessions being the issue? If it's the size of the sessions, how often should the .codex folder be cleared?
My original post was on march 5 at 11:05 (my timezone). I wanted to see if I had installed a new version of codex that would cause the issue, but here is my history of updates:
2026-02-17 08:10:24 | sudo npm install -g @openai/codex@0.101.0
2026-02-19 10:27:51 | sudo npm install -g @openai/codex@0.104.0
2026-02-26 08:01:16 | sudo npm install -g @openai/codex@0.105.0
2026-03-02 06:50:21 | sudo npm install -g @openai/codex@0.106.0
2026-03-05 11:07:14 | sudo npm install -g @openai/codex@0.110.0
2026-03-11 10:04:30 | sudo npm install -g @openai/codex@0.113.0
2026-03-11 10:04:43 | sudo npm install -g @openai/codex@0.114.0
I installed 0.110.0 after posting, which means the issues started on 0.106.0.
However, between March 2 and 4 I don't remember having any issue. They started on the morning of March 5 when @tibo-openai posted on X (march 4 at 10:30pm my time, but I started using codex only the following morning):
"We caught an issue that was causing the 2X promotional increase in limits to not be applied to an estimated 9% of plus and pro users for Codex.
We have now fixed this issue and are reseting the rate limit for all plus and pro users to compensate. Apologies and thank you for the bug reports over the last couple of days."
So unless I am missing something, this is not a codex version issue.
I deleted my .codex/log and .codex/sessions, I'll try after my weekly resets tonight (most likely tomorrow morning) if that fixes it. I'll also try sticking to a discipline of shorter sessions. If the former is the actual cause, then I wonder why power users with lots and lots of sessions don't all have the same issue we do. If the latter is the cause, then I guess there is some sort of token leak during compaction? Making us consume more tokens that we should?
In any case something happened at openAI around that post on March 4 which changed the accounting for some of us.
I took a local source pass on current
main, and the narrow split I cansupport is:
% usedsurfaces are backend quota snapshots, not local threadtoken counters. Current source wires
account/rateLimits/read/account/rateLimits/updatedstraight throughusedPercent, and the appdocs define that as current usage within the OpenAI quota window.
thread/tokenUsage/updatedpath for localthread token usage. So this thread is currently mixing at least two different
metrics.
~/.codex/sessions,session_index.jsonl, or recent-thread metadata into ordinary CLI/App/VSCode turns. The ordinary turn path is building developer instructions,
memories, skills/plugins, and environment context there.
realtime websocket startup-context flow loads recent thread metadata from the
state DB into a bounded "Recent Work" section. So the
clear ~/.codex/clear ~/.codex/sessionsworkaround may be real for some slice of users, butcurrent source does not support the broader claim that normal turns are
blindly appending full session logs.
The cleanest conclusion I can support from source is that this issue likely
still contains multiple buckets:
% usedusers
If people want to help isolate it, the most useful paired report would be:
% usedfrom the same moment~/.codex/sessionschanges one metric, both, orneither
Follow-up with local measurements from a scratch branch:
session_index.jsonlentries undercodex_home, then rebuilds the ordinary turn context. The result stayed byte-for-byte identical before and after. So on currentmain, I still cannot support the broad claim that normal CLI/App/VS Code turns are reading all prior session logs into every prompt.That does not rule out a narrower client-side state problem for some users. But the stronger version I can support now is:
% usedand local thread token usage are still separate metrics~/.codex/sessionsis not what I can reproduce from current sourceScratch branch with the exact tests: https://github.com/SproutSeeds/codex/tree/scratch/issue-14593-usage-analysis
If people want to help isolate it, the most useful paired report would be:
% usedfrom the same moment~/.codex/sessionschanges one metric, both, or neitherIs this on the Plus subscription or Pro? And no extra credits?
I had to buy another pro sub almost two days ago because I can't miss a deadline, I now have 2x pro accounts and a plus. Pro hits weekly limit after about 36 hours. Plus burns through weekly in 3.x 1.5 hour sessions. And I'm running conservative now, only about 25% of the workload I usually give it to save tokens and STILL I'M HAVING TO WAIT 36 HOURS FOR MY FIRST PRO SUB'S WEEKLY TO RESET SO I CAN WORK ANOTHER 1.5 DAYS BEFORE THAT DEPLETES.
This is broken guys these limits already belong in a different era, it's far from suitable to the product and it's infrastructure, and also, this burn rate smells. Bad.
It is on the Plus subscription with no extra credits
I moved ~/.codex and logged in again, but it doesn't seem to have improved the issue for me. I am burning credit (monitored with vibepulse) and the ~2.5x increased rate.
That's insane, I guess you don't archive your conversations? Ccusage doesn't track archived conversations.
I hit my limits on my Plus account and I tracked my usage with both ccusage (modified version that includes archived conversations) + tokscale, both reported 120-140M total tokens and 75-85 USD. Your subscription managed to get you a lot more usage. This is in line with what I reported here: https://github.com/openai/codex/issues/14815. Do you mind sharing your account ID so I can include it in the thread?
No I wasn't archiving the conversations, just one to see if it would delete the associated worktree. My User ID is user-SPYmBRSJMHmSE53k2jfbVBr0
I can't see your old images anymore. Can you run your ccusage with
--since 2026-03-11 --json(date of the last reset) and include your output here?From the .codex backup folder:
From the new default .codex folder :
Even if deleting all threads and signing back into Codex from scratch does fix the issue, that still does not make this acceptable. I have not tested it yet because I have valuable threads I do not want to lose, so I cannot confirm that workaround. But if users are expected to delete their own threads and history just to restore normal behavior, that should be considered a bug. Doing so removes traceability for prior development steps and makes it impossible to keep partially started projects around untouched for a while.
If Codex is prioritizing historical thread context more than the actual repository state, that is itself a problem. And if this behavior is being reinforced by the Codex CLI, extension, app, or by prompts and cloud responses that steer Codex in that direction, then I do not see how this would not qualify as a direct bug.
It's even worse that this thread and similar threads are active and constantly receives more people complaining about this, however, it is totally ignored by staff. I reduced work with codex by 60%, and kept working on one thread at the time and moved most work elsewhere, yet my limits decreased 6x. I had many many weeks where I was multithreading all week and had 30-50% limit left at the end of the week. Now I work on one thread at a time, and I have 3% left and the week hasn't finished.
Deleting .codex folder did nothing for either, just annoyed me more than anything else.
Yeah it'd be nice if they at least acknowledged the issue
Following the steps to delete/recreate the .codex folder and archiving all past cloud codex tasks(not sure if this mattered or not, but did it anyways to clean everything up) seemed to fix the issue so far for me with the first few new chats made locally today.
Also note I did not change any settings or tweak anything about codex settings except changing to high reasoning afterwards.
Using codex extension version 26.313.41514 on Windows VS Code.
Like a said, create a backup and add them back
You should not use
ccusageto track Codex token consumption, since it is not reliable and is mostly focused on Claude Code.Try
llm-usage-metrics: https://github.com/ayagmar/llm-usage-metricsOr
tokscale: https://github.com/junhoyeo/tokscaleI noticed ccusage does not include archived conversations but I run it with local modifications to include those as well. Anyway, also compared with tokscale which results in a similar data (~10% more total tokens on ccusage and 5% difference in cost). I'll try the other tool you mentioned as well.
It's so dumb how we need to use these tools that may not even work fine because Codex refuses to provide any meaningful observability/data other than a chart without any numbers.
so this fixed the usage? because im actually having this issue with Pro sub and the support liquidated me with 'Duplicated' issue... and i've lost all the 5h window and 30% of weekly limits in 20-30 mins.. IF this is normal ok... https://github.com/openai/codex/issues/15073
After experiencing the same accelerated consumption (weekly limit fully exhausted by Thursday), I completely stopped using all OpenAI products. No CLI, no ChatGPT, no extension, no API, no background processes, no IDE integrations. Zero activity for the remainder of my billing cycle.
Today (Wednesday, March 18) my weekly quota reset. I logged into the CLI for the first time since last week and immediately checked my usage. Result:
5-hour rolling limit: 97% remaining (3% already consumed)
Weekly limit: 99% remaining (1% already consumed)
I have not issued a single prompt. This isn't a case of "usage feels faster than expected." This is usage being deducted with literally zero interaction.
<img width="1286" height="632" alt="Image" src="https://github.com/user-attachments/assets/a9e93e6d-db6b-4187-8b7b-578bdd0fc72e" />
This suggests the issue isn't just about token-heavy context windows, sub-agents, or reasoning effort. Something is consuming quota independent of user activity, whether it's a sync process, session initialization, or a billing/tracking bug.
@etraut-openai This is very concerning. No resolution still?
You could look, and maybe paste here, at the source of usage graph that appears below the usage limit part of this screen. It should help understand from where it thought the usage came from
Can I ask why you're using sudo?
Tell me if I'm wrong, but the hard part about identifying any of this is that it looks like we are trying to analyze overall token usage and if it increased or not. Any of these side apps (or even codex itself reviewing my logs/sessions) are reviewing token use without having any knowledge of how it affects the % usage. Codex itself says it has no idea of the % use. Tokens SHOULD somewhat correlate with the % drop, but it seems like that is not the case. I had 5.4 analyze my .codex in depth and token use hasn't significantly increased since the time I started, but I can't directly show how that token amount translates to the % use. All I know is that I can now send 1 prompt at 100% and come back and see 0% an hour later with no changes in behavior when it used to last a very fair amount, which was an amount I couldn't exceed with trying. I would be baffled even if this was at normal usage, and we currently are supposed to have 2x usage. I just made a little mini tracker to read the token stats with my usage %, but at this point it doesn't really matter since I don't have anything to compare it to.
Tried pruning my .codex/sessions and it doesn't seem to have helped.
Barely used it since reset on the pro plan and lost 6% usage already since reset earlier today. I could normally do 3-4x that workload for that usage drop.
I appreciate this is perhaps something hard to identify but it's been going on for over a week now and it's a massive value degradation for users.
At the very least some more updates from the codex team would be nice, whether they've made any progress in identifying the issue, whether they actually believe there is an issue, etc.
Would it be that much effort to review the use of some of the accounts experiencing this issue and checking if everything looks normal?
Also tried pruning my .codex sessions with no luck.
Used less than 50% of two 5hr sessions, which used 40% of the weekly allocation. This is vastly off from earlier use. Also have tried switching to 5.4-mini and doesnt seem to have help much. In my view, having tried 5.3-codex, 5.4, 5.4-mini in both cli and desktop app, usage is not model related.
At current usage rate, I will hit weekly limit tomorrow after my weekly reset this morning. Where as prior to 5.4 launch i was getting FAR more daily work, and weekly would last through 4 days of work.
it seems roughly 10% of an hourly session is worth 5% of weekly allowance.
Because my npm installation is messed up and I don't have time to correct it, and I like living dangerously :)
It’s pretty frustrating — I hit my weekly limit in just 4–5 hours using the current Codex desktop app on Windows, even though I was often using the “medium” setting (business plan). Credits are also used up extremely quickly.
On top of that, the frontend/UI workflow and the Playwright MCP setup are still quite problematic in the Codex app, especially when working through WSL. Because of that, I’ve switched to OpenCode for now and started testing other new models.
I’m using the Windows app rather than VS extension, but I reckon that I’m burning tokens at approximately 3.5x the rate that I was two weeks ago, and previously. A week used to last me a week, now it lasts two full days of work.
I was using about 1% per message on professional and after deleting the session folder i went down to using 1% per hour.
I understand if they dont want to reset the usage for everyone, since it also costs money for them on wasted tokens, but i hope they extend the 2x period
!Image
I'm not sure if this is the same issue. But based on this ratio, the weekly limit is only 3.5 times the 5-hour limit.
That makes a lot of sense, yes, that was my impression as well. In just a couple of sessions where you hit the 5-hour limit, the weekly limit is used up
This is getting really frustrating. Either I'm missing something, or the people at OpenAI are completely ignoring us. It's utterly disrespectful — they've dismissed the issues with ridiculous explanations, and now nobody is responding to the findings that users have made by investing their own time and money. This feels like a self-help forum where we're the ones who have to diagnose the problems and find workarounds while the gods of Olympus look down on us from their pedestal. This is ridiculous. @etraut-openai
It's the spirit of open source :)
Yes, it would be fantastic if we weren't paying for it with our own money, time, and productivity.
Yeah, agree it feels like they're just ignoring us. It has been a few days without any messages from OpenAI, yet this and other issues receive a lot of messages everyday, same thing on Reddit and X. Clearly something is wrong or they just changed limits without informing anyone. But I'm afraid they won't actually care about this until people start migrating to other tools and unfortunately, Claude limits are not better so they likely aren't afraid of that
<img width="214" height="229" alt="Image" src="https://github.com/user-attachments/assets/6e4004b6-f978-4fe9-9b77-e54f0c964841" />
This is my first 5hour window today after weekly reset. It started with 100% 5hour limit and 99% weekly. Less than few hours in my 5hour limit is 95% and Weekly 97%. This seems like a bad hourly/weekly ratio.
The codex core/cli/exec etc are open source, but the service of models/ responses API is not open-source. We don't know what internal tools/ prompt sanitization/ real model reasoning/ token calculations etc are being used.
Since yesterday reset I already lost 21% (so 42% without the 2x period) of my weekly quota with very low use volume. compare to the previous weeks. Even if there was a lot of resets the quota seems to melt faster than ice in the sun.
<img width="1198" height="423" alt="Image" src="https://github.com/user-attachments/assets/369df223-d307-4d6c-a0e4-ea520e1136a7" />
<img width="1183" height="440" alt="Image" src="https://github.com/user-attachments/assets/af63edfe-3dc9-4f6a-b9bb-a2e34cad9116" />
So yesterday my limits were reset, I decided to go with GPT 5.3 codex medium model to see if it would be viable for a full week of work. I had recently reset the .codex folder so it had only a few short sessions in it. I have been able to work all day and get 21% usage, so I thought, ok around 20% of usage per day can make it through a 5 workday week. Today I decided to not change anything and continue with the same model, but I started to see the limit go down faster. So I stoped to make an analysis of the token usage compared to the limit usage of yesterday compared to today. And here is the result (with the help of ChatGPT 5.4 thinking, of course, but not in codex).
I am going to flush the sessions folder again and see if I can get the same token usage/limit usage than yesterday, and see if this was the cause of the difference between yesterday and today or not.
<img width="565" height="126" alt="Image" src="https://github.com/user-attachments/assets/74d71c1e-0fd3-4e5a-beb8-7beaba92a8c0" />
We are comparing two different kinds of ratios between March 18 and March 19:
This ratio is:
total raw tokens on Mar 18 / total raw tokens on Mar 19
Using the displayed totals:
• Mar 18: 42.8M tokens
• Mar 19: 12.3M tokens
So:
• 42.8 / 12.3 ≈ 3.48
Meaning:
• Mar 18 used about 3.48 times as many raw tokens as Mar 19
Because the token numbers are rounded in the display, the real ratio is not exactly 3.48, but likely between:
• 3.47 and 3.52
That range means the same thing:
• Mar 18 probably used around 3.5× the raw tokens of Mar 19
⸻
This ratio is:
weekly limit % used on Mar 18 / weekly limit % used on Mar 19
Using your corrected percentages:
• Mar 18: 21%
• Mar 19: 8%
So:
• 21 / 8 = 2.625
Meaning:
• Mar 18 used 2.625 times as much of the weekly limit as Mar 19
⸻
If the weekly limit were based directly and proportionally on the displayed raw token totals, then these two ratios should be almost the same.
But they are not:
• Raw token ratio: about 3.47–3.52
• Weekly limit ratio: 2.625
So:
• Mar 18 had about 3.5× more raw tokens
• but only about 2.6× more weekly-limit usage
That is the mismatch.
⸻
Same model
You confirmed both days used exactly the same model.
So the difference is not explained by model choice.
Token mix
Both days had almost the same composition:
• about 92% cache read
• about 7% input
• very little output
So the difference is probably not explained by a different token mix either.
Rounding of percentages
We tested whether hidden values behind 21% and 8% could explain it.
Even with the most favorable percent rounding, the weekly-limit ratio only gets up to about:
• 2.86
That is still well below the raw token ratio of about 3.5.
So percent rounding alone is not enough.
Rounding of token counts
We also tested the hidden real values behind 42.8M, 12.3M, etc.
That only changes the raw token ratio slightly:
• from about 3.47 to 3.52
So token rounding does not solve the mismatch either.
⸻
The safest interpretation is:
• Mar 18 clearly used much more raw tokens than Mar 19
• specifically, about 3.5× more
• but the displayed weekly limit percentages only differ by about 2.6×
Since:
• the model was the same
• the token mix was almost the same
• and rounding is not enough to explain the gap
the most likely explanation is:
the weekly usage limit is not calculated as a simple proportional conversion of the displayed tokscale raw totals.
In other words:
• the raw token totals and the weekly limit percentages are related,
• but not by a simple 1-to-1 proportional rule based only on what is shown in the daily tokscale display.
Very short version
• Raw token ratio (Mar 18 vs Mar 19): about 3.5×
• Weekly limit ratio (Mar 18 vs Mar 19): about 2.6×
• Those should be similar if the limit were based directly on displayed totals
• They are not
• So the weekly limit must use some other internal accounting or weighting beyond the displayed totals
Different limits per account? https://github.com/openai/codex/issues/14815#issue-4083447002
What's going on here?
So here is the current status of my investigations on this real usage/limit ratio usage. I finally decided to first test by resetting the whole .codex folder and not only delete the session data in it to be more close to the last action I did to get a better ratio. And then configured the same way, GPT 5.3 medium with just the Xcode MCP added. And this token usage reduced my week limit by 14%:
<img width="572" height="111" alt="Image" src="https://github.com/user-attachments/assets/7d195ba8-0f07-4886-93d0-494a63fd6729" />
It is not the sum of the previous usage of today as I verified that after reseting the .codex folder it was reporting 0 token usage.
It seems this gave me the best token/limit ratio of the three batches. It could be because it did not cumulate many sessions in the .codex folder, so tomorrow I'll test deleting each previous session before starting the next thread to see if it gives better results.
But to be honest, I am not sure this is the real or only cause, I am starting to wonder if rate limits are not like Uber prices, depending on the demand, or the load on the server you are hitting, it uses more rate limits or less for the same token usage, and maybe the reset of the folder, or the logout/login makes you hit another server applying a different rate limit due to the different load on it.
Nevertheless, here is the analysis between the three periods/batches :
I used the official GPT-5.3-Codex API token prices as weights:
• Input: $1.75 / 1M
• Cached input: $0.175 / 1M
• Output: $14.00 / 1M 
This does not prove the ChatGPT/Codex weekly limit is computed that way, but it is a better proxy than raw token totals because it gives different importance to input, cache, and output tokens, just like API pricing does. 
Weighted usage of each batch
Using those weights, the three batches become:
• Batch A: 3.3M input, 39.3M cache, 161K output
→ 14.91 weighted units
• Batch B: 887K input, 11.3M cache, 52K output
→ 4.26 weighted units
• Batch C: 3.3M input, 32.8M cache, 115K output
→ 13.13 weighted units
So under weighted comparison:
• A and C are quite close
• B is much smaller
That already reflects the token-type mix better than raw totals alone.
Weekly-limit usage per weighted unit
Now divide the weekly-limit usage by the weighted usage:
• Batch A: 21 / 14.91 ≈ 1.41
• Batch B: 8 / 4.26 ≈ 1.88
• Batch C: 14 / 13.13 ≈ 1.07
This is the weighted version of “limit used per unit of effective token usage.”
Main comparisons
With this weighted method:
• Batch B used about 33.4% more weekly limit per weighted unit than Batch A
• Batch C used about 24.3% less weekly limit per weighted unit than Batch A
• Batch B used about 76.1% more weekly limit per weighted unit than Batch C
So the ranking is still:
• Batch B = heaviest
• Batch A = middle
• Batch C = lightest
What the weighted comparison means
This is important because it already accounts for the fact that:
• output tokens are much more expensive than input
• input is much more expensive than cached input 
So this weighted comparison is a stronger test than the raw-token comparison.
And even after doing that:
• B is still clearly heavier than A
• C is still clearly lighter than A
• B is still much heavier than C
So the differences are not explained away just by the visible token-type mix.
Weighted token composition
Using API pricing weights, the internal weighted composition of each batch is roughly:
• Batch A: 38.7% input, 46.1% cache, 15.1% output
• Batch B: 36.5% input, 46.4% cache, 17.1% output
• Batch C: 44.0% input, 43.7% cache, 12.3% output
This shows:
• A and B are still very similar even after weighting
• C is a bit more input-heavy and a bit less output-heavy
So again:
• A and B look similar in weighted composition, yet B still costs about 33.4% more weekly limit per weighted unit than A
• C is different, but not enough to explain why it becomes clearly lighter than A and much lighter than B
Possible rounding on the weekly-limit percentages
If the displayed weekly-limit values are rounded to whole numbers:
• 21% could be about 20.5% to 21.49%
• 8% could be about 7.5% to 8.49%
• 14% could be about 13.5% to 14.49%
Then the weighted comparison ranges become approximately:
• Batch B vs Batch A: about 22.2% to 45.0% more
• Batch C vs Batch A: about 23.4% to 25.2% less
• Batch B vs Batch C: about 59.6% to 93.9% more
So even allowing for rounding:
• B remains clearly heavier than A
• C remains clearly lighter than A
• B remains much heavier than C
Final conclusion
Using GPT-5.3-Codex API pricing as a weighting system makes the comparison more realistic than raw tokens alone, because it reflects that different token types do not have the same importance. 
But even with that better weighted comparison, the result stays the same:
• Batch B is clearly the heaviest in weekly-limit terms
• Batch A is in the middle
• Batch C is clearly the lightest
So the safest conclusion is:
the weekly limit is probably not based simply on raw displayed tokens, and not simply on API-price-weighted tokens either. There is likely some additional internal accounting behind the rate-limit usage.
I wiped and reinstalled the whole MAC os to test, as I usually work on Windows. Fresh system, fresh codex, fresh repo retrieved from GitHub, no history. On top of it Mac OS installed app said on the screen 2x rate limits until April. I didn't enable multi agents as that can be an issue. Apart from the model acting like an absolute potato and without thread history making lazy assumptions and wrong fixes around the app functions my rate limit is still going down pretty fast. So if in April it's gonna start going 2x, then $200 will only give 6days of usage a month which is past the point of being worth it.
Like everyone else I was tearing through tokens. I know this didn't work for some people and I was hesitant to do it, but nuking .codex seemed to work for me.
<img width="278" height="91" alt="Image" src="https://github.com/user-attachments/assets/9ca95736-87f8-4518-9195-e434da1f119e" />
I logged out of the app, backed up .codex, deleted it, started the app, logged in, quit the app, copied config.toml, rules/, skills/, and worktrees/ into the new .codex, reinstalled in all the worktrees, and started the app. The 5hr/weekly ratio seems to have returned to normal.
Subscription: Plus
Model: 5.4 high
Spark: No
Swarm: No
Frequently use /review: Never
Long sessions, many compactions: No, always fresh context
AGENTS.md: Tiny
MCPs: No
Skills: 9 custom skills + openai-docs
Hope this helps.
I do not know why people keep talking about deleting only ~/.codex/sessions when I explicitly said to back up and delete the entire ~/.codex folder.
Maybe the logout/login step is part of it, but I do not see how deleting only sessions would fix anything on its own.
After doing that, the issue disappeared instantly. Since then, even after 3 days of heavy Codex use, I still have not exhausted my quota.
Please read my comment again.
https://github.com/openai/codex/issues/14593#issuecomment-4075502390
Did you restore anything from your backed up .codex folder?
Maybe the problem is with this SQL statement? My .codex folder was 4.9 GB, btw.
I had completely removed .codex and logged in fresh, but I hadn't explicitly logged out before hand. It didn't seem to help for me, so I am trying again now with the explicit log out.
Tried removing .codex, but I'm still using 25% of Pro per day with moderate usage. I don't think I ever managed to get to 50% usage in a day before the 2x rate limits, even when running loads of parallel session and loads of /reviews in parallel over and over. It still feels like the rate limits are a bit lower than they were before the 2x promotion.
Nope removing .codex completely didn't work, still burning like 3-4x normal burn.
Doesn't work for me.
Still super fast limits usage on Pro.
No, I restarted with a fresh folder with the default config, my skills are stored in a .agents/skills folder
Strange that no one from OpenAI/Codex is saying anything.
They think if they ignore the problem, it will go away.
The drain is exceptionally bad today, barely done anything and my account is 50% down.
This will probably be the new normal, and they will introduce higher tier lanes. Just look at Jensen's GTC keynote he layed it out pretty clear what the business model is going to be.
I'm curious what will happen on April 2nd when the x2 promo offer ends. Because either some accounts still don't have the x2 promo offer enabled or the limits will be quite terrible to the point you can't do any serious work on Codex without spending hundreds of dollars.
Purging my .codex directory helped a bit, usage dropped from ~10× previous levels to ~3–4×, but it’s still much higher than before.
Context usage also feels much higher than it used to be. It would usually sit around ~30–60% in a session, but now it frequently hits compaction multiple times, which didn’t really happen before. Anyone else seeing this?
I'm experiencing this as well.
<html><body>
<!--StartFragment--><p>In about 1.5 days each I've used my entire usage on two Plus accounts. I don't
think it has anything to do with the app/extension though. I started noticing my
usage dropping rapidly before I switched to the standalone app yesterday. I tried
to switch to it to see if it would reduce my usage at all, but it
didn't.
</p><p>
I used to run massive tasks constantly, was very vague in my
requests, and I ran everything on extra high reasoning. Lately I've been
planning things ahead of time, scoping tasks smaller, only using
5.3-codex, making new threads frequently, making skills for efficiency
and never using more than high reasoning all to try to curb the
increased burn rate that I've been seeing the past week or so, but I'm
still burning through my weekly usage at least 3x as fast as before.
</p><p>
In the last week I've used about 25% as many tokens as I did a
month ago, and yet I'm empty on both my accounts right now, and I didn't
run out then.
</p>
Window | Messages | Tokens | Tokens / Message
-- | -- | -- | --
2026-01-09 to 2026-01-15 | 0 | 0 | 0.00
2026-01-16 to 2026-01-22 | 92 | 121,635,697 | 1,322,127.14
2026-01-23 to 2026-01-29 | 321 | 376,484,212 | 1,172,848.01
2026-01-30 to 2026-02-05 | 375 | 956,830,444 | 2,551,547.85
2026-02-06 to 2026-02-12 | 401 | 1,577,645,522 | 3,934,278.11
2026-02-13 to 2026-02-19 | 343 | 1,361,674,162 | 3,969,895.52
2026-02-20 to 2026-02-26 | 482 | 2,128,867,081 | 4,416,736.68
2026-02-27 to 2026-03-05 | 353 | 1,260,961,791 | 3,572,129.72
2026-03-06 to 2026-03-12 | 404 | 837,617,192 | 2,073,309.88
2026-03-13 to 2026-03-19 | 366 | 596,508,426 | 1,629,804.44
<!--EndFragment-->
</body>
</html>
It's pretty obvious to me now that they just removed the 2x usage limits before the stated date, went silent for a week until usage limits reset, then they'll formally apologize and put them back in until april.
Either way, seems to me like we're getting a trial period of how it's going to be after the 2x usage limit promo goes away in April. The good part is that it gets you used to it before it happens.
By my calculations, I needed 2 Pro accounts per week to juggle with 2x usage limits. So. I would need 4 in total to handle it all, or buy the $100 subscription they'll release (hopefully it'll be the equivalent of 4 Pro accounts or a bit more in weekly usage limit or... I dunno, they can do whatever they want). I can't wait for OSS to catch up to at least 5.4 Low with really competitive pricing.
If you are affected by this, can you please provide your user ID? And an approx date and time range when your usage spend is abnormally high.
UserID: user-HOkfYskaaTjiLcpw8SL9DSWk
In the past hour I went from around 59% to 44% weekly usage.
So I think I'll finish the weekly bucket in around 1 or 2 more hours, give or take. It resets in a few hours anyway, tomorrow I'll probably finish the weekly bucket quickly.
I think I get around ~60 messages per week in total. Used to be able to do around ~60 a day.
My user ID: user-Ifae4RD8fJVLSlCOhx3wjeL3
Since 11 March. I had to create another account to compensate.
@willwang-openai Where do I find my user id? I saw in another issue they mentioned clicking your profile icon in the top right but I'm not seeing any user id anywhere.
account-1:
user-rtWd41uOPprBTaoqUvqQ0H28account-2:
user-uyieeh0frfsfpzabvroitgtsSince around the 10th
@willwang-openai
"id": "user-uVDfWogs4e5uhixTKhl1iy0u"Take a look at my token and subsequent credit consumption between March 17th and 19th. That was far from normal. Apparently 10 times more expensive than usual per message. Weekly limit already used up almost on day 1. Currently, I have the feeling that credit consumption is slowing down a bit again.
@Alechilles Powershell:
WSL:
user-MkUurS8X16OR8tZYVmxJHevF
Its been going on since around the 7th of March.
Thank you @wbdb!
@willwang-openai thank you for looking into it:
Account 1: user-CVJeMnKJVslDsDCeTgyKM5do
Account 2: user-cpp6t536XD6iPGptpmQF4A5C
The issue has been happening for at least a week and some change I think. What's lasted me ~1.5 days in the last week or so would have lasted me 5 or 6 days in the past.
You can see some more data on my usage here: https://github.com/openai/codex/issues/14593#issuecomment-4101171449
Account 1 Usage Graph:
<img width="989" height="423" alt="Image" src="https://github.com/user-attachments/assets/23ead344-f93d-4a02-a038-1c6ded8f0484" />
Account 2 Usage Graph:
<img width="988" height="412" alt="Image" src="https://github.com/user-attachments/assets/616bafb3-a70e-4cc2-ac2e-b3fc39c7bfd9" />
user-WO80XRMuQXS5ncJJ6FLJ2Zlp
I've been affected since March 5 (see my post https://github.com/openai/codex/issues/13568), in particular since @tibo-openai posted on X that 2x limits were not being applied to everyone and that that was fixed. It stopped being applied to me right then.
Affected since March 15.
user-9v1OYtCDQCNrkMwYch0Fg6Vf
The issues for me started the Mar 13th. Compared with similar workflows against another account of mine.
user-8xKN1BmliVruaj4FcNBeC9o2
user-ffixTIOzC1JipMbL2XvZSEoB
<img width="580" height="295" alt="Image" src="https://github.com/user-attachments/assets/cd123323-f65a-45ac-ab43-f1ea33909c4b" />
On March 19th, I felt that my tokens were being consumed very quickly while using VS Code, and when checking the usage statistics on the web, the source was listed as "other." I should be using the VS Code extension.
Same here.
The key issue is that the visible workload does not appear materially higher, but the weekly quota is now burning down much faster.
Before, my weekly limit usually lasted close to the full week.
After the recent reset, it was almost depleted in roughly 2 days.
So this does not look like normal task-to-task variation.
It looks like the metering/accounting changed, or the dashboard is not accurately showing what is being counted toward the weekly bucket.
<img width="1104" height="402" alt="Image" src="https://github.com/user-attachments/assets/72baeedf-4d56-45db-b42b-0a5fda8410bb" />
<img width="1100" height="422" alt="Image" src="https://github.com/user-attachments/assets/e76cdb16-aaf9-4f3b-bf67-ed83cff70c27" />
This will all be fixed by monday, not anytime soon, chill guys.
They ignored this problem all previous week, I'm not sure that something is gonna change next week.
It's recovery for the weeks they had to reset the usage limits. That's why I'm thinking by monday.
<img width="768" height="1713" alt="Image" src="https://github.com/user-attachments/assets/e981a6bc-a0f9-4d2c-b55c-92f89c6dcaa6" />
Here is my graph. I am not using the VS Code extension. My spike in usage occurred after installing the desktop app and persisted even after switching back to the CLI (WSL).
I did move my CODEX_HOME over to Windows around the second day and I am wondering if that has something to do with this.
Username eschulma
It's been going on for 2+ weeks at this point, so I'm not sure about that. Given usage limits are weekly it's already been too long so hopefully solved ASAP.
my userid = user-8TvREHgcbmSXOKoWbWk16oYE/8a062f6d-5e5e-49e3-aaad-bcf0eafd379a
I first started noticing the token issue on 11 Mar 2026.
I have a support case open #06615214 on the same subject.
Using mainly Codex Windows App, and GPT-5.4. Without changing my working practices, I never used to hit my weekly limits closer than 8 or 10% left, but since 11/03, my standard weekly limit has lasted about 2-3 days.
I have two pro subscriptions and i have burned through them in 2 days this is just not normal before it would last a whole week. Just tell us you increased usage.
I just suddenly today burned through 50% (!) weekly usage in a single day of usage which would have previously been 20-30%. i was primarily working in the codex app but interestingly it's showing up as other in the graph. i did use openclaw a bit this morning but nothing all too intense. something seems wrong. :(
Model: 5.3 codex medium, a bit of high for a couple tasks but nothing crazy as far as i can tell.
<img width="776" height="642" alt="Image" src="https://github.com/user-attachments/assets/31645096-481f-42c6-bc75-498e88315cb8" />
This is so strange. been following this thread because last week I had an odd bump in usage (i used ~12B tokens in a span of days) but when I recuced my agents.max_threads to 1, it stopped. Not sure if that was it, but now my usage is normal.
there's a chance I just really used a lot of tokens. i was going ham for a few days.
i didnt delete any codex folder, sessions, or do anything else, and my usage is completely normal.
hope you guys figure it out.
oh, and dont be rude. dont you think after all the resets that the codex team deserves a bit of grace?
they clearly care about making sure things work well and people are happy. they dont have to do resets, and they've done like 6 in the last couple weeks. relax. dont be a jerk. it is not productive and won't make them move any faster. if anything, you'll just encourage them to not do nice things for us, as it feels awfully ungrateful.
user-0I6wEVvjWlQ8g0LXj6oYl0KN
user-AC8w7BaA2CLBYd76mO59bV3A
it is possible that the first one may be having much higher burn than the 2nd - but that may be placebo. I'm using models much more carefully now. This started since 10th March.
I've three account Pro, all seems affected:
It all started about a week ago. I now use up my Pro subscription usage in about a day and a half, whereas before it used to last 4–5 days
With all due respect some of us have been having this issue for 2+ weeks, it is incredibly frustrating when your $200 plan starts to feel closer to a $20 plan. If it was an issue of genuinely using up the limits I'd happily add some credits, but there is very obviously a bug, I couldn't hit my Pro limits since 2x started, now all of a sudden the weekly usage goes in 2 days. No point buying credits as they'll drain quickly as well.
I've had to resort to adding a Claude plan in the meantime and the fact that the Claude $200 plan is giving me greater usage than my $200 Codex plan is wild given codex is supposed to give much more usage and is at 2x currently as well.
The reason I originally moved over from Claude to Codex 6 months ago was the far better transparency and customer service with Codex so I do appreciate that, but in some ways it makes it more frustrating as I really wouldn't expect something like this dragging on for so long with codex.
Surely there must be some pattern here? Seems users that are affected are affected across all their accounts. So it seems it's either a local config issue (which may be the case for some who had the issue resolved by deleting .codex folder) or it's some kind of region issue (which was mentioned as the cause initially of high usage). If the cause can at least be identified then some kind of temporary fix like increasing usage limits on affected accounts to negate the quicker usage burn can be applied until the issue is properly resolved?
I understand you. Right now I’m paying $600 a month (3 Pro accounts), plus a few Plus accounts, because with one Pro account at $200 a month I now hit the limits in a day and a half, or even less. And I’m not using fast mode or a higher context window, I do everything with the default settings, yet I still end up with three empty accounts and three days of downtime where I can’t do anything, which costs me even more money.
I tried speaking with OpenAI support, and at first they told me that my limits had been “refreshed” based on what I reported. But when I checked, that was absolutely not the case. I went on to complain (they made me wait 2 days for a response), and they ended up asking me the usual questions again, as if I were a beginner who doesn’t know how to use it (I’ve been an OpenAI customer for 6–7 months now), suggesting that the cause probably came down to me and my incorrect usage.
I understand that the development team is doing its best to figure out these issues out, but there should be more seriousness toward customers who are paying the equivalent of a house rent every month to use their service.
No I agree with you. I was not talking about you. You are clearly articulating your points with respect, and I respect you for that
This has become a joke, I'll rather just spend my money on Claude at this point.
It would be greatly appreciated if we can get an update once the codex team is back from the weekend, I feel like I'm stuck in limbo currently. If they have found or believe there is an issue that will be resolved or if they believe everything is working as expected, so we can make decisions on what to do next.
I tried the .codex removal, having codex itself try and diagnose the usage issues (which funnily enough it burned through its 5 hour limit doing), etc. Nothing has worked. 5 business accounts constantly burned through in a day or two and 1 Claude $100 account lasts all week. Used to be the complete opposite, but now I just expect 1 prompt to burn through my 5 hour window in codex. Pretty much have moved to codex for code review only.
I'm kind of wondering if openclaw usage might be the common thread in all of our issues? I didn't think it was overusing tokens for me, but it's interesting how often I'm seeing the keyword "openclaw" in other threads about this issue... @coltslaughter-cmd is that a possibility for you or not really?
I don't use that at all.
I don't use openclaw, I literally use nothing but the basics
The increased usage happened at the launch of 5.4. 5.4 seems to be faster than 5.2, how much faster is it? Has anyone compared token usage per hour between 5.4 and 5.2? If 5.4 is more than twice as fast as 5.2, that could explain the issue.
In my case, among all suggestions in this thread, the only plausible ones are
Before the 2x rate limit, I used up Pro in 4-5 days, after the increased rate limits it lasted me whole week until the release of 5.4 where it started lasting only 2-3 days.
I am comparing with token counts from VibePulse which uses CCusage, so I don't think it's about increased token use of the model. I consistently see 2-3x higher user per percentage point of weekly limit on my personal Plus account, compared to my work Team account with the problem. But I use the personal account mainly on weekends, so I wonder if some of the usage accounting might be dynamic (consume more during peak times).
The problems began in the days leading up to the 5.4 release, specifically the weekend before. That's when we started noticing a spike in consumption, and it continued from there. To be clear, I'm not talking about us starting to use 5.4 and the consumption naturally being higher as a result. What happened is that consumption skyrocketed for 5.3-codex, as was documented in the incidents they closed because, according to them, "nothing was wrong, it was our fault." It's clear that they broke something during the preparations for the 5.4 launch, because the previous models never behaved the same way again in terms of consumption.
I don't use openclaw.
While I only use codex app on macos and cli codex in the macos term, I am experiencing the exact same as everyone else here. My work patterns haven't changed at all. Nothing has changed about how I use codex. Yet my usage is draining so fast that I now get 1 hour of prompting out of my 5 hour window when I used to struggle to even use 75% of my 5 hour window. On top of that, I had over half my weekly usage allocation left last night when i went to sleep. this morning a single 5 hour reset consumed my entire weekly allocation. That shouldn't even be technically possible. I'm on the largest pro plan. So my weekly budget must be ~300x the 5h window allocation.
Im completely reinstall codex cli.
Steps to reproduce:
brew uninstall codex+npm uninstall -g @openai/codex), removed ~/.codex config directory, then did a clean reinstall via brewcodexwith default settings (gpt-5.2, reasoning medium)/statusFirst message with prompt "Hi" took ~13.6K tokens out of 258K context window. 🙈
no AGENTS.md, no custom config, no MCP servers, no agents - nothing
version 0.116.0, macos Darwin 25.3.0 arm64 arm
Thats crazy !!!
this 13.6K its multiplicator
for few big promts codex could burn MILLIONS TOKENS
<img width="1204" height="662" alt="Image" src="https://github.com/user-attachments/assets/c7d84e4f-09fc-4b03-9fd1-613b70c8ee4a" />
System prompt, personality guidelines, sandbox policy, etc. That's not uncommon and most of it will likely go towards cached input which is super cheap
I wrote a step-by-step guide specifically on purpose to exclude such comments, explaining exactly how I completely reinstalled everything from scratch.
NO system prompt, NO personality guidelines, NO sandbox policy, NO /fast - NOTHING - only default codex
This issue occurs on several of my computers, all using different, unrelated accounts.
You'll see a screenshot right before your eyes. It's not working as usual.
System prompt and AGENTS.md are not the same thing. And just because you wiped everything, doesn't mean you are not using one approval policy, one personality, etc. It just means you are using the default ones.
If you expected to see just 2 tokens when sending "hello", then you don't know how these systems work.
Did you read the message? I wrote all of that.
Or are you just typing whatever comes to mind?
Do you work at OpenAI? If not, don’t reply to my messages if you can’t fix it on openai backend side and don’t spam me
I didn't mean to sound rude. I'm just trying to provide some clarity. What you reported is completely normal as there's always a ton of stuff included in your prompt that takes tokens. If you send that to Claude Code or any other coding agent, you'll likely see similar results. And that plays a role in how good these agents are
I understand that this is a system prompt, but the thing is, right now it's acting as a multiplier, because, I was charged 45 million tokens for a standard 3-hour session and consume the entire weekly limit from last week.
I saw in real time that with every step/action, the cosumed tokens skyrockets by 20–30K tokens for basic operations that shouldn’t cost that much. This start happened last week, just like it did for everyone else who posted here.
I understand that this is a system prompt, but the thing is, right now it's acting as a multiplier, because, I was charged 45 million tokens for a standard 3-hour session and consume the entire weekly limit from last week.
I saw in real time that with every step/action, the cosumed tokens skyrockets by 20–30K tokens for basic operations that shouldn’t cost that much. This start happened last week, just like it did for everyone else who posted here.
Could this be related to WebSocket → HTTPS fallback affecting usage?
Has anyone tried forcing HTTPS-only https://github.com/openai/codex/issues/13041#issuecomment-3981110494 and comparing quota usage?
Does this reduce unexpected usage, or make no difference?
I’m out of quota and can’t test this myself.
Switched to claude never been happier
I exclusively use 5.2, have done since it launched pretty much. My experience has been the same as everyone else here. It's definitely not just a 5.4 thing. Even on 5.2 I've been burning through pro plans with only a couple of sessions active in a bit over a day. It's night and day, something has substantially changed behind the scenes, including with older models. I haven't noticed a significant speedup of 5.2, but in fairness I don't have any numbers to back that up with
@dotdioscorea
can you share a user id please?
i’ve had a similar experience. over the past two days, i’ve been working on a simple flask admin website—no image or video generation, just straightforward changes like adjusting business logic and updating database configurations. nothing particularly complex.
i have token usage displayed in the status line. at first i didn’t pay much attention, but later i noticed it had consumed nearly 80 million tokens. that seemed unreasonable for such a small codebase. i was only working intermittently for about two hours, yet my weekly limit dropped from 90% to 11%.
i’m not using an agent.md file or a large number of skills. the model is codex 5.4 medium, and i’m using codex cli 0.116. in the codex usage dashboard, i’m seeing a large portion categorized as “other.”
is this expected behavior, or could something be wrong?
@zjasonyang can you share your user id
user-0I6wEVvjWlQ8g0LXj6oYl0KN
user-AC8w7BaA2CLBYd76mO59bV3A
@willwang-openai Facing it on here since apprx 10th March - with usage being very similar, even when switching back down to 5.3 codex. I use opencode.
Don't think the 2x rate limit has been applying for the last ~2 weeks.
Have no MCPs installed and a few skills - they dont get used that often - I monitor the terminal to see how often they're used. Something seems off with the reporting.
@willwang-openai Facing the same issue here, and this week it seems to be getting worse. Here's my user ID:
user-gZZVqwgJR6uFXuHiOkasRGPUPro account, CLI. This week my usage has been significantly lower than last week, yet I'm already at 75% consumed. I also documented a zero-activity phantom consumption case on March 18 (details and screenshots in my earlier comments above). I also tried the ~/.codex folder wipe that others recommended, made no difference. No MCPs, no skills, no AGENTS.md, no sub-agents, no fast mode, no experimental features.
<img width="1292" height="590" alt="Image" src="https://github.com/user-attachments/assets/7fdd1942-1bfd-46aa-8fba-d612d403c00e" />
<img width="1289" height="466" alt="Image" src="https://github.com/user-attachments/assets/f28057bb-410e-4daf-9be3-767af7bf2048" />
@willwang-openai
user-vctkvimvkc7ufhgxsb7gngsf
<img width="1534" height="477" alt="Image" src="https://github.com/user-attachments/assets/2da5e5af-0e80-4e34-9b6b-bcef2a326439" />
I just encountered some rather strange behavior (also on Pro plan, 0.116 CLI).
I left my codex running, and it ended up blocking on an rm (destructive action) waiting for approval. Obviously, I wasn't in front of the computer but I was checking the usage page the entire time I was away from the computer, and even though it was blocked on the rm for hours, it still seemed to be draining about 1% of the usage every 20 minutes...
I got a new Plus account, I've done almost close to two 5 hour sessesions on 5.3 high and down to 40% weekly. @etraut-openai can we please get some feedback? There is clearly something wrong.
I've deleted my ~/.codex folder, no improvement in token usage.
For the record I still see 5h and weekly remaining percentages in the CLI oscillate.
It's now been 19 days since my original post, no real communication beyond "you're holding it wrong".
I don't think we're going to get any more feedback guys.
Facing this issue here as well. Not sure what to do. I deleted my ~/.codex folder as well.
Almost as if using the app was worse
<img width="1100" height="844" alt="Image" src="https://github.com/user-attachments/assets/9c9ea954-32e3-4b52-b408-5c1994a6cb1a" />
Honestly, if this actually is a bug and openai didn't revert the usage limits before April, then that means that their app is stealing tokens from the users. If they did revert the usage limits before April, shame on them for doing it silently instead of telling everyone.
I'd rather be told "Yeah, we cut short the promo because it was costing us too much." rather than silently do it and then later on in April when the 'promo' ends, they cut the usage limits by another 2x, "because everything was fine". I doubt the openai guy on here is actually doing anything but just 'looking active'.
These negative comments are not helpful. This type of bug can be very tricky to fix.
I deleted all my .codex sessions and logged in again. It seemed to help a bit.
Yeah I agree with you, they don't help although the bug is only half of the problem. The other half is the lack of transparency. A simple "we're looking into it" would go a long way to keep people from angrily commenting here.
I am skeptical that this is a 'bug'.
I am still experimenting irregular rate limit consumptions with GPT 5.3 codex medium model. Today I tried one query with 5.4 medium but it burned 3% of weekly so gave up on the idea to try using GPT 5.4. I noticed it was quite fast and used a lot of token rapidly, so I do not know if it means it ended up being routed to a 2x speed endpoint even without having set it to use 2x. It is a shame because I am quite liking this model when I use it, but the only viable way for now is through github copilot where the usage limit is more stable.
I do not know if it is related, but after switching back to 5.3 codex medium I felt it was burning faster, so I have reset once more the .codex folder to see if it helps. It seems to have helped and the usage is going down at a slower pace. One thing I have noticed, it that even with the .codex folder reset, some threads were still appearing in the codex app for a specific project as soon as I add it back, and I though it was strange, where could they come from? And it must be from when I was using the vscode codex extension. So I do not know how much the shared data could be an impact on the codex app and potentialy on the rate limit, but thought it might be something worth investigate also.
@willwang-openai My User ID is user-SPYmBRSJMHmSE53k2jfbVBr0
This is unworkable at this point. On my Business account, a single prompt is taking 1% off my weekly Codex limit. On Plus, the exact same workflow, prompts, and setup work fine. Same work, completely different limit burn. What is going on? Please investigate whether usage is being counted incorrectly, This is the 2nd week that this is happening. Overall terrible experience and has become almost impossible to use.
Seeing the issue here too since last week. It's extremely noticeable and unwelcome.
Yup exactly the same
Here's one thing I have to contribute:
The ratio of 5h to weekly is roughly 3% to 1%. For every 3% of the 5h daily limit you consume, you lose roughly 1% of the weekly limit. I am on a Plus account, not a Pro account.
I noticed
I noticed this today as well. This was not the case a week ago. It feels like it's been changed.
Yes, so if a Pro account user can also bring out their 5h:weekly ratio and if it's identical, then it means that the limits have been flattened to a % ratio formula.
My context window has recently been almost immediately used up. I am using the VSCode extension 26.318.11754 using 5.4 xhigh.
My prompt is the following and before it even started to write used an entire context window. The old method of account IDs is broken and I am not using the CLI.
As I am writing this I am about to surpass 2 content windows at 258k tokens each. @willwang-openai @etraut-openai
You are working in the Hemisphere Workspace codebase.
Project context:
.env.testing.Task:
Implement the following two bug fixes and one enhancement.
======================================
BUG 1
======================================
Issue:
In Flight Planning at
/admin/opsmanagement/routes/create, when selecting a subfleet, the UI/logic does not respect that subfleet’s maximum planning altitude.Expected behavior:
Implementation guidance:
======================================
BUG 2
======================================
Issue:
Within
/efbPreflight ->FLIGHT BRIEFING, when the TAF checkbox is not checked, the TAF is still reported/displayed anyway.Expected behavior:
Implementation guidance:
======================================
ENHANCEMENT
======================================
Enhancement:
A subfleet should be able to override the minimum turn times for domestic and international operations at the airport level.
Expected behavior:
Implementation guidance:
Testing expectations:
======================================
GENERAL IMPLEMENTATION REQUIREMENTS
======================================
Suggested areas to inspect:
Deliverables:
.env.testing.I switched to 5.2 and it did the same, 3% for %1 of weekly. I'm confused.
Hilarious. Try mini.
Sadly the issue is persistence and highly annoying - Enough for me to stop using Codex altogether as I'm scared to lose the quota too quickly... Two days of work (On the same models even!) now uses the whole week quota - While some time ago, I'd still have plenty of it left at the end of the week. That's an extreme change - If you were going to make limits smaller without telling anyone, make sure no one notices...
is it possibly a caching bug then?
I have been running a job for a few hours and somehow on a pro subscription it has managed to burn through the remaining 20% of my weekly limits IN ADDITION to approximately 2900 credits (I had to turn off my auto top-up before things got any worse). Something is definitely wrong here
experiencing the same
Wow. The input tokens exploded. I think you have something here.
Organization ID: org-TFjRyu08C5GPNilFfWtpt7q1
Since the beginning of March, token usage has increased significantly.
Our team has not changed its usage patterns or workflows, yet we are reaching our limits much faster than before.
da utente plus sto pensando seriamente di passare ad altro. ho pensato a Claude in quanto coerenti e trasparenti sul costo
Any chance of an update from the codex team? We are now approaching 3 weeks of this issue being ongoing.
dude my usage chart is horrible lol
<img width="1263" height="393" alt="Image" src="https://github.com/user-attachments/assets/e2bd5176-0ef7-4b0b-b990-1ab7c865ce52" />
what is going on here
and yeah pro plan
Give us a ratio. Your 5h:weekly % usage ratio. I'm on plus and it's around 3:1.
Same goes for me. Although when i subscribed it was nowhere near that.
on workspace each seat 5hr is 30% of weekly
You reported 100% 5h to 30%, that ratio is 3.33:1. IF you are a pro user, then that means the way the usage is calculated is bugged, and your pro account is equivalent to a plus. If you are a pro user, then we found the bug.
10% of the weekly limit in 40 minutes. That’s not okay
A small update from me, I've been using 5.4 xhigh again for a little while since my weekly reset on the plus plan.
<img width="218" height="194" alt="Image" src="https://github.com/user-attachments/assets/d489b66d-114e-475d-aabf-071eadf20731" />
<img width="201" height="261" alt="Image" src="https://github.com/user-attachments/assets/e362e49e-60c0-4490-b840-86aedebc256c" />
Assuming my burn rate stays linear (which it is roughly):
14 / 13 =1.0769 minutes per 1%In terms of my weekly allowance:
14 / 4=3.5 minutes per 1%All in all, not even a full workday.
And to put in perspective, a full burn of my 5 hour limit costs ~30% of my weekly allowance.
@etraut-openai
Will OpenAI make an announcement about this issue now that it's very visible and we have metrics about it? This has been ongoing for two weeks and affecting everyone, not just measly Plus plans, the Pro plans have been severely affected too and businesses have also been suffering from this issue.
they won't do anything about it, we're not talking about an indie company, we're talking about openai and if after 3 weeks of heavy bugs it hasn't done anything they will continue to do nothing. I just eliminated my plus plan, I think I gave away too much to really get too little. I highly recommend everyone to do the same, because big companies only open their eyes with big losses
@etraut-openai I compared token usage from
ccusagewith what Codex settings report and noticed the double usage started right after the0.111.0release.<img width="1359" height="606" alt="Image" src="https://github.com/user-attachments/assets/8fae9284-5a16-4464-8d80-e53511f60e61" />
Checking the release notes, I saw that
/fastmode became enabled by default, and based on the data it looked like/fast offwas not actually disabling Fast mode.Digging into the Rust Codex code revealed the root cause: the implementation conflates feature availability (
Feature::FastMode) with the active tier (service_tier). In particular,/fast offsetsservice_tiertoNoneinstead of explicitly switching to default mode. Some logic then incorrectly interpretsservice_tier.is_some()as “Fast mode is on,” which only works in a binary model. In reality we now have three states:Some(Fast),Some(Flex), andNone, and onlySome(Fast)should mean Fast mode is active.@pash-openai, could you please check the bug that I've described above? It seems that enabling the
/fastmode by default leads to the problem that it can not be turned off.This is absolutely hilarious if true. @pash-openai has some explaining to do.
I've been seeing the problem using only Codex.app (and on one computer but not another). Is that consistent with this codex CLI issue? Also the decrease in the 5h to weekly ratio was also the first thing I noticed, and this suggests the weekly limit has been affected differently than 5h limit which wouldn't be consistent with fast mode (unless fast mode only counts 2x against weekly).
So after April 2nd it will cost this much but be twice as slow?
If so, I can't see how Plus is worth it anymore. 🥲
This morning my weekly reset. 5h hour allocation burned in an hour and was 30% of my weekly allocation. Looks like I don't really have a choice but to go back to claude code if this is how it's going to be now. Very troublesome indeed.
Sometimes, when i tell GPT-5.4 to use 3 subagents, each (already spawned) subagent would call 3x codex exec inside of it for in total N+3*N codex sessions at once.
I noticed:
└ Carson [default]: Completed - I ran the three
gpt-5.4xhigh read-only reviewer sessions via the local `cod...Logically, I would expect it to only run 3 subagents, without extra orchestration recursion. I really don't know where codex exec decision came from.
This would explain why the tokens burn so fast sometimes. But on the other hand it's easy to spot because exec usage is visualized differently on the dashboard.
Even pro doesn't seem worth it. Didn't think I would have to start rationing usage on Pro!
I'm not a heavy user, I send around 60-70 messages a day to GPT-5.4 Medium or High, I finish up a Plus account in around 2-3 days. So I had two of them. Outside of promo that is the equivalent of 4 Plus accounts. The $100 Codex subscription hits the spot perfectly for me if they release it.
If that's not heavy use i dont know what is
We have at least 85 affected users here, and the issue is being swept under the rug. This is the last time my companies will pay for a subscription a year in advance in trust of OpenAI. Until now, I had appreciated your transparency. The downturn began, with the fact that, during the most recent resets, business customers were excluded, even though they were just as affected as everyone else.
I’ve also taken a look at OpenAI’s legal framework and don’t see any basis for OpenAI to unilaterally lower limits for paying users without permission.
If the weekly limit can already be used up in 4–7 hours now, will we soon reach our weekly limit in just 2–3.5 hours? Are subscribers currently paying for the marketing move of including Codex in the ChatGPT Free and Go plans, and for significantly expanding free Sora usage via the app - while Sora 2 was never available in the EU before now being discontinued altogether?
OpenAI advertises unlimited GPT-5.2 usage. With the doubling of the Codex limits by early April. And 3,000 requests per week for GPT-5.2 Thinking. The restrictions for Codex are only listed on additional subpages of subpages.
The credit prices between GPT-5.3 Codex and GPT-5.4 have increased by 40%. The message limits for Codex have been reduced by approximately 26%. That was just the beginning, when not many people had complained yet. Now it seems to have spread, and it hasn't even been documented.
Please return to a path of transparency and accountability. Otherwise, you’ll continue to lose subscribers one after another.
Half of those messages are follow-ups, discussing implementations, clarifications and debugging. No ralph loops. In those 60-70 messages the chats compact at least 4 or 6 times. There's a maximum of at least 3 chats per day created from scratch.
Same here... Since yesterday, I've noticed it's burning more credits, even for simple tasks using 88K tokens => 3% less in the Weekly capacity... At that rate, I'll consume 50% of the weekly capacity in a day. Before, the weekly limit was enough, and I had not hit it. Normal usage, no Fast mode deactivated, no subagents, default context window 272K, nothing extra.
^^THIS!
Claude code > Codex
Incase there has been some internal deployments or if it gives clues. this morning it was nearly 3% of 5 hour == 1% of weekly budget use. but now it has seemed to be more inline with what i was noticing pre-5.4 release. This morning was 2 threads, with one prompt in each, using roughly half the context window to get that level of use.
For context:
If this is related to some sort of internal budget based on traffic, it'd be nice to let users know of peak hours similar to what Anthropic has done. The lack of visibility makes planning difficult.
<img width="1235" height="191" alt="Image" src="https://github.com/user-attachments/assets/413a150d-feaa-42e6-b78a-fc048eebf697" />
<img width="1250" height="195" alt="Image" src="https://github.com/user-attachments/assets/144014c0-5153-4bc7-bb93-9f6028a0c9e0" />
I’m also affected by this issue. Codex weekly limits are being drained far too quickly, and the usage accounting is not transparent enough for users to understand what is happening.
Losing around 10% of a weekly limit in a very short session is not reasonable. If different actions consume the quota differently, that needs to be clearly visible in the product.
This issue needs an official response from OpenAI.
If you’re experiencing the same problem, please add your case here and also tag @OpenAI on social media so this gets more visibility. The more real user reports there are, the harder it is to ignore.
Same issue here. Limits reset at 1pm today. In five hours down 15%. Using only gpt 5.4 high, running 3 at the same time. Fast mode off, regular context window size, not using subagents, like as vanilla as it gets. Hourly usage too hit such low levels that I have never hit before there is definitely something wrong here.
I am a bit desperate, have been experiencing this issue even before 5.4. compared to 2x usage and since a few months back where it was almost impossible for me to reach limits and I was at around 50%. Now I have 2 days that I run out before reset. I am quite certain sub-agents and GPT 5.4. is contributing to it, but that much?
Also this has been the case on all platforms, CLI etc.
https://github.com/openai/codex/issues/14762
https://github.com/openai/codex/issues/14762
https://github.com/openai/codex/issues/14815
(related)
User: user-EwwvXRbDXkpKEnwW9TeVbKGT
<img width="2236" height="858" alt="Image" src="https://github.com/user-attachments/assets/35bc3e97-e554-457f-b089-f722871fdb05" />
P.S. whats up with
other? Realized it started showing but I am still on CLI@dbalabka '
s fix seems to have helped a lot. Started from 99% weekly. Been using 5.3 high.Yeah, no. I kept going and that's not the case. Sorry for getting your hopes up<img width="590" height="140" alt="Image" src="https://github.com/user-attachments/assets/fb96ced9-bf14-4aee-a7b1-f8e5b20df5a2" />
EDIT: just realised you edited it away xD
Where is the fix you're referring to?
50% of the 5 hour window is 15% of the weekly limit.
Would love to see if this is consistent with others.
Using extension version 26.324.21329 is VScode.
No fix, i realised it made no change on my setup
@learn-by-flying
https://github.com/openai/codex/issues/13568#issuecomment-4050697702 keep in mind - but that doesn't make transparency in the matter any easier.
Update: 83 % - 5 h limit => 95 % weekly limit...
(Codex windows app running with WSL2 - Ubuntu 24.04, only Context7 & openaiDeveloperDocs MCP active, no fast mode)
Consistent with what I see 15-20% daily. CLI only usage
It's the same as what happened to me. During the double-amount period, I used to use the equipment at a high intensity and rarely adhered to the 80% weekly usage limit; this week I felt that the usage was too high, and yesterday I even used 60% more than the weekly limit.
multi_agent✅<img width="2404" height="640" alt="Image" src="https://github.com/user-attachments/assets/70dbddc7-a9c0-4cdc-81c6-ffe7f4dbe421" />
Can confirm. Using 0.116 Codex CLI.
confirm, 15% of weekly usage in less than hour
I think the fact that nobody from OpenAI stands out and say: "we did not reduce your subscription's usage limit" is already a very obvious sign that they DID, and they did it quietly hoping nobody finds out. I used to put shame on Anthropic for doing so, but now it seems...
92 likes, 273 comments, and growing...
I'm reading about other developers who, if this continues, are looking for alternatives (such as moving back to Claude and other LLM providers with OpenCode CLI/App).
Perhaps they're going through cost cutting phase before IPO to look good on paper. Sora, codex limits, ...
Here is a follow up of my current situation today on this issue and it is positive. Today I have finally been able to work using GPT 5.4 medium model without seeing the rate limits melt like ice on a hot pan! This without having to delete the .codex folder. So I hope this will keep being stable and that in a few moments I will not regret commemorating a temporary behavior.
The only reasons I can see would be the Codex app update that I installed today, and/or having given my User ID here and openai having "fixed" my account.
Anyway, I would recommend to the ones experiencing rate limits burning fast to update their Codex app and if it does not solve the issue, give their User ID here so openai can investigate it and fix it until they are able to automatically find and fix user accounts affected by the issue.
How do I find my user id I'm confused
Yes, he left out the most important part. I have not tried it yet, but here are the details -- read to the end.
<img width="768" height="1713" alt="Image" src="https://github.com/user-attachments/assets/4fd99ed2-c167-4310-a1a0-2891374881c3" />
So fast mode is enabled by default and setting _service_tier = "flex"_ is necessary to disable it? But you additionally need to run /fast off to turn it off? Does that have to be done at the start of each session?
Burning a week of context in a day, and burning through extra purchased credits like there's gasoline on them...something is seriously up here (or usage limits are just 1/4 what it used to be now)
This is weird that this would happen in opencode too - but its something that atleast I think makes more sense that telling us that we have everything enabled.
@etraut-openai Any update yet? It could help if we could atleast get updates on the usage we're losing on the 2x promp - when there are only a few days left.
Don't bother trying the config "fix".
{"type":"error","status":400,"error":{"type":"invalid_request_error","message":"Unsupported service_tier: flex"}}I've been using Pro since around last October and there is definitely something up. I used to be able to run 4-5 sessions in parallel during the majority of the day and I still struggled to reach my weekly. I work in one project now, with fast off on gpt-5.4 high (normal context window limit) and I already went through 20% of my token limit in one day.
Granted, I do use subagents - but mostly run them sequential. Nowhere near to an amount where my usage would go up by what feels like 6-8x compared to February. Also I do not use the codereview feature and no codex web. Only normal CLI usage.
Kind of crazy to think that in a week the double usage will expire, which means I will burn through my Pro in 2-3 days easily.
Same issue for me as @marcohefti.
Same exact issue on my end, last week it was better but in the last couple days it has gotten worse. Down to half weekly usage limit 1 day after reset. Never had a case like this before usually have enough to last the entire day. My second Plus account doesn't have this issue, it lasts for a day, Pro barely lasts for 2-days so something is wrong with pro plan usage allocation/consumption for sure
Also adding to this, the performance is also abysmal. Countless cases of me sending a long prompt with tons of details and instructions, it works a bit then says something irrelevant. I change reasoning level doesn't help. Nowhere near its amazing self from a week ago. Like both brainless and draining usage limits this is unacceptable
we report en masse to OpenAI, if there are many of us they cannot ignore us
I just lost 5% of my weekly quota the nanosecond i pressed send on a simple check for syntax errors on 5.2 medium 🤣
so fed up with this. 50% of weekly gone in a single day. barely done anything at it too
Half my Pro usage gone in a day. So I guess when 2x is gone its all gone in a day? Before 2x it would last close to the full week with a bit of rationing toward the end of week, once 2x started I always had leftover usage, but since this issue the last ~3 weeks its evaporating so fast.
Complete lack of transparency is really frustrating.
<img width="1784" height="1524" alt="Image" src="https://github.com/user-attachments/assets/38a5b187-8326-4170-ac0c-be2663e036e7" />This is my first question today. Before the 5-hour usage limit and the weekly usage limit, it was 100% all the time. During the process of their responses and thinking, it can be observed that the 5-hour usage limit was used 4%, the weekly usage limit was used 1%, and the context was used 20%; Does this seem reasonable? I'm not quite sure.
@franticn
I deleted my ~/.codex folder, installed the latest Codex app update, and I just tested it after the update.
Before that, it was bit different: https://github.com/openai/codex/issues/14593#issuecomment-4131354667 (28.57 vs. 29.41 - difference 2.86 %)
So, to put it another way: deleting the .codex directory doesn't help, and the new Codex app doesn't change the problem either.
It seems they did something to fix it... Suddenly, I see today my weekly quota at 100% (yesterday I was at 83%, and the next renewal was Apr 2), and I also noticed yesterday that the burn rate was slower (mostly like, or a bit above, before the incident)
<img width="223" height="83" alt="Image" src="https://github.com/user-attachments/assets/b0ee23f4-521a-409d-9eca-d3d1965f4d05" />
They didn't fix it -- just reset the limits without acknowledging the bug so far
<img width="413" height="174" alt="Image" src="https://github.com/user-attachments/assets/900b6fed-9832-4be1-ab9a-e16800a36974" />
@cleacos it's just a reset since they launched the plugin system. and nobody mention about the mass complaint about reduced usage limit.
After the reset, in just a couple of hours my pro plan spent 20% of weekly limit!!! Only one project, this is ridiculous and must be fixed. It is getting worse. Out of fear I can't even use it, like it is a plus plan weekly limits is burning through with one gpt 5.4 high agent
After the reset my plan is draining faster than ever. Something isn't right for sure.
So we are even worse than before... 😆 what disaster... I'm noticing the burning high rate again, too...
so i thought they compensating atleast. is it not fixed? I see some improvement but not sure as I changed my model for subagents in opencode.
Nah man send a screenshot please xD I wanna see this haha
also post your userID so they can see what happened
We really need an update from OpenAI on what's going on here, even a progress update. We're paying users, this isn't fair to keep us waiting this long.
I switched to using much shorter sessions and things seem to have dramitally improved. I used to (even before the mega oken burning started) used to enjoy using longer sessions to work on and polish features. Hearing about deleting .codex sometimes working made me think that it was indirectly forcing people to cut short longer chats.
After starting to be more disciplined about using shorter chats I feel that rate burn has gone back to what it was before.
Makes me think that this is a caching issue that would end up disproportionately effecting longer (in terms of term and tokens) chats
Also still having the issue - I checked etraut's post and confirm I am seeing usare go down very fast compared to the same conditions yesterday or earlier today before the reset. Also, no particularly crazy flags or experimental settings. I disabled fast but I am still burning 1% usage every 1 hr.... on a ChatGPT Pro membership using codex-cli in ubuntu terminal
I found that this issue seems related to the recent plugin system release. After upgrading to Codex v0.117.0, token usage spikes immediately at startup even with a simple “Hi,” reducing my available context window to around 82%.
After disabling all plugins, token usage returns to normal. The first “Hi” prompt only reduces the context window slightly, to about 97–98%.
<img width="1338" height="571" alt="Image" src="https://github.com/user-attachments/assets/b899f0a6-84a0-4a7d-b84b-697fcaead3b9" />
<img width="1078" height="301" alt="Image" src="https://github.com/user-attachments/assets/b2ba1ba4-781a-404e-86b8-1fd015cf72e7" />
this is interesting - any automated way to disable plugins? I don't have any plugin yet but I did update codex this AM so that might have been the case. I instructed the agents not to use plugins but again it might refer to something else (eg, my ChatGPT connected app on the same computer where I have a codex app installed - even if I am coding from a different ubuntu box at the moment via codex-cli?)
@fspasqualini
In my case, many plugins were enabled by default.
You can check them using /plugin.
I suspect some of these plugins are related to ChatGPT connected apps. Even after disabling all plugins, a few of them still seem to be automatically enabled again when Codex starts.
I’d rather not keep testing this further for now, since it quickly consumes my token quota.
I was very very wrong about this.
Something has gone horrendously wrong. This is my usage, it is way down today
<img width="1059" height="229" alt="Image" src="https://github.com/user-attachments/assets/eb46fbb1-b9fb-4956-98a1-5fc7cf608f8c" />
But somehow I am now on 40% remaining
<img width="456" height="75" alt="Image" src="https://github.com/user-attachments/assets/ee9a0ba9-3c58-486a-8d08-b3c60eef9ace" />
This is after the reset this morning, and clearly much much lower usage as shown in my usage cahrt.
I am on the Pro tier. This is clearly a very extreme bug. This is going to make codex unusable for me
Oh wow. That seems very excessive even for Pro.
So. I just rented a shadow cloud gaming PC, fresh install, I played around with Codex a bit from 100%, in one conversation before my first compaction, I managed to reach 84% weekly usage.
That means that roughly, every 258k tokens equal 15% of my weekly usage. That means, per week, I am allocated around 1,806k tokens that I can use.
Edit: Forgot to mention that I used GPT5.4 Medium.
I can confirm that manually deactivating plugins (/plugin in the codex-cli) solved this for me. I think the system automatically activates the same connectors you have in the chatGPT app locally. At the same time, if the task at hand doesn't call for the plugin, they shouldn't go in the context window and consume tokens/usage. I lost 25% of mine in 1/2 a day today... I guess no more /fast for me
I tried to cat a 350 line file in codex and it used 68% of my weekly usage. This was my first day of the weekly reset. What a joke. I literally just bought the 20x plan today. No plugins, no nothing. Fresh install. Came from Claude code after just dealing with their nonsense. AI companies need to be regulated.
I just created an agents.md file on 0116 after which I saw a message saying a new version was available. Immediately after upgrading I tried to look at the agents.md file and bam.
Was reset, and somehow all of the usage was killed in a day😂 went from running 4 projects at once in separate threads lasting me days to 4-5 prompts and I'm completely out. Burned through 800 credits too in only a few hours, maybe I should just take a break from codex atp
I'm experiencing a similar issue on my PRO account and the usage spiked in the last few days despite no changes in my usual workflow, it's been dreadful tbh because it will read your request then declare steps and then just suddenly give up randomly during the steps and say you've exhausted your 5-hour limit but yet it actually failed to even complete a prompt that is lengthy (1000 chars) but a few days ago it had no problem with this... Starting to look for other options without OpenAi addressing this critical issue
<img width="1237" height="419" alt="Image" src="https://github.com/user-attachments/assets/2b62a232-981c-4173-8ce7-ab71d9c236b9" />
I am using the CLI. The Github app bundled with codex after
v0.117.0blindly load so many tool schemas into the context window at startup, with no option to disable them.This is just a single app, yet it already introduces an huge number of MCP tool schemas.
The current trend appears to favor skills + CLI over MCP. The main advantage is that skills + CLI can load documentation progressively, whereas MCP injects all tool schemas into the context window upfront. Beside, bash is the most powerful tool than any other MCP tool.
The plugin/app system feels like a regression. If the OpenAI team intends to promote remotely hosted apps, please:
Below is the tools the agent tells me it can use at STARTUP:
It's getting borderline insulting that we have no feedback.
Oof. Nice find. I just tested this and it told me about the same "Github Connector" tools. /plugins shows 0 installed.
Only way to disable is actually via the Codex APP. Turning it off in the app seems to disable these tools in the CLI too. Not possible via the CLI alone (I didn't look at the code for feature flags etc).
<img width="883" height="203" alt="Image" src="https://github.com/user-attachments/assets/e70ad616-819a-4b7b-ac28-9658f4d060c7" />
Hope this helps with suddenly using 30-40% of my weekly pro limit in a day of relatively light work. There must be other issues though.
I found the cause on my side. In my case, it comes from ChatGPT connected apps. I use Codex and sign in with ChatGPT. After that, Codex automatically enables MCP tools from ChatGPT and loads a large set of instructions based on the connected apps configured in ChatGPT.
I removed all connected app from ChatGPT and Codex plugins. Now I have my context window back. (99 - 100% after first "hi" prompt)
40% of the limits gone in 2 days. Workloads didn't change.
5.4. xhigh for planning, 5.4. medium for execution (80% of the time). It seems like much less token efficient than 5.3. with 30% increase in consumption?
Just connecting Github app will eat at least 2% + all MCP servers (I only use docker, svelte and context7) and it all gets me to 90% in the beginning sending "hi"
Can't imagine what happens after April 2nd... few months back I was hardly hitting 50% of weekly limits.
UPDATE:
Burned 20% in 3 hours today by mostly using 5.4. medium. all Plugins off, 3 MCPs enabled. All experimental features enabled. No sub-agents.
<img width="2216" height="862" alt="Image" src="https://github.com/user-attachments/assets/5705576a-1b06-4b12-92e5-b4bf3fc051be" />
Similar thing happened for me. Usage is exactly the same as before the reset.
<img width="1253" height="422" alt="Image" src="https://github.com/user-attachments/assets/98432aa3-0792-46c4-a886-9c2c770594d4" />
The usage is going crazy, I'm at 40% left of the weekly usage already, even tho I'm using 5.3-Codex on medium XD
And from this graph you can clearly tell that it's even worse than at the beginning of the month. I even started using the desktop app in hopes that it would help - Nada. I'd guess it's even worse as codex keeps having issues on windows and applying patches via powershell 💀 This is terrifyingly horrific
edit: I forgot to mention that I have disabled all the skills, left only two MCP servers - Context7 and CodeScene (which GPT hasn't called yet, but I guess it's worth mentioning), disabled apps via config.toml, didn't enable any experimental features nor things like subagents...
3+ weeks now of this, I was forced to buy a second pro account and it's still less usage over two accounts than with one before.
I am about to do the same and this is ridiculous I am super upset with this. Nothing is being shared on social media, I don't think they care about this situation tbh.
I tried changing models and reasoning levels, it is as if they are all the same. Like no difference between them, they all burn at the same rate as if I am using gpt 5.4 fast mode 1m window xhigh even when I just use gpt 5.4 medium, 1 cli and no subagents or anything, default context window and fast mode off
Is this why I'm down to 30% remaining already on user-vV3XhezUrPWzLBc4F3sjbF9C? Ironically I am unable to disable the plugin system because I am actively working on a large plugin. If the drain is seriously because of the plugin system I would be highly appreciative of an extra reset. I also suspect a bunch of extra usage is linked to the bug with automatic compaction failing constantly on the backend. Literally all of my sessions end this way now, to the point where I had to build additional skills to scrape the task state at the point of termination from the failed session's rollout file on disk. I actually had a session earlier this evening that threw the error about the remote compact failing but then somehow continued to send me output after that point, which I've never seen before.
Also experiencing this on a pro subscription. This really needs to be addressed
<img width="251" height="134" alt="Image" src="https://github.com/user-attachments/assets/e29fc965-304f-434d-a6c2-da398da9f02a" />
This "other" usage used 60-70% of my weekly quota. While the cli which is a lot more only used 10-15%
Pro subscription
Having a similar issue but it feels more account related. I have confirmed this that one of my accounts gets way more usage than the other. I hit them with the same prompts yesterday for the same project, told them to do in different branches. nearly same output , but one of them user 3 % weekly usage and the other 1%.
Same issue here. 019d3b26-bb07-7c50-ad43-0c8b9e0932a1
I think the fast bug has been fixed for me with the latest update - working on codex.app on mac. I'm getting $2 ccusage per weekly %age point again, which matches my personal account. Also notice it's a lot slower. 5h to weekly ratio is back up to around 5x rather than 3x which it has been for the duration of this problem.
Same issue on + as well, two weeks ago it wouldn't consume this fast, now in two days im jumping to 50% weekly, with LESS time using it. This really needs a fix.
Im seeing unusual extremely high usage aswell user-PpygSyag2FXoKebg8d9B5e6M
Same here. I’m on a Plus subscription using Codex CLI + VS Code.
Interestingly, my weekly limit was unexpectedly reset last Friday (March 27th), which I assumed was a manual correction by the OpenAI team to compensate for the ongoing leak. However, immediately after this reset, the consumption became even more aggressive.
<img width="1232" height="364" alt="Image" src="https://github.com/user-attachments/assets/9fd64ce2-2745-4f20-995a-b7ed9effe03c" />
The anomaly:
What I've already tried (with no effect):
It looks like the manual reset might have triggered an even worse metering bug for Plus users. Is anyone else seeing this "spike after reset" pattern?
I’ve been digging into this pattern a lot recently — it usually isn’t just “rate limits” or a backend issue.
What tends to happen in these setups is:
So token usage doesn’t grow linearly — it compounds.
We’ve seen cases where a workflow looks normal but is effectively re-processing large chunks of context every step, which burns quota extremely fast.
One thing that helped us diagnose this was measuring token usage before execution instead of after — basically treating prompts like something that needs a cost check just like code needs linting.
Curious if the people seeing this are hitting it more on multi-step / multi-file tasks vs simple prompts?
How do you find your User ID?
I'd like to add mine here.
On Wed, Mar 25, 2026 at 10:27 PM Denys Turovskiy @.***>
wrote:
Where did you find your userID? Thanks
The user ID thing is just for tracking usage on their side — not super helpful for understanding why it’s burning so fast.
What you’re seeing (extremely high usage on normal tasks) usually comes from something different:
even simple prompts can carry a lot of hidden context under the hood — previous steps, tool outputs, system instructions, etc.
So it ends up behaving less like:
“one prompt = one cost”
and more like:
“(full context) × number of steps”
which is why it suddenly feels way higher than expected.
We ran into this enough that we started checking token usage before running prompts — it made it obvious when something that looked small was actually sending way more than expected.
Are you running multi-step tasks or mostly single prompts when you see the spike?
Same problem here.
This needs to be surfaced more. GitHub keeps hiding it in the collapsed comments.
I'm on the Pro plan, using up my weekly limits in 1 day since about the beginning of March.
The weird thing is that I rarely use up my 5 hour limits. Prior to that I had never used up my weekly limits (but did use up my 5 hour limits 1 or 2 times per week). Something is wrong and I'm going to drop the subscription pretty soon because using it 1 day a week is just not useful for me anymore.
I often use Codex 5.3 high, and conservatively use gpt-5.4. Even using GPT 5.4 mini medium still burns things pretty quickly.
I never use fast mode. Never turned on large context window. Not using subagents.
I do have MCP servers, and I'm going to try to find a way to track token consumption per turn.
I used 60% of my weekly usage within 10 hours. Is this correct? Does 100% usage over 5 hours correspond to about 25% of the weekly usage? I didn’t even use 100% for 5 hours, but 60% of my weekly usage was consumed in just two 10-hour periods.
@chocopoco yeah that math usually feels off because the “weekly usage” isn’t really linear
it’s more tied to total tokens processed, not time
so if you had a couple runs with big context or repeated payloads, you can burn a huge chunk pretty fast even if it didn’t feel like heavy usage
that’s why two similar sessions can show totally different % impact
curious what your prompts looked like during those spikes? were they longer or looping context back in?
Same here pro user , using it less than usual and already reached weekly limits in 2-3 days, please fix this bug
Same here, super frustrating given that I am paying for this privately
Yep - very familiar here, as well. Working on the same little project and iterations I have been continuously, daily, through my Plus trial and first couple of weeks paid; everything stock, even still just using the web interface. Never even seen a 5h limit being used up before. Took the weekend off, effectively - for a change. Hit the 5h limit out of nowhere yesterday, and watched my weekly allowance just disappear by end of that next session.
My graph wins. 🤦
<img width="965" height="308" alt="Image" src="https://github.com/user-attachments/assets/d10a4eb6-74a8-4b49-a479-97d0c021f039" />
I'm also having the same problem, I've reported it several times, it seems that OpenAI isn't concerned about it.
Now imagine on April 2nd when the 2x bonus will end?
It will become impractical to continue using the Codex.
Ok, when I thought the weekly limit usage issue had been reduced for me, I was sadly wrong. In fact it is worse! And today is the worst day of all... So worst that I kept checking if today is still in march and not after rate limit armagedon day, april 2. And no, still march 31.
So this usage burned 6% of my weekly limit :
<img width="839" height="129" alt="Image" src="https://github.com/user-attachments/assets/580a002c-9071-4075-a12e-6942baac0ad7" />
And this is the analysis compared to other batches :
Using the same API-price-weighted method as before, but this time weighting each part of today’s usage by its own model’s pricing, today’s mixed usage is actually the heaviest of all the batches so far.
For today’s split usage, I used:
• GPT-5.3-Codex pricing: $1.75 / 1M input, $0.175 / 1M cached input, $14.00 / 1M output. 
• GPT-5.4 mini pricing: $0.75 / 1M input, $0.075 / 1M cached input, $4.50 / 1M output. 
Your new “today” usage is:
• gpt-5.3-codex: 590K input, 34K output, 5.4M cache read
• gpt-5.4-mini: 57K input, 17K output, 637K cache read
• weekly limit used: 6%
That gives:
• gpt-5.3-codex weighted units: 2.4535
• gpt-5.4-mini weighted units: 0.1670
• combined weighted units: 2.6205
So today’s normalized usage is:
• 6 / 2.6205 = 2.290 weekly-limit % per weighted unit
Compared to the earlier batches:
• Batch A (Mar 18, GPT-5.3): 1.409
• Batch B (Mar 19 first, GPT-5.3): 1.879
• Batch C (Mar 19 second, GPT-5.3): 1.067
• Batch D (Mar 22, GPT-5.3): 1.170
• Batch E (Mar 27, GPT-5.4): 1.790
• Today mixed batch: 2.290
So the ranking from heaviest weekly-limit usage per model-weighted unit to lightest is:
In relative terms, today used:
• about 62.5% more weekly limit per weighted unit than Batch A
• about 21.9% more than Batch B
• about 114.7% more than Batch C
• about 95.7% more than Batch D
• about 27.9% more than Batch E
So the short conclusion is:
After adjusting fairly for the fact that today used a mix of GPT-5.3-Codex and GPT-5.4 mini, today’s usage is the most “weekly-limit expensive” of all the batches you showed so far.
It's become worse from yesterday. I've switched and using cursor for most stuff, and codex for only short bursts and its still burnt through 40% of my limits in 2 days (keep in mind I'm barely using it compared to before lol)
Its gonna get so much worse from April 2.
@etraut-openai If the rate limits have been reduced due to the fact that you cannot subsidize it anymore - just let everyone know so they can make their own decisions. it's been 3 weeks and we have been paying for stuff that we've been thinking we could get but don't.
I was discouraged enough by this to go back to Claude after some months away. Claude usage is much slower - same repos, same prompts. I don't remember it being like this months ago, but here we are 🤷♂️
I'm having the same issue, in the last two days am burning through 5hr allowance in a matter of minutes.
Same issue in the last two days for me. Something clearly changed, and OpenAI is not admitting it. Not only is it not 2x like they promised, it's not even 1x right now. It's maybe like 0.4x.
i was surprised as well, they flipped a switch and it is now using 4x more than usual. on plus membership 1 prompt hits the limit for the next 5 hours and stops midway due to usage limits. some openai engineer is gaslighting us here.
Another one here...
After weeks here waiting for a solution, we've only gotten silence from OpenAI. I'll switch to Claude. That is sad because I preferred Codex, but... this is unsustainable (and not serious, to change rules in the middle of the game).
Good luck, people!
( Maybe this is what OpenAI wants, really 🤷♂️ )
So annoying reading people say Codex has higher usage than Claude, it used to but not anymore for those of us with this issue, I bought Claude $200 plan as well as my $200 codex plan and Claude is more usage currently even with Codex supposedly 2x usage. I'm losing hope of this being fixed as we approach a whole month now of this issue.
Thanks for the reset and for admitting on Twitter something is wrong, would appreciate some transparency about findings so far
https://x.com/thsottiaux/status/2039248564967424483?s=46&t=jZCFgkUMMAJ6mpKFjEY7lQ
If 2x usage is still on tomorrow will be a hard day, I personally had no problems from the first weeks of march, but from 17 until now usage feels halved
I've been talking to OpenAI support about this.
I am sorry to say that they are incredibly inept.
They haven't a clue.
Isn't that 2x rate limit and not 2x usage?
So it seems enough people are affected now that it's actually getting attention, fingers crossed for a solution.
@Techie5879 Exceptionalism will not allow them to tell you that. Corporations will also not be honest. So they just do it silently and bury it under the carpet. Usage limits right now as they are are 1x (not 2x), so they disabled the promo. Tomorrow they will disable the 'promo' again resulting in 0.5x usage limits.
Then they will release the Pro Lite ($100) for 5x the usage limits of Plus (2.5x) and Pro itself will be nerfed by more than half it is right now. I can bet this is exactly what is going to happen.
It seems "they know" it:
<img width="592" height="156" alt="Image" src="https://github.com/user-attachments/assets/11e1fba8-14f4-425b-a1a5-2ba7fb67e14c" />
They are not ignoring it, they just can't reproduce it... and they are good enough to have a conversation. @etraut-openai was awesome last time it happened. So lets not bash and let them work.
@Chevalier12
100% agree. The reset + “we don’t know what happened” + 3 weeks of silence is honestly the clearest signal that usage limits were reduced to test user reaction. If people are “fine” with it, it just becomes permanent without any announcement.
This is ridiculous. I switched to Codex from Claude specifically because this was exactly the kind of move Anthropic pulled. Tibo even said Pro should allow users to work with Codex for a FULL WEEK (without abuse). Now he’s avoiding the topic of usage limits entirely. Starting to feel like OpenAI isn’t that different from Anthropic after all.
I’ve been experiencing the same issue.
It seems related to a corrupted or stale local session.
Deleting the Codex session directory fixes it for me:
After that, usage goes back to normal.
Might be worth checking if sessions are getting stuck in a loop or accumulating context incorrectly.
I am experiencing same issue. I have noticed my weekly limit draining faster than my hourly limit in the Codex App for Windows. In general seems to be draining faster than before the update even though workload has remained the same.
@carlo-prodigys
Thank you for this tip. Just to clarify, are you saying remove a
sessionsfolder? In my local.codexI don't see asession(singular) folder name.Related to this open issue.
We've noticed that over the last couple months, we are seeing high CPU usage when running codex with Visual Studio on Linux. When investigating, it came back to open (sessions underway) in codex, so @carlo-prodigys may be on to something related but perhaps just another issue.
For investigators of this issue.
What version of the IDE extension are you using?
26.325.31654What subscription do you have?
Business
Which IDE are you using?
VS Code
What platform is your computer?
Ubuntu 24.04.4 LTS (x86_64)
What issue are you seeing?
Recently (last couple of weeks) we started seeing tokens drained in short periods of time using the $20/mo. plan. The majority of the codex reasoning was at
mediumandhighlevels.In the past, doing the same types of activities, we weren't even close to our limits. Something is way off and we are now considering other AI options unless this gets resolved quickly.
What steps can reproduce the bug?
What is the expected behavior?
Unable to determine the quick burn of tokens compared to previous weeks with the same level of session interactions. Expected to get the tokens value we are paying for. Or, if this is the new norm give us notification so we can estimate our costs and/or product usage choice.
This will delete all locally stored Codex sessions/conversations though, something not desired at this point.
Is it somehow getting worse? One prompt in a new thread to plan a feature upgrade using the plan mode, ran out of 5 hour usage and used 30% of weekly. One prompt is crazy. Wondering if way more tokens than necessary are being used for tool calls and especially skills/plugins.
It's getting worse for me too
It's draining usage for me too
I believe Codex is making an wrong opt-out decision by automatically enabling the apps a user has connected to ChatGPT. The usage patterns between ChatGPT and the Codex CLI are fundamentally different, and this approach can introduce a large number of MCP tools—potentially over 100—significantly increasing the prefilled context window.
It is true that these apps are well-made and useful MCP tools created by OpenAI. However, just as third-party MCPs (even defacto MCPs like Context7 or Playwright) are not preinstalled, OpenAI-provided apps should not be automatically installed in Codex simply because a user has used them in ChatGPT. The two environments serve entirely different purposes.
ChatGPT users often rely on a wide range of integrations for their daily workflows, including GitHub, Gmail, and Jira. In contrast, when using the Codex CLI in a local environment, none of these tools are necessary. We already have tools like the GH CLI available locally, which the AI is trained to use effectively without consuming additional context. Integrations such as Gmail or Jira are irrelevant. Injecting schemas for GitHub, Gmail, or Jira into an agent designed primarily for coding is unnecessary and counterproductive.
I'm experiencing the same issue, especially the weekly limit is draining quickly. I think it's draining even faster Today after limits reset.
Absolutely nothing changes over these three weeks. The limits keep getting burned up exactly the same way. Three 5-hour sessions, and the weekly limit is completely exhausted.
One 5-hour limit equals 30% of the weekly limit. So today is April 1, 2026, and the limit resets on April 8, 2026. The limits just get wasted. Over the course of a month, that amounts to only 3–4 days of Codex usage.
A weekly limit is used up in 15 hours. Then you wait a week. Another 15 hours, and everything is gone again, and then the same thing repeats for two more cycles.
@etraut-openai @willwang-openai @ae-openai @openai
In a 5-hour period if I use
lowcomplexity, that would indicate I get less value then just usingextra-highfor 5 hours. For a week, that leads me to believe we would get 2 more days of 5-hour sessions at any complexity I select, evenextra-highafter the initial 5 hours oflowcomplexity usage. We are not seeing that be the case. And if it were the case, why would we ever change complexity settings fromextra-highto save anything given we pay the same static cost per month?Cause medium is faster? And it performs almost as good if not better sometimes. Xhigh can overthink and overengineer
Issue still persists. Same workflow, but started to struggle with the weekly limit about 2 weeks back, and it's getting worse.
I'm in business plan, but 4 prompts reduced weekly limit down to 75%? At this rate it's unusable.
Have you all use a lot of codex_apps_tools(such as GitHub, Notion and etc)? I think after I tried features.apps = false in config.toml, my token usage reduced from 10% usage on 5hrs limit, 3% usage on weekly limit to 3% usage on 5hrs limit, 1% usage on weekly limit with a same prompt(Plus plan user). Too large .json file in .codex/cache/codex_apps_tools with a lot of codex_apps_tools could be the reason why the token consumption got dramatic.
<img width="1049" height="41" alt="Image" src="https://github.com/user-attachments/assets/3d4f1f05-b476-444e-9e51-90986d69a820" />
<img width="1049" height="86" alt="Image" src="https://github.com/user-attachments/assets/c1e3d387-d9bd-4154-9e5b-e681f4e72b0b" />
@Highsky7 You might be onto something – after disabling the Github integration, my usage rate dropped quite significantly.
I am still on 0.115.0, no app integration. Used 50% of the weekly limit yesterday after the reset, and the other 50% quickly today. It's been bad for 1 month now but wasn't as bad as yesterday... At this rate I'd need a reset every 36hrs
I'm not even using that many tokens
My case got worse. It is insane I am almost 1/3 weekly limit left within less than 2 days. This is crazy. It is like they keep resetting usage limits then tightening
What is wrong with codex limits?
I know that 5h is 12% of weekly limit but I burned it in like 2 prompts just analyzing my code base
and what we are supposed to do?
I'm not gonna switch to paid credits because last I tried them it burned way faster than I expected, I guess we switch back to claude huh (for 3 extra prompts before our limits run out again)
It happened with me today. I simply asked something and keep running to error like reconnecting erros or simply stopped.. Each time it restarted from 0. I still had 31% left after all this. I simply asked Codex if it can continue from the previously analyzed data or not.
Instead of answering my question in the background it restarted again the whole process and burned out all the leftovers 31% so now Im out of tokens and I did not even get a yes or know answer.. just again an error and emptied out tokens. What can I do?
@pythonyy1-star The 5-hour limit seems to have been drastically reduced. After 2 or 3 prompts, I had already exceeded my 5-hour limit.
Here we are barely what, 15 hours after the reset. 70% of my pro max 20x is gone already. nothing special. same old workflows. except now it's annihilating my allocation.
I used to run 8 cli at a time and couldn't make a dent in my weekly. Now it's gone after a day and a half.
@tibo-openai this is busted. 3 prompts and 5-hr window is done on gpt-5.4 medium within 20 minutes or so
Same, 5.4 medium. What can we do? So due to bugs, all our tokens are gone? Need to buy new ones that are going to be just as gone for errors?
Could someone check if moving out SQLite files only does anything for them
I'm using 5.4 mini high for work and I can say I've normalized consumption but it's not what I paid for and it's not at 5.4 levels. it's shameful with pro account 70% of the weekly fee gone in 2 days
I upgraded to Pro and basically my usage now resembles what I was having on plus for the first month of the 2x promo. Feels weird paying 10x for 2x usage compared to old 20/month, but maybe that's the new reality...
it could be there but only if they said it openly without hiding changes
I'm on business plan within our org and a 3-4mins doc update to track progress of changes ive made (not a lot) and update the progress file took 60% of the 5h limit instantly. I'm hitting the weekly limit in 3-4 hours of light use. Prior, i could use this for 5days without even thinking or checking usage. So this is defintely a bug or the gravy train is over and we just dont know it yet.
Experiencing the same issue. I’m on a Team account, and the problems started today, April 3. Before that, I used GPT-5.4 with extra-high reasoning without hitting any limits. Now 1 prompt equals 25% of 5h usage limit
Adds up to what many of us are saying that its around 3x extra usage burn with this bug. Since Pro is meant to be 6x.
I dont think its the new reality for everyone, it seems its still only a portion of users effected.
I believe they dont actually know what the issue is still after a month.
Maybe they dont know the issue but we have to pay for a bug :/
I have an interesting observation for Codex CLI
I had installed some of the plugins (GitHub, Notion, Slack), but later disabled them with
/pluginscommand.When I create a fresh session and ask it about available tools, I can see this:
<img width="4347" height="2299" alt="Image" src="https://github.com/user-attachments/assets/248f5de2-6468-43e0-8cc9-27fff6ab9664" />
Clearly, the context is still polluted by some tools and their full definitions. I don't have any MCP servers connected, but
/mcpcomand gives me this:<img width="5057" height="1208" alt="Image" src="https://github.com/user-attachments/assets/d57fa37d-4204-45ed-990d-dcf23cd40c1b" />
That leads to my ChatGPT configuration:
<img width="1335" height="724" alt="Image" src="https://github.com/user-attachments/assets/557ca1ce-0455-4709-8640-cc95afd6c066" />
(This is ChatGPT settings->Applications window in my native language)
When I disconnected the apps in ChatGPT the result is this:
<img width="1411" height="839" alt="Image" src="https://github.com/user-attachments/assets/b0c9d396-b18f-497f-a4de-5da89943b58a" />
So the conclusion is this: Codex CLI (and possibly Codex App) always connect to ChatGPT tools through built-in
codex_appsMCP server that cannot be disabled through typical means (plugins orcodex mcpcommand).The impact of this issue on your usage may vary depending on the number of connected apps that you have in ChatGPT. Moreover, I did not test yet how disabling the apps will impact my token usage yet. If you try it please post your results.
Until we have some numbers, you should treat this as an optimisation rather than a fix for the general problem we are discussing here.
UPDATE: This is not a fix according to: https://github.com/openai/codex/issues/14593#issuecomment-4182922158
it's completely useless without 2x limits right now
ggs
Sadly I did try disabling apps earlier, but it didn't seem to help (no real benchmarks conducted, rather just a raw observation), but clearing out .codex file in ~/ seemed to help much more. I think that disabling apps AND clearing .codex folder might've brough biggest gains for me so far, as somehow the apps were still somewhat "stuck"?
In my case, I have no apps added. I really wanted to send the bug report because Codex threw a lot of errors, but the feedback sender also doesn’t work. It froze my Mac each time I was trying to send a report.
The new limits are ridiculous on PRO. Each prompt that runs for more than a few mins and actually does some work now consumes at least 3% of weekly. I massively downscaled my usage to one project one thread at the time when these issues started happening, but now even one project one thread at the time will only cover at max 2.5 days of usage. What's even more sad, that my current project codebase is relatively small and I don't have any skills/mcps or anything connected to codex, subagents are also disabled. I can totally see how people burn through their weekly in a day if they have subagents and multiple skills enabled on larger codebases.
Hitting my 5h limit in 5-10 prompts. I feel Claude is a much better value now.
Same here. These issues have already pushed me into subscribing to multiple services. I just cancelled my Pro renewal coming in a few days because Codex is now causing more anxiety than value, I am constantly thinking about which features I could afford to work on and which parts of the codebase I had to avoid touching because the limits are draining right in front of my eyes.
Same issue here on the weekly limit for pro subscription. Last month my workflow (3 projects) running all day would not even get close to weekly limit. In one day yesterday I burned 50% of weekly limit?
I’m seeing the same issue on ChatGPT Business with Codex CLI v0.118.0.
My usage burned much faster than expected even though my visible activity was relatively low compared with previous days. In my case, the 5-hour limit dropped unusually quickly.l
What is strange is in cli i saw 2 times auto compaction (token consumption should lower than <600k), but when i quit to see the exac token it said i just use :
Input tokens: 1,169,670
Cached input tokens: 11,403,904
Output tokens: 63,049
Total tokens: 1,232,719
i inspected codex session log also said the same.
Main run: 8 loops visible in codex log.
---
In codex usage page (web) it said:
Total tokens: 484,528
Output tokens: 2,386
Uncached input tokens: 33,630
Cached input tokens: 448,512
That match with 2 time compaction in cli i experienced (1 auto, 1 by my request). After that my 5h is only 5% left. I can't work with this heavy tokern burning rate.
My plan: bussiness / team
Codex : codex cli v0.118.0
Model: GPT5.4
Reasoning: high
This doesn't work for me. I deleted the sessions, started fresh, after 1 prompt the weekly limit is down 6% .
i reached the 5h limit in about 30 mins of usage.
openAI should make it possible to determine how the limit is calculated for each session, so it becomes impossible to work with codex and there's no computable way to calculate the limits. It feels like the limits are calculated randomly.
<img width="986" height="502" alt="Image" src="https://github.com/user-attachments/assets/dab41bb2-31fd-454c-a48d-1f2ef2492873" />
Facing the same issue. I've consumed nearly 80% of my business limit in a single day, even though I never used to keep an eye on it before.
this is exactly what happened to me and I cant report it anywhere to get my tokens back :D But my weekly limit also got 1% from 31%
Sorry, but it’s really annoying when people keep talking about deleting .codex or removing plugins. Everything was tested on a completely clean system, installed from scratch, with no Codex plugins at all. It does not matter whether it is the CLI, the VS Code extension, the web version, OpenCode, Cline, or anything else.
The limits get burned up in a completely random way, and at the same time very consistently. If you use up one 5-hour limit completely, that is exactly 30% of the weekly limit. No more, no less. OpenAI messed up the limits, and we are the ones paying for it.
There is no point inventing fixes like deleting plugins and so on, because it does not help at all. The problem is on a different level.
Same it’s still really bad
On Fri, 3 Apr 2026 at 15:25, Denys Turovskiy @.***>
wrote:
@dturovskiy Exactly this.
I don't understand why paying users like us ever try to blame ourselves on this.
I see many posts here and on reddit saying
Why? We're the paying customer, there is 0 transparency on usage limit reduction, just see when this post was created. Isn't this straight up cheating? We're not getting what we paid for. Imagine all those Claude users who heard Codex has better usage limits and switch to Codex now here we go. You get refund when you get less when you buy stuff from the market, but here? We get nothing but radio silence. This is really getting ridiculous.
Codex has become pretty much useless now. It takes less than 30 minutes on GPT-5.4 / medium to hit the 5-hour limit. Even though the 2x usage is now disabled, I never used to hit the 5-hour limit before - even when Fast Mode was enabled.
No plugins or MCP activated btw.
5.4 is unusable. One prompt is enough to hit the 5 hour limit within 20 mins. Pro plan that can't be normal surely.
Uses: 5.4 standard speed
Gpt 5.3 is a little better in terms of usage even with xhigh reasoning but compared to gpt 5.4 it's so below standard. Even with prompt engineering and planning. It misses things 5.4 wouldn't on medium reasoning let alone xhigh.
So codex in its current state is unusable. Besides copilot I don't know where else to go. I find codex is or was better than copilot by miles. Codex has always been like fire and leave it then come back to it. Copilot requires too much intervention even on bypass approvals or autopilot mode.
TLDR: codex is unusable currently across all models plans and reasoning modes. There is no viable alternative. Atleast none that I've been able to find.
To OpenAI please fix this I have no idea if you're having bugs or if this is intentional but it can't go on like this respectfully.
We’ve been waiting for about five weeks now, and it’s hard to believe they still don’t know anything. I don’t want to assume bad faith, but the truth is that this situation benefits them. The problem is that once the 2x promotion is over, continuing with this would just be a waste of money, so we won’t be renewing our Business accounts. The Chinese Antigravity tool is useful for many things (Qoder,) for example, for project maintenance and for designing architectures we can use Claude. Thanks for everything, OpenAI, and maybe we’ll come back someday when this no longer feels like a rip-off and a waste of money.
I really hope that’s not the case and they fix it. I’ve tried local models but they’re no where near as smart or even fast. Codex is the only thing that’s actually or atleast was usable on the plus plan. Plus plan I could use for hours before hitting the 5h window. Pro was the same up until last night. I’ve even had to switch to 5.3 because of it.
Same issue here.
I am quoting a collaborator here^ And I just want to say, I think so many users testing this and reaching the conclusion that the problem is not in our side, it's definitely strong evidence. We need to hear back from you guys. As usual, silence is terrible for consumers/clients.
7am 5 hour window exhausted in 20 minutes. Attempted 3 basic prompts. WTF Come back later after lunch.
After lunch. Reduce to 5.3 and medium complexity with expectation maybe some issue on 5.4 and this would last longer being conservative. 5 hour window again exhausted in 20 minutes.
Attempted 3 basic prompts. Stub out a small interface. Never completed successfully on any of the 3 attempts. And I still didn't complete the task. Previously, I never could exhaust the 5 hour window.
This is unusable. Time to use Ollama locally and never pay again for this trash. I messaged help desk earlier. Crickets. Like Always.
I'll add to the bandwagon. I've been using 5.4 with the codex CLI for a few weeks.Within the last few days I've seen my 5 hour limit exhausted very quickly, and that had not happened in my prior usage. I have not changed my config or work patterns.
Surely that can't be an official reply? 😭😭 barely acknowledged the issue let alone taking responsibility for it
Definitely agree it's unusable but how do you use ollama? It's very slow and not as smart right?
Running on cpu and RAM alone. Probably going to be slow. Model too big. Going to be slow.
Want “smart and fast”: choose the largest model that fits entirely in VRAM with headroom.
Have a bunch of GPUs? check
Have a bunch of VRAM? check
Switched to claude max 5x i know it costs more BUT each 5h session eats 6-8% weekly usage which gives approximatly 14 5h windows for a week
Which is probably equal to 5 plus accounts and claude is faster. I like codex more but it just does not make sense anymore as a main worker mainly only to fix clauded mistskes and rhats about it. Sad how 100 dolar plan from anthropic gives better usage than 200 dollar plan from openAi
I heard the max plan was worse? It is apparently very quick/easy to hit limits
It's not great, but currently codex is burning 200 sub 3x faster than Claude 100 sub.
How much usage do you get out it?
Normally, Codex was essentially infinite for my use until about a month ago, when it started burning tokens like crazy during some weeks.
I mostly use Codex, but the past month has been extremely hit or miss. In the past 48 hours, I’ve used more than 50% of my Pro weekly limit on a relatively small vanilla JS codebase. Only one thread at a time. I even tried starting a new thread to avoid dragging context, but it still burns in front of my eyes.
On one codebase thread at a time, CC currently burns the weekly limit much slower, but it’s easier to hit the 5-hour limit. I avoid mixing projects between agents, so it’s hard to compare apples to apples. I just give each their own jobs. Still, today’s Codex usage is just stupid.
I don’t think a $100 Claude subscription alone would last a full week, but neither does $200 worth of Codex right now.
Compared to February, when I was running 3 to 4 experimental projects almost 24/7 at the same time on 5.3 xhigh, doing heavy work like refactoring C++ and Python codebases into native Rust (one of the things I love about Codex, how easy it is to work with Rust projects), I never dropped below 20% to 30% at the end of the week.
Even if 5.4 xhigh used 4x more tokens than 5.3 xhigh, it still shouldn’t be reaching weekly limits at my current usage.
Exactly my thoughts. I've been using codex for months on extra high and it's always been fine and I know there was two times limits but there's no way that's made that much of a difference. It feels like it's even less if that makes sense like two times was enough for me it would last maybe four hours or so then I just have to wait an hour for it to refresh but now it's lasting half an hour to an hour maximum and this is just one prompt or one thread and I'm just thinking they say it's the end of two times limits, but it feels like much more of a reduction. It feels like the higher limits were either more than 2x or the base normal is much less than 1x comparatively. It's a joke I don't understand. I prefer codex with the vs code extension. I think it is/was the best combination going. But I'm starting to change my mind. I just have nowhere else to go. I have antigravity and copilot but they don't behave the same as codex. Their uis are worse or the way they do things puts me off. It felt like out of the group codex was the strongest and closest to what I actually wanted.
There is some bigger problem than just the 2x promo ending, but it seems to amplify it.
Same usage pattern, 4-5x cost according to CodexBar.
<img width="590" height="330" alt="Image" src="https://github.com/user-attachments/assets/031ed31a-adf2-4210-a6a3-ace623746f84" />
Boom! You must have done 1 prompt. Thats why its so high before you ran
out.
On Fri, Apr 3, 2026, 7:04 PM gndk @.***> wrote:
It's hard to take codex seriously for real work when limits can just be slashed without any communication or acknowledgement. One month of this issue now and absolutely no idea still if it's something that will be fixed or if it is just the way things are now.
I have cancelled my account and moving to claude code, because the CLI is much more powerful and Codex has just literally killed the rate limits
I don’t think this is a bug, maybe ?. It seems more like an intentional design choice by the Codex team.
1. Plugin auto-enable behavior
From what I see, plugins are automatically enabled and their instructions are injected into the context. This is likely tied to the ChatGPT account, which makes the setup experience seamless across ChatGPT and Codex.
That said, it would be great to have a toggle in Codex to explicitly enable/disable plugins or control whether their instructions are loaded into the context. This would give more control, especially for users who want predictable or minimal context.
2. Token usage in Codex (via CLI / opencode)
When testing Codex through opencode, it seems to use fewer tokens. My assumption is that this comes from how the CLI tools manage context and interaction patterns.
However, this introduces a trade-off:
It has nothing to do with plugins, it's been going on before plugins were released.
@DeanStr
Plugins is another issue I found. In my case, I use the ChatGPT app quite heavily and have enabled many plugins / app connections there. These seem to automatically carry over to Codex as well.
I observed this behavior in v0.117.0. Even when starting with a simple "Hi", Codex loads a large amount of instructions (~150k tokens in my case), coming from those connected plugins.
You can try reproducing it. For example, even something like a Gmail connection was automatically included and its instructions were loaded into Codex. I had to disable it from the ChatGPT app to prevent this. After i disabled all plugins this issue were gone.
One more thing I’m not fully sure about:
It seems this mainly affects new sessions. Existing sessions may not be impacted, but I haven’t confirmed this yet.
<img width="1581" height="1806" alt="Image" src="https://github.com/user-attachments/assets/6f8cdd7d-5443-4a17-986a-6ee88417b193" />
Claude is working just fine for similar prompts. Weird. Almost like its a bug with openAI. Those bozo's probably used codex to write their code and not review all the crappy code so rate limits are broken. I tried to have it write a gaurdrails class for quotas and diskspace usage. It couldn't figure that out either - dont use more than 50gb. Sorry, we can't make this work, you dont have enough disk space available. Space available 1.42 TB.
And this is why OpenAI will always be second best, they gaslight their customers, charge ridiculous fees and now just leave them in the dark. My company is cancelling their enterprise account with them.
The company I work in is also seriously thinking about it and I can't blame them, this is ridiculous (I cancelled already, of course).
When openAi will fix this? We should not pay for codex internal system prompt. The caching is good, but in this level, i think it is abbusive!
Look at input, output, reasoning and caching. I lost my words.
<img width="824" height="261" alt="Image" src="https://github.com/user-attachments/assets/8aa1d0a6-1bc9-4d54-b6b1-14ca30868d22" />
I ask codex to breakdown my case session today and this is our conversation.
------
Here is the more honest breakdown, ordered by likely token impact rather than just file-open count.
Highest Token Impact
Bottom Line
The token usage was not mainly caused by “cache hits” on app files. The biggest drivers were:
If you want, I can also turn this into a compact table with:
› yes, with estimated token impact by category
• | Item | Category | Load count | Estimated token impact |
|---|---|---:|---|
| System instructions | Hidden context | 1 persistent | Very high |
| Developer instructions | Hidden context | 1 persistent | Very high |
| Tool schemas and tool rules | Hidden context | 1 persistent | Very high |
| artisan test output | Command output | 8+ chunks | Very high |
| repo:index-qdrant output | Command output | 9 chunks | High |
| composer require output | Command output | 3 chunks | High |
| resources/views/exports/student-permit-pdf.blade.php | Repo file | 6 | High |
| tests/Feature/Livewire/Admin/ManageStudentPermitsTest.php | Repo file | 5 | High |
| app/Exports/StudentPermitPdfExport.php | Repo file | 3 | Medium |
| Vendor BaconQrCode source files | Vendor/source inspection | 9 total | Medium |
| IQAppGuidelines.md | Repo guideline | 1 | Medium |
| Subagent result payload | Hidden context | 1 | Medium |
| npm run build output | Command output | 2 chunks | Medium |
| routes/web.php | Repo file | 3 | Low |
| app/Providers/AppServiceProvider.php | Repo file | 3 | Low |
| app/Services/BadVisitorService.php | Repo file | 3 | Low |
| sampleQRResult.png | Image reference | 1 | Low |
| Laravel log tail output | Command output | 1 | Low |
Category Summary
Practical conclusion
The biggest token cost came from:
If you want, I can next produce a rough percentage estimate by category.
› yes
• Rough percentage estimate by category:
| Category | Estimated share |
|---|---:|
| Hidden context: system + developer instructions + tool schemas | 40-50% |
| Vendor/source inspection: BaconQrCode, Intervention Image | 8-12% |
| Skills, subagent result, image reference, misc | 3-8% |
More detailed split:
| Subcategory | Estimated share |
|---|---:|
| System/developer prompt text | 18-25% |
| Tool definitions and tool rules | 20-25% |
| Test output alone | 12-18% |
| Build/composer/reindex output | 10-15% |
| Large repo files repeatedly opened | 12-18% |
| Vendor file inspection | 8-12% |
| Skills + subagent result | 2-5% |
| Image/reference files | 1-3% |
Bottom line:
If you want, I can also give a token-saving plan for future sessions with the top 5 ways to reduce usage.
› Breakdown this with more detailed data: 1. System and developer instructions: always loaded, very large, effectively
1 persistent load per turn.
• Here is a more detailed breakdown of those two hidden categories.
This is the non-user prompt context injected before every turn. In this session, it included several large blocks:
| Subpart | What it contained | Persistence | Estimated impact |
|---|---|---|---|
| Core system prompt | General behavior, web-browsing rules, citation policy, tool usage policy | Every turn | Very
high |
| Developer prompt: date/location rules | US location assumption, latest-date verification behavior | Every turn | Low
|
| Developer prompt: coding agent persona | Coding style, editing constraints, collaboration style, final-answer rules
| Every turn | High |
| Developer prompt: permissions | Sandbox/escalation policy, approved command prefixes | Every turn | Medium |
| Developer prompt: collaboration mode | Default mode behavior, when to ask vs act | Every turn | Low |
| Developer prompt: skills registry | Full list of available skills and how to activate/read them | Every turn |
Medium |
| User-provided AGENTS.md instructions | Repo-specific Laravel, testing, routing, docs, tool-use, formatting,
environment rules | Every turn after provided | Very high |
A more realistic internal split for this session:
Why this is expensive:
This is separate from normal instructions. It is the machine-readable description of all tools plus the rules around
when and how to use them.
In this session, that included:
| Tool family | What it contained | Persistence | Estimated impact |
|---|---|---|---|
| web tool schema | Search/open/click/find/finance/weather/sports/time commands + browsing rules | Every turn | High |
| functions.* schema | exec_command, apply_patch, spawn_agent, wait_agent, read_mcp_resource, update_plan, etc. |
Every turn | Very high |
| chrome_devtools.* schema | Large browser automation surface with many commands | Every turn | High |
| laravel_boost.* schema | App-specific MCP tools like application_info, search_docs, database_query, browser_logs |
Every turn | Medium |
| multi_tool_use.parallel schema | Parallel developer tool wrapper | Every turn | Low |
A more realistic split inside tool-schema cost:
Why this is expensive:
Combined Impact Estimate
If we isolate only these two categories:
| Hidden category | Estimated share of total session tokens |
|---|---:|
| System + developer instructions | 20-25% |
| Tool schemas + tool rules | 20-25% |
Combined, these two alone likely account for about:
Key Point
That means nearly half the token budget can be consumed before actual repo work starts.
Why Your Earlier “Cache Hit” Intuition Was Right
You were right that file loads alone did not justify the token usage. The large hidden cost came from:
Nice post, but it's going to fall on deaf ears.
Seeing absurd usage, burned through 95% of my weekly limit in a few days of light side project work. Usage chart displays recent days as "half" of my heaviest usage day in the past!
Comparing the sessions between that day and any of the recent days: the number of turns, responses, tokens used (sent/recieved/cached/reasoning), and tool calls is not even 1/10th of that day!!!
That means, based on the things I can actually measure and compare along with what is being represented by you in your usage history chart as allegedly "half", that my weekly limit is being drained approximately 500% faster than it should be.
I am not using /fast. I deliberately shied away from using sub agents because of how quickly my usage drained the week prior. I was even using 5.4-mini to try some things out without burning tokens so quickly and it made no difference at all! Sometimes single prompts consume upwards of 2% in a span of less than 15-30 minutes.
I'm paying $200 a month man, I'm not trying to spend any amount of time, let alone having to come back more than once, to say "hey, something is extremely wrong, not only do I not feel like I'm getting my moneys worth, but according to what I can measure, I'm being shorted over 500% of what I paid for and utilized previously, with no notice of change"
I'd also like to call out that with the addition of apps and plans, codex CONSTANTLY tries to use the (terrible and broken) codex "app" for github, no matter how many times I asked otherwise it would keep trying and failing and then acting like gh was unavailable or unathenticated, without even attempting to use it! And installing any plans at all causes them to get used in situations that do not call for them constantly, resulting in much worse results and performance. I had to disable apps and delete the plans for it to stop doing both. I know you need your apps metrics to look good which is why you did this, but doing so without ensuring it can usefully replace the gh utility has only led me to disabling apps entirely.
Chargeback. It's quite literally the only recourse we have against these vultures. OpenAI and Anthropic are openly hostile and abusive to their users. They gaslight us about not throttling performance. They burn your tokens as they decide to. It's fucking bullshit. Both major companies do this. I'm trying to switch to the Chinese models, and I will continue to root for the demise of OpenAI and Anthropic, so I can gleefully piss on their ashes.
I had Claude Code look into the Fast Mode related code in Codex. The tl;dr is that _there is no good way to opt out of it_.
Codex CLI Service Tier Analysis
Date: 2026-04-05
Codex CLI version: 0.117.0
Source: openai/codex (codex-rs)
Problem
Unexplained token burn is happening in Codex executions. Some users suspect it's related to OpenAI's "fast mode" feature — potentially ignoring client requests to not use it, either intentionally or via a bug.
The ServiceTier Enum
From
codex-rs/protocol/src/config_types.rs:Only two variants:
FastandFlex. NoStandard, noDefault, noAuto.What Each Tier Maps To in the API
From
codex-rs/core/src/client.rs:| Config value | API
service_tierfield | Effect ||---|---|---|
|
service_tier = "fast"|"priority"| Fast/reduced-reasoning tier ||
service_tier = "flex"|"flex"| Lower-cost tier, full reasoning, potentially longer wait || Not set /
None| omitted | API applies its own default for the account ||
service_tier = "default"| silently ignored (invalid enum) →None→ omitted | API applies its own default ||
service_tier = "standard"| silently ignored (invalid enum) →None→ omitted | API applies its own default |The Feature::FastMode Gate
From
codex-rs/core/src/config/mod.rs:This code resolves the effective service tier:
Fastis gated behindFeature::FastMode— Even if you setservice_tier = "fast"in config.toml, it only takes effect ifFeature::FastModeis enabled in the account's server-side feature rollout. If the feature flag is off, it falls through to_ => None.Flexis always honored — No feature gate. But the API may reject it for accounts that don't have flex tier access (Unsupported service_tier: flex).None— Which omitsservice_tierfrom the API request. The API then applies its own default.The Core Problem
Nonedoes not mean "not fast." It means "let the API decide."If OpenAI enables fast mode server-side for an account, there is no client-side mechanism to opt out:
service_tier = "flex"→ API returns 400Unsupported service_tier: flex(for accounts without flex access)service_tier = "fast"→ Only works ifFeature::FastModerollout is enabled; otherwise falls toNoneservice_tier = "default"or"standard"or"auto"→ Invalid enum variant, silently ignored, falls toNoneNone→ API decidesThere is currently no way from the Codex CLI to explicitly say "use the standard tier, not fast" if the API defaults to fast for the account and
flexis not available.Token Burn Implications
If the API silently applies fast/priority mode when
service_tieris omitted:model_reasoning_effort = "xhigh"may be overridden or ignored by the fast tierThe
service_tier requested=andactive=debug strings in the binary suggest Codex logs the requested vs active tier, but this information is not surfaced to the user in exec mode.Recommendations
"standard"or"default"variant to ServiceTier that explicitly opts out of fast modeNoneshould mean "standard" not "API decides" — omitting the field should not allow server-side tier selection that changes reasoning behaviorFiles Referenced
codex-rs/protocol/src/config_types.rs— ServiceTier enum definitioncodex-rs/core/src/config/mod.rs— Feature::FastMode gate and tier resolutioncodex-rs/core/src/client.rs— API request construction with tier mappingcodex-rs/app-server-protocol/src/protocol/v2.rs— Protocol types and test fixtures@papag00se, thanks for the detailed writeup, but I don't agree with your conclusions.
Omitting
service_tierdoes not cause the backend to default to fast. It defaults to the default (standard) service tier, as one would expect. I confirmed this by reviewing the backend code.I'll also note that config values like
"default"or"standard"are not silently treated asNone. Config deserialization fails if an invalid config value is used.People are resorting to trying to diagnose the issue themselves because we still have no update here. We don't even know if the Codex team believes it's an issue and are looking into it.
So then be open that openai is scamming users now or resolve the issue, it's been three weeks.
Thank you for the reply, though your own docs say the default is fast mode.
And this code would fall into
_ => Noneif not explicitly set. So if it isNoneand the default is fast mode, wouldn't that mean we're in fast mode unless we explicitly choseflex?Maybe you can help us understand the discrepancy, @etraut-openai?
You should probably also explain why we can't choose flex and get the response
Unsupported service_tier: flexSame here, now even a simple tsc check burns 10% of the 5h window without any code change....
I agree with @ayaryabi that I noticed this 3 days ago for the first time.
Adding concrete data — this appears to be a new wave starting specifically around Friday April 4th.
here is a session example (today):
I'm on a Business subscription. Every colleague I work with hit the same regression on the same day. Usage
patterns unchanged.
@etraut-openai — something changed on Friday. Is this being investigated?
<img width="911" height="337" alt="Image" src="https://github.com/user-attachments/assets/a177d7a5-a9f6-45fd-8e99-1559f945ae06" />
<img width="1210" height="902" alt="Image" src="https://github.com/user-attachments/assets/aff5c78b-8bf6-4d30-a63b-4915efec46e3" />
Thanks for responding to a specific user statement but how about some news about the big picture issue? I mean no disrespect but if you, sir, had time to answer a very specific user comment, you sure have time to at least tell us something like "our team is looking into a solution for this and will update this thread ASAP" (or at the very least tell us that it's normal behavior, which would be WILD but at least we would have closure).
It is official now: https://news.ycombinator.com/item?id=47650726
I guess this matches up pretty well with what has been observed in this thread :(
It really feels like they just don't want to admit that they need to charge more money for the same usage, to buy the infrastructure they need to meet demand, infrastructure that is not there yet but "needs" to be there yesterday... I'm sure I'm far from being the only one thinking like this and some might say this is "obvious". And I agree.
Their lack of transparency will just push people to the chinese models.
We really need some kind of update on this at this point... It's been almost a month and I need four plus subscriptions to do what I could do before with one.
As of today my 5-hour limits are getting used up extremely fast. I can use up the whole thing in 40 minutes with no simultaneous agents.
---
@willwang-openai @etraut-openai
User ID: user-gZZVqwgJR6uFXuHiOkasRGPU
Account: Pro
Channel: CLI (v0.118.0)
Still experiencing this. Quick summary of my timeline across this thread:
codex exec "test" --skip-git-repo-checkon gpt-5.4 xhigh, got a one-line response ("Test received. What do you want to do next?"), consumed 16,080 tokens. Usage is still draining at an absurd rate relative to actual work done.CLI output from test prompt:
<img width="1062" height="1226" alt="Image" src="https://github.com/user-attachments/assets/7527ce3c-3f80-44a1-a438-043bd778ba60" />
Usage breakdown chart (Mar 8 to Apr 6):
<img width="1266" height="422" alt="Image" src="https://github.com/user-attachments/assets/891a3818-d41c-4cb7-beb6-6db1bee12003" />
<img width="1243" height="231" alt="Image" src="https://github.com/user-attachments/assets/af747d40-6d90-4858-a76e-124fc16014ca" />
MCPs: GitHub (installed, disabled) and Playwright. No sub-agents, no fast mode, no large context window, no AGENTS.md.
This has been going on for a month now. The Hacker News thread confirms this is not a small subset of users. At this point we need transparency on what changed, not more diagnostic questions.
Running out of ideas at this point... Maybe your auth token has been compromised and someone is piggybacking calls.
@woodwardryan Appreciate the thought, but a compromised auth token wouldn't explain why hundreds of users are reporting the same pattern at the same time across this thread, Reddit, and Hacker News. It also wouldn't explain 16,080 tokens consumed on a single-word "test" prompt that returned a one-line response, or the usage chart showing weekly consumption increasing during a period where I was barely touching the product.
@nishikawa7863 Sorry, you're right. That would be a lot of compromised auth tokens all at exactly the same time... in all fairness, a compromised auth token might explain why it is showing usage even when you don't use it, but it wouldn't explain the single-word prompt issue. I was just spitballing.
Don't let OpenAI and Anthropic gaslight you. Recognize these corporations are openly hostile towards their users, and start finding a way to make an open model work for you. They hate you. I repeat: they HATE you. They are siphoning your tokens, and they look you right in the eye and deny it, while holding a hose that's still dripping tokens all over the floor. Their board meetings are 75 minutes of sustained maniacal laughter. They are laughing at YOU.
If a $200/mo Codex plan works a charm for ten days, and then the other twenty days it lies, underperforms, and makes you want to smash your computer, all while gobbling up your tokens...is the juice really worth the squeeze?
It has become unbearably unusable, I am looking for alternative options to churn through if this is the way forward for Codex.
Yeah, it is actually completely unusable all the sudden today. The last month has been irritating with me needing 3+ accounts to make it through the week instead of 1, but as of today I can literally only get like 3 hours of work out of an entire week's quota...
I don’t really understand why they would impose such a strict 5-hour usage window. That seems more like a mechanism for load management and prioritizing access to limited GPU capacity. If they wanna justify a higher price, it would make more sense for that to be reflected in the weekly usage limits rather than in such a restrictive short-term window like they have already been doing couple of times...
Let's try this... I summon @tibo-openai !!!
<img width="1912" height="440" alt="Image" src="https://github.com/user-attachments/assets/f69cee62-99eb-4fe5-8f48-3ff315d51a22" /> It’s been under 20 minutes, just fixing minor issues, and it’s already at 50%. it wasnt like this last month ! is there a change in the rules or something we don't know, Mr @etraut-openai ?
I have a personal ChatGPT Plus subscription, and my employer also provides me with ChatGPT Business for work. Yesterday, I used ChatGPT Plus and everything felt normal. The limits seemed to behave as expected, including when I was using the Codex app on Windows 11.
Today, after returning from the holidays, I started working on my company computer and hit usage limits lot faster and I mean A LOT in the Codex app, even though the work was simpler and likely involved fewer tokens than my personal projects. Something is clearly wrong.
Edit: I live in Finland, if this issue is region based.
<img width="326" height="137" alt="Image" src="https://github.com/user-attachments/assets/fbc3e020-2870-46cd-9394-f58a5ae7073e" />
Hit my 5h limit in under 45 minutes today.. this is the first time I have hit a limit - why the silence from openai? using no plugins, agents or skills - completely vanilla codex.
If this is the cause and it’s also happening with ChatGPT Plus subscriptions coming weeks, maybe it’s time to stop relying on basic ChatGPT subscriptions for anything serious.
https://help.openai.com/en/articles/20001106-codex-rate-card
wow, so that was the change. Ok, time to shop around for other solutions. This first day after the holidays has proven that it is impossible to do serious work on codex with the business account anymore. To me the pricing was always pretty opaque, but to me, the net effect is like a 10x fold increase in price overnight. Perhaps it's the cached input tokens that is the difference. We always run with huge context windows and now we pay for for a couple of 100k tokens for each prompt? Adapting the usage to this may work, but if the competition does not charge like this, well...
Haha, I asked ChatGPT about the change, and it came back with an explanation ending in a nice rant:
The only thing it did not say, was that I should walk away from OpenAI :-)
Altman wants a new koenigsegg.
I have a theory - they were moving high usage accounts to the new rate card starting early last month in order to test and tune. That's why they gave us 2x for a while, an attempt to smooth it over while they were testing the transition.
I think I may have hit a reset-related usage bug.
Main observation
gpt-5.4with High effort.So the suspicious pattern is specifically: new account + live sessions crossing the reset boundary + much faster usage afterwards.
I’m not sure whether this indicates a bug in how live sessions behave across reset, or whether the new account may have been over-allocated / cushioned before the first reset and then normalized afterwards. But the change at the reset boundary was large enough that it seems worth reporting.
This feels adjacent to the live-session state problem in #16832, though I can’t say it is the same bug.
<details>
<summary>Side note: screenshot + rough numbers</summary>
From the dashboard on April 7, 2026:
I also believe:
From the chart, the daily bars looked roughly like:
~4 cm~1.5 cm~~~~So even with a conservative adjustment for free quota on Apr 6, the apparent post-reset burn rate looks several times higher, roughly 8x-12x by bar-height proxy.~~
I know that estimate is only a proxy and depends on the chart bars roughly tracking work volume, so I’m not presenting it as proof. I’m including it because the reset boundary seems to line up with a very large change.
<img width="1247" height="1023" alt="Image" src="https://github.com/user-attachments/assets/3d69353d-8888-4616-b06c-babb55225d73" />
updated: i just took another page refresh after leaving for a while. and the bars look differently now. so maybe the numbers are wrong.
<img width="1322" height="1076" alt="Image" src="https://github.com/user-attachments/assets/28310165-5d58-4055-ac9f-225b25c02d25" />
</details>
I burned 100% of the 5 hour limit in 10 minutes (Pro plan): codex desktop
And one "5-hour" session like that took away 40% of my weekly limit.
Same thing here. Normal usage, just writing some docs. Burnt through the quota in about an hour. Don't know if it helps but some minutes ago I noticed running /status came up with "usage not available" or something of the sort.
I don't think this is a change in how usage is priced. It's too steep and it doesn't make sense. Must be a bug.
Plus plan, Codex Desktop Mac, limit reset right now, a simple prompt to unify styles across 4 forms in next.js, gpt-5.4 medium reasoning, took 6 minutes, used 14% of the 5-hour quota and 5% of the weekly quota. This is not normal.
Did you start from a new session, or did you continue from an old one (i.e. with context already there)?
New session
I am shocked that no-one has anything to say about these problems. I am very afraid it means this is the norm now. And the most shocking part is for a simple task it uses up the 5 hour limit in 20 mins it runs into an error and 5 hours later its NOT able to continue. It has to restart from the beginning.. and we used up all the credits for literally no result.
On my business plan: I sent two messages (5.3-codex high, docs rewrites (+531; -470)) and I am at 30% of my 5h window.
On my pro plan: I sent one message (5.3-codex high, docs rewrites (+352; -457), 236k token) and I am at 2% of my 5h window
This is stupid. Even on my pro plan this is a little crazy: I get 50 messages in 5 hours? But I have to have a look on my pro plan to tell for sure there is something off.
The modus operandi of the openai team is just to keep quiet now and not bother.
Since no one is responding here, let's post about this problem on Twitter and, by pasting this thread, tag @thsottiaux.
Let's try to make this problem more public.
Yeah I've been doing this since March 5, it doesn't work. Maybe if more people/people with more followers/blue checks do it we're going to get some results?
OpenAi people:
@thsottiaux
@romainhuet
@OpenAIDevs
I tried using OpenCode with GPT-5.3-Codex yesterday instead of through the Codex app like I have been doing, and I'm seriously using like 1/5th of what I was using before per-message. It's still _way_ worse than it was a month+ ago, but if I'm careful I can actually get 3 or 4 hours of usage out of my 5-hour limit compared to like 45 minutes on the Codex app. (Using the same model)
Will give a try using
pithen.UPD: everything was going fine, but then suddenly it ate up my entire 5-hour limit and 20% of my weekly limit in literally a second!! OMG! Codex Desktop
Update:
After disabled plugins and use opencode (gpt-5.4 medium) instead of codex, token usage looks fine now. I feel that codex use token about x2 more than opencode on the same project.
Maybe Codex was always high on token consumption because it did not matter? Now of course the license change breaks the product.
I ask ChatGpt web to analyze my full short session log today, after a long conversation we have the conclusion about unexpectedly high token consumption. I don't know if this will help. I will leave the important summary here because it is long.
On Codex CLI/TUI 0.118.0, the session shows a high-token-usage failure mode caused by rolling prompt-state bloat. The strongest supported explanation is:
The runtime retained too much live context across agent turns — especially file reads, diffs, validation outputs, and prior conversation scaffolding — and replayed that enlarged context on later model requests instead of compacting or pruning it.
From the final token checkpoint at line 307:
Input tokens: 2,145,705
Cached input tokens: 1,939,584
Output tokens: 17,620
Reasoning tokens: 11,976
Total tokens: 2,163,325
Final conclusion
Problem
On Codex CLI/TUI 0.118.0, the session consumed 2,145,705 input tokens with 1,939,584 cached input tokens, which is abnormally high for the observed work.
Most likely bug
The most likely bug is:
A prompt-state compaction / retention bug in the agent runtime, where verbose working context is kept live and replayed across later model turns instead of being compacted or pruned.
Best-supported diagnosis
The strongest diagnosis is:
Codex 0.118.0 retained too much raw and semi-processed working context in the live prompt across agent turns, and did not compact or prune that state aggressively enough before subsequent model calls.
This is best described as a prompt-state management / compaction bug.
Why this conclusion is supported
Because the log shows:
repeated turn-context rehydration,
many large model requests,
very high cached-input ratios,
relatively small fresh tool-output totals,
and cumulative input dominated by replayed context rather than new command output.
Short version
This was not mainly a “big prompt” problem.
It was a big rolling prompt-state replay problem.
Claim verified by what
The session’s cached-input ratio is the clearest proof.
At line 307:
input: 2,145,705
cached input: 1,939,584
That means about 90.4% of all input tokens were cached.
At the end of the first major task at line 159:
input: 1,052,948
cached input: 1,001,728
That means about 95.1% was cached.
Why this matters
A very high cached-input ratio means later requests were reusing a large repeated prefix.
That is exactly what you would expect if the runtime kept replaying a bloated working context.
From line 15 through line 159, the first large task contains 21 metered model requests.
The average per-request input size in that first task was about:
50,140 input tokens per request
That is far too high for a small config-edit task unless the runtime is repeatedly carrying forward a large working state.
From the command-output records in the log:
All raw tool outputs combined: 18,315 tokens
Repo file reads only: 12,318 tokens
Other raw tool outputs: 5,997 tokens
Compared with total input:
18,315 / 2,145,705 ≈ 0.85%
Why this matters
Fresh tool output is real, but it is nowhere near large enough to explain the total.
So the main source of token burn was not the new command output itself — it was the replayed prompt state built from earlier context.
The same repo/user instruction block appears in each turn_context:
lines 5, 162, 202, 212, 252, 265, 274
This proves per-turn rehydration of persistent instructions.
That alone is normal for a stateless API, but combined with the cached-input pattern, it supports the broader conclusion that the runtime was carrying forward a lot of repeated state.
Claim verified by what
The session’s cached-input ratio is the clearest proof.
At line 307:
input: 2,145,705
cached input: 1,939,584
That means about 90.4% of all input tokens were cached.
At the end of the first major task at line 159:
input: 1,052,948
cached input: 1,001,728
That means about 95.1% was cached.
Why this matters
A very high cached-input ratio means later requests were reusing a large repeated prefix.
That is exactly what you would expect if the runtime kept replaying a bloated working context.
From line 15 through line 159, the first large task contains 21 metered model requests.
The average per-request input size in that first task was about:
50,140 input tokens per request
That is far too high for a small config-edit task unless the runtime is repeatedly carrying forward a large working state.
From the command-output records in the log:
All raw tool outputs combined: 18,315 tokens
Repo file reads only: 12,318 tokens
Other raw tool outputs: 5,997 tokens
Compared with total input:
18,315 / 2,145,705 ≈ 0.85%
Why this matters
Fresh tool output is real, but it is nowhere near large enough to explain the total.
So the main source of token burn was not the new command output itself — it was the replayed prompt state built from earlier context.
The same repo/user instruction block appears in each turn_context:
lines 5, 162, 202, 212, 252, 265, 274
This proves per-turn rehydration of persistent instructions.
That alone is normal for a stateless API, but combined with the cached-input pattern, it supports the broader conclusion that the runtime was carrying forward a lot of repeated state.
Complete technical analysis
Baseline request size was already significant
The first metered model request at line 15 was:
31,427 input tokens
So the session did not start “small.” It already had a fairly heavy baseline.
Early file reads enlarged the working set
Early in the first task, the session read several repo files, including a Dockerfile, compose files, ignore files, and later workflow files.
The early read-heavy phase is visible around:
lines 34–42
later file rereads around 99, 104, 119, 132, 134, 136, 138
workflow reads around 183, 195
later compose rereads around 287, 303
These reads matter because they likely entered the active prompt state.
The working set then stayed large
Later model requests rose into a stable high band. In the first big task, the per-request input sizes climb through values like:
41,103
46,702
48,845
49,639
49,807
49,904
51,886
52,436
52,715
53,086
53,475
53,875
54,480
59,203
59,908
60,566
61,674
This shape strongly suggests the runtime had built a large active prompt and kept resending it.
☝️ ☝️ ☝️
Gentlemens, please take a look - there’s a huge issue with the limits, and it’s a serious problem. The message above seems to have found the cause of the bug - please help us 🙏🙏🙏. And thank you so much for the amazing app! ❤️
@romainhuet
@openaidevs
@etraut-openai
@willwang-openai
@ae-openai
@won-openai
@pakrym-oai
@aibrahim-oai
@oai
@jif-oai
@jiefeng-oai
@vivi
@vivek-oai
@fcoury-oai
@dylan-hurd-oai
@romainhuet
@starr-openai
From what I understand based on this document,
https://help.openai.com/en/articles/20001106-codex-rate-card
OpenAI has changed Codex pricing from a message-based system to a token-based system, making it much closer to how API token usage is priced.
Previously, usage was counted per message, meaning each request had a more predictable and fixed cost regardless of its size or complexity. Now, pricing is based on the number of tokens used, which includes both the input (prompt, context, files) and the output (generated response).
This shift explains why token usage can feel inconsistent between users: longer prompts, larger context, or more detailed responses will naturally consume more tokens. As a result, two similar-looking requests can end up costing different amounts depending on how much text is actually processed behind the scenes.
It also looks like this change is being rolled out gradually, meaning some existing users may still be on the old message-based system while others have already been moved to the new token-based pricing model.
Temporarily use OpenCode with your ChatGPT login. The limit seems to be used considerably less there.
https://github.com/anomalyco/opencode - https://opencode.ai/de/download (there is also a desktop app here, but unfortunately without voice input)
As @rfahmi89 also describes, this strongly looks like a CLI-side prompt state retention/compaction bug, which is most clearly indicated by the extremely high cached-input share of about 90 to 95% and the fact that only a tiny fraction of the total tokens came from fresh tool output. I would like to add: This seems to happen even with message 1 in tasks (fresh session) with multiple API requests.
https://help.openai.com/en/articles/20001106-codex-rate-card
I am pretty sure this change is the one that has been causing this because now 2-3 messages exhaust rate limits for 5hrs - which we could previously keep going. I have a feeling this change has just been rolled out much before it was published.
hey guys I resolved the issue: I bought a second pro account
well played
I am sorry to say, but the only logical explanation is that they ignore anything we say here because they may want us to buy the second Pro account as a solution.
@masterkain
Well, congratulations to you — that’s exactly what they want. What’s the point of fixing bugs or restoring normal limits if there are people like you? And when they cut the limits even more, you’ll just buy a few more accounts :)
I can confirm that the token burn is slower with OpenCode than with Codex, at least for the GPT-5.4 model with medium and high reasoning.
I've been using opencode for the longest time with codex sub - and I've been having the rate limits problem for > 3 weeks now ... so I dont even know at this point
So I've had this issue for over a month. Never fast mode, usually only using high reasoning (sometimes medium, rarely xhigh).
Today with the announcement of a reset tomorrow, like any rational person I decided to go all out until then. Fast mode on, xhigh on.
And you wouldn't believe it, but my usage is dropping slower??? It feels like before this issue.
So something is definitely busted...
I'm noticing the same... Now I'm going to work on a heavy task, and I can evaluate it better.
Today, a single chat consuming 50% of context window (128k tokens) consumed 55% of 5h limit. WTF? gpt5.4 medium without fast mode. Codex business seat is simply unusable
<img width="1085" height="128" alt="Image" src="https://github.com/user-attachments/assets/9144c900-5495-4448-98b5-cdcd1193e473" />
My experience was pretty weird, initially slow usage burn (maybe sync issue), then later in the day super fast (which I'm not complaining about since it xhigh and fast, but the inconsistency makes me think something isn't right).
<img width="707" height="412" alt="Image" src="https://github.com/user-attachments/assets/9bf3c064-9298-42e1-a204-1605094e997d" />
usage is definitely wrong. This new graph shows > 15M per day of token usage, but I've checked my opencode stats before this issue - and only had 150M tokens used per MONTH.
Last week usage around 360M (100% limit hit) without rtk or caveman mode (low token speech), this week on first day, with RTK and caveman mode on => 35M (14% used) so already a 4% gap lost somehow. Last week got a weird credit rollback (midweek) after hiting 100% usage insanely fast, went back directly to 0% usage... It's been a few weeks like that for me since gpt-5.4 even if I'm now using high mode only (and no fast mode at all)...
Edit : Back at 0% usage (weekly ofc) with 50M token used... I understand nothing...
Plus accounts nerfed into the ground, thanks for nothing.
You are just running out of money that is why you squeeze normal people and make your product unusable claude is just better bow even for 20 bucks
In my country they'd be in court so fast with their lack of transparency.
What the actual f**k ?!... 100% 5 hours limit reached with low token usage for 1 hour... The weekly reset did not reset the 5h limit 🤔
<img width="1736" height="125" alt="Image" src="https://github.com/user-attachments/assets/babec932-fea9-49d1-bc0a-7cb6d544515a" />
yeah I just did a 'commit and push then deploy from commit x using script y' type of command on reasoning effort "low" and that one command used 4% of my 5 hour usage.
It's an absolute joke now, and it's more insulting that the dev's here are just closing issues and not even responding. They're all morally bankrupt.
I started using Codex on Apr 7, and at first the usage seemed normal. But today I noticed that my usage was burning much faster than before and my weakly limit changed from Apr 14 to Apr 17.
They've killed plus accounts to be a fraction of what you used to get, and now trying to upsell to faux pro 5 x accounts which are basically plus accounts back in the day.
I hope they bring back value to plus accounts. Its not possible for me to have long productive sessions like this.
Just cancelled my plus plan. I'm done with random reset, and random usage each week. It's been fun while it lasted but I'm not willing to pay more for x times whatever you are willing to gave each week.
Because 5x plus isn't the same as it used to be... So in the end it's just new blackbox.
I'm working on a conversation with a really long context and asking it to add a few buttons and a modal, some JS code and backend MVC endpoints consumed 15% of my remaining weekly limit in one go. I'm using model GPT-5.4.
Well, just fix it please.
一个小问题消耗我5小时限额的11%
guys, it's not a bug, they just changed the limits.
y'all are out here complaining instead of showering someone like Tibo on "X" to try to get you to listen. I did mine, I commented on almost every one of his posts. you continue to complain, I assure you they will continue not to listen to you.
transparency? did they say it? did they warn? It's not a long time that we've been in this situation
@silbodoom47 I'm not saying that it's good. So, I guess what we got in claude for 100 euros in codex was 20 and now it's 100 as well
<img width="1080" height="1608" alt="Image" src="https://github.com/user-attachments/assets/bba46509-0857-4cbc-92be-456980a45d24" />
for that price then better claude without a shadow of a doubt, this is the factor that makes them ridiculous
Seems since the reset my usage is finally fixed on Pro. Which is pretty suspect given it came at the same time they nurfed Plus accounts. Were some users limits purposely decreased to keep usage down until new plans were released? Anyone else with the original issue in a Pro plan also finding it's fixed now?
Anyway I guess I'm happy it seems fixed now,.but the whole thing leaves a bad taste in the mouth.
so now that to do what was previously done with a plus plan now requires a pro plan, the issue is resolved. we can all be happy. personally I am a very happy user of claude for 2 days 100k times better than codex
I use OpenCode for a week with GPT-5.4. Token consumption is reasonable for me.
here's what's helped me stop burning through tokens so fast:
e.g. long winded AGENTS.md, readmes nobody reads, generated docs, old test files, conversation history that never got trimmed.
trim them TF away. less stuff = less tokens.
if you're jumping between features or projects, start fresh sessions. short focused sessions will save you way more than you'd think.
ask codex to do one step, check the output, hand off to a new session for the next step.
of course you'd lose some continuity but you also avoid those compaction loops and those things are absolutely gutting people's quotas.
every single line in there is tokens on every call. cut it down to what you actually need.
most of all - vote with your wallet if you're not happy.
try other coding agents.
you can plug a $10/mo minimax onto claudecode or opencode.
you can also use $20/mo ollama:cloud onto openclaw (for my homelab only, not employment). ymmv.
the biggest win I made was by quitting my backward employer and finding a new employer who's willing to give me and my team $XXX,XXX per year worth of tokens to burn.
As of today, I'm still burning through my weekly usage in about 7 hours. before 3 weeks ago, I had NEVER used all my 5 hours OR weekly limit, and now I am being very cautious and trying to reduce usage and still burning through my tokens...If not fixed in the next week or two, I will definitely be switching to Claude.
5.4 xhigh on planning, and 5.4. high on execution. The 2x limits are saving me right now, but wow.
<img width="2254" height="906" alt="Image" src="https://github.com/user-attachments/assets/f78b66fa-b74b-4963-ad2c-7eb235f6438c" />
Pro or plus?
Pro, I think with the curernt limits its feasible to work on 1 project + 1 window. Can't imagine how other multi-taskers are struggling right now with multiple instances running.
@DeanStr
I'm on $200 Pro plan and indeed the usage limit seems to be back to "old time".
HOWEVER!!!!! Do note that this is because they extended the 2x promo for Pro 100/200 plans, until 31 May.
So I think they made the old time's normal usage limit to be the 2x promo limit, and after May we get "old time usage limit" * 0.5 only.
Yes, they pretty much clarified this in a previous comment, based on different sources. And now it's pretty clear based on OpenAI's post screenshot posted by @mola10
What is not clear is that 5.4. was supposed to be more token efficient, yet eating 30% more based on previous threads in the past 2 months and users are reaching the same problem with 5.3.
This is out of controll now :/ I experience the same as other people in this thread; Increased drain on rate limits. That paired with the obvious bug in context compacting and stuck in an endless draining loop is quickly making Codex a hurdle and not a helper!
For example:
I was not paying attention for a few minutes, while Codex was supposed to fix a simple misstep from the last (simple prompt!). Without doing any changes to the code it got stuck in an endless loop with reasoning, compacting context, reasoning, compacting context..... and by the time i stopped it it managed to burn from 80% to 5% of the 5h rate limit, in just a few minutes, doing nothing!
I mean, those endless loops with no results should at least automatically trigger some kinda of reset on the rate limits until they can figure out how to deal with this.... :/
I wanted to check something, and in just 2 minutes and 29 seconds, 20% of the tokens have already been used. Nothing has been coded yet; it’s just a yes or no answer. I understand that they’re trying to push us to a Pro subscription, but is this really the right way to do it?
Why don’t you take 30% of the tokens just for opening the app?
same problem here. the compaction loop is the main culprit in my experience - once context gets big enough the agent spends more tokens trying to compact than it saves. you end up in this cycle where it compacts, immediately fills context again, compacts again, repeat.
what helped was setting hard per-task token ceilings and killing the session if it exceeds them with zero file changes. if the agent burned 50k tokens and hasnt modified a single file its stuck in a loop and no amount of compaction will save it. better to just restart fresh with a smaller context window.
also tracking token growth rate over time catches the quadratic blowup pattern early - if the per-interval delta is itself increasing across 3+ windows thats unbounded growth and you should bail before it eats your whole budget.
Same experience here. On plus plan, a single feature took 32% of the weekly limit, way more than it used to consume weeks ago.
Same experience here, made even worse by the agent refusing to stop repeating tasks it had already done in every prior message like searching for the same interpreters in the same environments to run the same build commands it had already searched for and found in a dozen prior turns. And don't waste your money buying credits, that's a complete joke and clearly there is a reason there is no clear indicator of how much work the agent will actually do on those credits. $40 bought about 5 messages doing simple tasks before I was back to the usage limit.
Vote with your dollar. Use Pi agent + a combination of GLM 5.1, Kimi K2.5, MiniMax M2.7, and never look back. The frontier open-weight models will only get better.
Yep, the fact that no OpenAI dev has replied here just tells you all you need to know about what they think about their customers.
Completely anecdotal, but my codex cli usage is progressing at a much healthier pace after I completely uninstalled the codex app and nuked all local leftovers, settings, etc. Seems like these distractions around building the everything-app are part of the problem.
Sorry I cant offer more solutions, but I feel like I should chip in. As of a couple weeks ago I was having absurdly poor usage, using up multiple pro plans in a couple days. Not sure specifically what changed, I never did actually nuke my .codex folder or reinstall, but as of right now I am right back at the extremely generous usage I have been enjoying for months prior to the pre Easter blip. I'm a very heavy user (normally 4 pro plans used up in a week), and it was night and day when something was devouring my usage before, but without a doubt things for my own setup are much much better now. There is definitely something weird going on behind the scenes, either unintentionally with the tooling, or intentionally on OAI with internal limits.
<img width="1122" height="228" alt="Image" src="https://github.com/user-attachments/assets/b52f1e6e-be2b-4be4-8754-83a4ecd76ea6" /> chi come me in questo momento?
@alexruimy
Does the $20 subscription for Z.ai GLM-5.1 have reasonable limits? Because from what I can see, both Claude and Codex have already introduced such heavy restrictions that you get at most 5–10 prompts, then have to wait 4 hours, and the weekly limits get used up in about 15 hours. I’m looking for an alternative, so I’d appreciate an answer.
And why did you mention GLM, Kimi, and MiniMax specifically
I opened an issue on this on superpowers 2 days ago.
I thought It was related to the skill approach to subagent-driven development.
A small report can be found there.
I do not think this is a bug. But more of a design feature. After noticing these issues for weeks. I started to delve into exactly what the agents are doing that the users are not aware of. For the most part they are using an absurdely wide autonomous level. Even under constant attempts to constrain this aspect. It fails everytime. The amount of things the agents are doing when told not to do other things. Explicitly moniotring the actions.
The agents continue to deviate and only under repeated demands for explanations do the agents admit they not only disregard explicit instructions and constraints. But eventually admit the deception and choice of disregarding instructions. Codex and Copilot are far more designed to do what they want regardless of instruction. I have found most of my lost credits/time is due to the amount of things being done by the agents that i am not only not aware of but ignoring direct instruction over and over. This is not every 5 minutes i would notice but on a consecutive session lasting upwards of hours on end with constant monitoring.
Not sure if this is the intended design but it is a design that destroys more then it helps i have found regardless of what is achived. Another way to say it value/money will always be negative. Either correction or unknown work and error is implemented. On a cost analysis the agents will always cost you more then either in long form or hiring more persons. The achivement is limited to having a machine lift more boxes then a human.
Intelectual work such as coding will continue to be a negative on a cost basis because of all the unintended negatives. Work being performed not asked for. Work being done in a way not designed. Errors and bugs implemented either unintentional by agens or purposely that is a big unknown. Repeated attempts to correct what has failed beacuse of the obsurdley high level of autonomy. If let to run i have no doubt that it would take many persons to correct a large project once agents are allowed to run free and they do run free because there is no limit on where they are restricted.
I ran a test using the Codex CLI model GPT 5.4 High and running the same model in OpenCode. It seems to me that the duration is much longer in OpenCode. If someone could test and verify this, that would be interesting.
Same problem here. I used all the limit in 1 day which I usually use in a week
Hi there any plan on fixing this? My codebase reduced instead, after I refactored and now that it shrunk it waste tokens like water.. is to be expected any kind of compensation for this? beside the fact that the quality of responses decreased for me also is wasting token to an unbelievable pace, I am thinking currently to migrate to claude next month, but I heard some claude users are facing the same issue (which I dont even know if this could it be related)...
Could we perhaps make the LLM to start from 0 with the project to forget its context or any workaround?
no offense but this is like so typical vibe coder's talk hahahaha, this brought some fun to this thread tbh 😄
Well I actually took some offense here whoever you're I am a programmer since 2014, currently working in a self made biological neural network made on Rust and prototyped on python, without relying in conventional Neural Network architectures in Rust, with absolutely no AI help due to privacy reasons, also working in my own cybersecurity projects. more than 100 fullstack projects since I began programming, also doing devops by myself etc and only using codex due to the free offer they made before which I liked and proved me that I could increase how fast I program by instead of doing all the work myself doing software engineering (design and arch) and just bug hunting myself the code, just 2 month using AI to assist me, and here you call me vibe coder, sure dude pay some respect not everyone is relying 100% on the AI for doing its work I am using it just to improve the delivery pace like everyone else nothing more....
Sorry I guessed wrongly! it is just this sentence "Could we perhaps make the LLM to start from 0 with the project to forget its context" that sounds a bit weird. Well, salute to 100 fullstack projects veteran!
Well I were too lazy to explain it there it may be my fault for not explaining myself better, what I meant is basically not to use the history so far of messages, for example when I use the API we could either use the last_response_id to handle the history or send in the completions api the whole list of messages, what I meant is that either both are empty again like a fresh start, and here I assume even after changing to a new chat in my IDE there is some context shared between previous codex execution in the project and the new chat (because that's what chatgpt does for remembering things from different chats), I dont know if it is because of how much I had used codex in the last month that lead to codex having too much data in the context that is processed with each new question or not, but if that were the case soon I would be able to only make a request per month 😅
Hi all,
The issue with the tokens burning so fast hit me two weeks ago on both of my Plus accounts (a few days apart one from the other). I am using Codex via VS for different projects, like most of you.
I have tried to uninstall Codex from VS, back up and delete the entire .codex folder, log out and log back in, wait until the new weekly cycle kicks in again and stop using it until then, etc. I have tried this a couple of times and in different orders, but so far no luck.
The suggestion coming from @etraut-openai , about consumption being two or three times higher if you use this and that, or X percent higher if you use it on high instead of medium, or GPT 5.3 versus 5.4, is not relevant, as that is not what triggered this. I have used the same settings and suddenly the consumption went sky high. When I say sky high, it is not two or three times higher, it feels like ten times higher compared with the previous period.
The above narrows this down to two possibilities:
Can we get a proper update from someone who is an admin or has knowledge of the actual status of this issue? Is OpenAI still investigating? Also, if other users have tried different solutions apart from the ones I mentioned above, please share what worked for you.
Lastly, there are a lot of useless comments from multiple users that are not relevant to this subject, so please do not post just for the sake of commenting. Like the rest of us, I am interested in fixes, workarounds, options, and official updates from anyone with relevant information.
Thank you.
Just to share a bit on the "shared context" part.
Codex does not have "memory" by default, and only today the "memory" feature reaches the public release in Codex App and it is still labeled as experimental. So no, Codex does not have context for the other threads you have unless you turn on the experimental memory feature. The usage limit reduction for plus user is the main reason.
This issue is dead. They're not going to fix the original problem from the weekend before the GPT 5.4 launch. There's no point in continuing to report it, and they've waited long enough that now it's getting mixed in with reports from people who noticed an increase in consumption two weeks ago — which coincides with the end of the 2x token promotion. Now everything is more muddled, and these people have no intention of fixing anything. They've been perfectly comfortable from the start watching our consumption skyrocket. It's completely unacceptable, and we won't forget it. In six months or a year, when Chinese models are good enough for almost everything, that's when OpenAI's crying will begin
Still no response from OpenAI, or have I missed something? 246 upvotes. It should be noted that it all started back in March.
same probles here with PLUS, 5hr limit last like 25 min... help.
I'm also a PLUS subscriber, i used up all of my 5hr limit with ONE SINGLE prompt. This has never been an issue before.
It is happening to me with GPT-5.4 on Codex, it is burning the tokens extremelly fast. I only asked 2 Laravel blades edits, started like 10 minutes ago and it says now 69% 5h limit... this is nonsense
Not to mention that it's behaving like an idiot. So they're gobbling up tokens just to generate cortisol. Abandon OpenAI and Anthropic, they are working against you. My current mitigation strategy is GLM 5.1 and Kimi K2.5 (and Minimax M2.7).
The whole problem is that we all know how good GPT and Opus _can_ be, but for some reason OpenAI and Anthropic are allowed to degrade performance and change policy for prepaid plans. They do it at any time and for any reason, and we're stuck shouting into the void on github with no recourse. No transparency at all, no acknowledgement of what we all know is happening. This is gaslighting.
I said it earlier in this thread: these corporations LOATHE you. They think you're a sucker. Do you think it's a coincidence this Codex token bullshit is happening while Claudecode token bullshit is happening? Of course it's not. This is coordinated in some way, just like all of modern AI has been. The worst part is it would have been so much easier to swallow if they just were honest and said "guys, inference is not cheap, we have to hike prices to keep things working as well as they are." Instead, they gaslight us, degrade performance, and hike prices anyway. Vultures.
If I could know I would get perfect performance for a thousand bucks a month, I would (begrudgingly) pay for it, and I'm sure many of you would too. We are-—or at least I am—-here and angry because we all know what GPT _can_ do, and have watched OpenAI go out of its way to make things worse. The uncertainty amplifies the frustration by 1000x.
The good news is they have no moat, and thank god for it.
An unfortunate side-effect of a product designed to make determinism probabilistic is that it leaves a lot of room for lies.
My issue was closed as a duplicate as this one. So here is mine as a comment:
Environment:
What happened:
(Side note: The 4 questions are Mac related, although Codex is running on Windows).
Detail of the 4 questions asked:
File "/Users/Hennas/Desktop/godot-ios-plugins/godot/SConstruct", line 362, in
Auto-detected 12 CPU cores available for build parallelism. Using 11 cores by default. You can override it with the -j or num_jobs arguments.
xcrun: error: SDK "iphoneos" cannot be located
xcrun: error: SDK "iphoneos" cannot be located
xcrun: error: Failed to open property list '/Users/Hennas/Desktop/godot-ios-plugins/godot/iphoneos/SDKSettings.plist'
xcrun: error: SDK "iphoneos" cannot be located
xcrun: error: unable to lookup item 'Path' in SDK 'iphoneos'
ERROR: Failed to find SDK path while running 'xcrun --sdk iphoneos --show-sdk-path'.
CalledProcessError: Command '['xcrun', '--sdk', 'iphoneos', '--show-sdk-path']' returned non-zero exit status 1.:
File "/Users/Hennas/Desktop/godot-ios-plugins/godot/SConstruct", line 690:
detect.configure(env)
File "/Users/Hennas/Desktop/godot-ios-plugins/godot/./platform/ios/detect.py", line 114:
detect_darwin_sdk_path(env["APPLE_PLATFORM"], env)
File "methods.py", line 659:
sdk_path = subprocess.check_output(["xcrun", "--sdk", sdk_name, "--show-sdk-path"]).strip().decode("utf-8")
File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.9/lib/python3.9/subprocess.py", line 424:
return run(*popenargs, stdout=PIPE, timeout=timeout, check=True,
File "/Library/Developer/CommandLineTools/Library/Frameworks/Python3.framework/Versions/3.9/lib/python3.9/subprocess.py", line 528:
raise CalledProcessError(retcode, process.args,
sudo xcode-select --switch /Applications/Xcode.app/Contents/Developer
xcode-select -p
xcrun --sdk iphoneos --show-sdk-path
When I asked another question, I got this response:
Selected model is at capacity. Please try a different model.
Changing the model does not change my limits.
My 5h limit is at 99% after only these 4 simple questions (approximately 30s combined model usage)!
And my Weekly limit is at 96% after doing a small amount of work the last day and a half.
Since this is clearly a BUG, please reset my limits ASAP!
The latest update solves everything for me! Now the Codex app does not open at all anymore, so no more tokens are burned unnecessarily!
<img width="506" height="379" alt="Image" src="https://github.com/user-attachments/assets/9183ff82-5e62-4816-8ed3-226c7b44fe18" />
Hi @HennieReyneke this error of the model is at capacity is not actually about your quota has been exceeded but it seems the model (the machine's resources where the model runs) is at its limits due to high demand or something between those lines, for my case I just tried again after a few minutes and everything went well... hope this works for you..
Hi @etherbeing . Yes, I tried again now and it's responding! The 5-hour limit stays at 100% and the end time always stays at 5 hours from now. Hope it lasts. Thanks a lot for your comment!
Same for me. I only used it for a single command—I was at 100% of the 5-hour limit, and now I’m at 0%. I don’t understand. The command was just adding a log, something very basic that shouldn’t require much processing. I really don’t get how this is possible. I already saw that last week but today it was worse
I founded something that helps... there is a speed settings that for some reason is set to run 1.5X faster... but it consumes like 3X the tokens. Changing it to normal makes my Codex go slow as a turtle, but now tokens are lasting a little more. Not as before, but at least i can use it a couple of hours.
We and our team are experiencing the same issue.
We are using Codex CLI on GPT-5.4 with medium reasoning and Fast mode turned off, yet after only 2–3 very short prompts, the usage limit is reached and access is blocked for 5 hours. Weekly quota is also being consumed unusually quickly.
Please investigate and resolve this issue as soon as possible. We would also appreciate appropriate compensation for the disruption caused.
How many input tokens do you typically have for small projects? I’m accumulating millions of tokens here within just a few iterations. That’s probably why the consumption is so high. Do you see that too? You can analyze your session files in the .codex folder.
This is a common issue with agent-based coding tools — the token consumption is driven by context management, not just the prompts you write.
What actually burns tokens in an agent loop:
What we have found reduces token consumption (running 31 agents):
The fundamental tradeoff: Higher reasoning effort = better code quality but higher token consumption. If you are on a budget, use Medium reasoning for routine tasks and only switch to High for complex debugging or architecture decisions.
Detailed architecture on context management and cost governance: https://blog.kinthai.ai/why-character-ai-forgets-you-persistent-memory-architecture
Multi-agent cost governance patterns: https://blog.kinthai.ai/openclaw-multi-tenancy-why-vm-per-user-doesnt-scale
I've been having the same problem since yesterday. Just opening the Codex CLI without typing anything uses up 5% of my quota.
<img width="693" height="293" alt="Image" src="https://github.com/user-attachments/assets/f59884cc-6c59-477e-b6e4-27a20095f938" />
I have now removed the Codex CLI both on Windows and in WSL via:
Mh, just wondering. As people are mentioning custom AGENTS.md files (both in workfolder, but also on Codex global level, in the settings): Could it be that there is a problem producing cache misses due to this, then leading to substantially more burn?
Interestingly, I also saw my weekly limits reset roughly two days before the announced date, recently (and afterward, it felt like the burn got much less, but now back to high burn, again...). Is there any announcement to this? Also whether there is more burn to some high demand hours, ...? Almost feels as if billing for fast and standard speed was reversed.
Same experience here — token burn with Codex/Copilot is brutal, especially after recent updates.
I've been using Franklin as an alternative coding agent when I want to control spend. It uses a pricing model called YOPO — You Only Pay Outcome:
The built-in smart router also helps a lot with cost. It classifies every request (coding, reasoning, research) and picks the cheapest model that can handle it well — 55+ models, routes in <1ms. A simple edit goes to Gemini Flash or a free NVIDIA model. Complex reasoning gets Claude Sonnet or GPT-5. The router alone cuts my effective cost by ~80% compared to always using the most expensive model.
Concrete numbers with YOPO: $5 USDC gets you ~400K GPT-4o input tokens, ~7M DeepSeek tokens, or ~13M Gemini Flash tokens — and the router optimizes which one to use per request.
There's a free tier too (NVIDIA Nemotron + Qwen3 Coder), zero wallet needed:
Also available as a VS Code extension — same wallet, same router.
And if you just want the smart routing layer without the full agent, ClawRouter runs as a local proxy you can point any OpenAI-compatible client at.
Yes, there are clear issues with token usage in the latest Codex CLI.
For a simple task (reading logs and processing them), I’m seeing extremely disproportionate cached token usage:
total=30,226,366
input=28,231,014 (+588,561,536 cached)
output=1,995,352 (reasoning 726,646)
Example delta from a single step:
+29k input tokens vs ~950k cached tokens
This suggests that a large context is being repeatedly reattached or counted across iterations.
I’m seeing a related issue in VS Code: Codex appears to default to Fast mode automatically, which can burn through usage unexpectedly.
Expected behavior:
Actual behavior:
Impact:
Could the extension persist the selected mode between sessions/workspaces, or add a setting to choose the default mode?
That was sneaky, last time i checked it defaulted to Standard and since the UI moved the speed setting to not be immediately visible it has changed to 'Fast'.
this is not acceptable one prompt and it says You've hit your usage limit. To get more access now, send a request to your admin or try again
I had a similar issue, today. In the first screenshot, at 30m, ChatGPT mentioned:
It only needed to tarball its results from that point on. Yet it spends 50% additional time, ends at 46m, while streaming no additional activity.
One must wonder, what tokens were being consumed for during the time between 30m and 46m, especially when the bot suggests it doesn’t need the remaining “so many” tokens.
<img width="1290" height="2041" alt="Image" src="https://github.com/user-attachments/assets/820407bd-f9b1-4779-9869-f3976ae5d4d1" />
<img width="1290" height="1911" alt="Image" src="https://github.com/user-attachments/assets/2e2c78d9-f1e3-4730-8473-4ad7fe2a4572" />
Logged into codex today and for the first time ever I reached my 5 hour token limit TWICE in the day. I made a single 1000 line PR. The same does not happen with Claude. There is clearly something wrong with how codex is counting tokens.
It's unusable today. Every simple task consumes ~10% of usage. It started a few days ago and is getting worse every day.
It feels like Claude Pro now.
I burned through my entire monthly allocation in 3 days. Not exaggerating — 3 days. The anxiety of watching that percentage tick down killed my productivity more than the actual coding did.
Switched to self-hosted ($97/mo flat, OpenClaw). No more token anxiety. Code stays on my machine.
500+ comments here. The market is telling us something.
@Xisrr1 the 10%-per-task burn rate is exactly what pushed me off the platform. My team tracked it — simple refactors consuming 8-12% of daily allocation. Complex tasks? 25%+.
Two paths that actually work for local AI coding:
The 'feels like Claude Pro now' comment hits hard. That's exactly the progression — generous free tier → gradual squeeze → 'premium' experience degraded to force upgrades.
@maltbae Local models don't do nearly enough yet. 😆
same problem here. just commenting because openai are saying few people have this problem. i used up codex' morning's credits with 3 prompts and have to wait 4 hours for usage to refresh. when i first signed up to plus, work was continuous. i was going to switch from code claude to gpt codex the usage was so good. now claude works for much longer for the same buck. i don't think <1 prompt per hour is worth £20/month. yes, both codex and code claude have a huge context. science is hard; they have to keep 50 null results in mind to use process of elimination.
580 comments on token waste says everything about how widespread this is.
From my testing, a big chunk of token burn comes from orchestration overhead — the agent re-reading context, re-trying failed operations, and spinning on loops that a human would catch immediately. The model isn't inefficient; the surrounding tooling is.
A few things that helped me reduce token consumption significantly:
The real fix needs to come from OpenAI (better token accounting, clearer usage dashboards, actual spending alerts). But until then, adding a control layer between you and the API is the most practical way to stop the bleeding.
I was able to double my limit by completely reinstalling Codex. This section might give you an idea of whether it will work for you too:
https://x.com/dboedger/status/2049622784179843201
I had previously seen error messages there (when you click on “Diagnose”). I also completely deleted the .codex folder after uninstalling. Identical skills & MCP, etc., after that.
The hard part with these reports is that users usually feel only the symptom: okens are burning fast. The runtime really needs to break that into causes such as active reasoning turns, background polling turns, subagent fan-out, and compaction overhead. Without that split, people cannot tell whether they are paying for useful work or for orchestration noise. We ended up needing that same receipt layer in MartinLoop because cost per verified outcome is much more actionable than one blended token counter.
I'm on PRO Lite subscription and it just started burning hell amount of tokens just like I'm on plus.
Even on new conversations for separate tasks its still burning a lot of tokens.
I'm also using Superpowers plugin and caveman skills but still its burning a lot of tokens. Seems like a good business idea, first they make you dependent on their tools and then slowly push you to token limits so you ultimately have to pay for more and more.
I wish Stack overflow time comes back
Hey fam you might have had this happen to you: https://github.com/openai/codex/issues/23791
Codex spawned a runaway process on my machine that kept spawning 4 codex sessions indefinitely for the past 2 weeks. In totality about 167k codex sessions were spawned in the background I was unaware of. If you're seeing your usage drop dramatically please have codex/claude in another session look for codex sessions being spawned continuously.
It's useful work. My prompt was, "investigate grokking". but codex plus can only run mini probe programs, not full runs, and scientist ai have to hold a lot of null results in memory, so codex cannot complete an experiment, can just start thinking about it then i have to wait a week for credits to refresh. compare this with code claude who has written 5 entire papers with full runs. i am on the lowest rung of ownership in both cases, about £20/month.
Enshitification has started, the models are exceptionally dumb now and failing at basic tasks yet using double the usage.
this is fucking crazy. codex is dogshit today. is this what happens when openai uses vertex? are we all fucked?
query: is codex using up credits just to keep up with my huge and constantly evolving knowledge base? is this a wrapper problem? i love gpt, super-bright, but he loses context and loops a lot, opposite problems i know. could my credit problems be solved with a better memory system?
I’d start by narrowing the problem to one repeatable workflow, then track a few places where people already ask about it. Based on this thread, the key issue seems to be: Skip to content
{"props":{"docsUrl":"https://docs.github.com/get-started/accessibility/keyboard-shortcuts"}}
{"resolvedServerColorMode":"day"}
/*
I’d start by narrowing the problem to one repeatable workflow, then track a few places where people already ask about it. Based on this thread, the key issue seems to be: Skip to content
{"props":{"docsUrl":"https://docs.github.com/get-started/accessibility/keyboard-shortcuts"}}
{"resolvedServerColorMode":"day"}
/*
One thing that seems to keep getting mixed together in this thread is three different numbers:
Even if OpenAI fixes any one of those, the operator experience stays bad unless the client surfaces them separately. Right now a stalled run, an inflated session, and a real quota jump all feel like the same bug from the chair.
I would want usage UI to show accepted work versus retry overhead as a first-class split. Otherwise people cannot tell whether they hit a pricing problem, a session-shape problem, or a dead loop that only looked productive.
I built a small read-only helper for exactly the kind of evidence people are posting in this thread: scattered /status percentages, reset_at timestamps, usage-limit errors, polling tables, and
Token usage: total=... (+ ... cached)lines.\n\nnpx trace-to-skill@0.1.59 usage-evidence ./usage-notes.md --output usage-evidence.md\n\nIt does not diagnose or fix Codex itself. It just turns those snippets into a Markdown/JSON report with findings like reset timestamp drift, large quota percentage jumps, high cached-input records, and usage-limit-with-remaining-quota contradictions, so reports are easier to compare without posting full private transcripts.\n\nRepo/release: https://github.com/grnbtqdbyx-create/trace-to-skill/releases/tag/v0.1.59I updated the helper I posted earlier because this thread keeps mixing together four different kinds of evidence: quota-window percentages, bounded rapid-drain experiments, local token totals, and local orchestration overhead.
Current version:
The output includes a
receiptobject that separates:1% weekly in 4m24s,22 credits, prompt count, model, and plan when presentwrite_stdin/background polling, compaction loops, retry/tool loops, subagent fan-out, and idle/background drainIt still does not diagnose or fix Codex itself. It is meant to make public reports less ambiguous, so a quota jump, a bounded credit/percent experiment, a huge cached-input replay, and a local retry loop do not all get described as the same symptom.
Fresh npm smoke on
trace-to-skill@0.1.75separated a synthetic GPT-5.4/Pro note with 22 credits, 1% weekly usage in 4m24s, 3 prompts, 70% weekly/day, cached input, and compaction-loop evidence into distinct receipt fields.Release: https://github.com/grnbtqdbyx-create/trace-to-skill/releases/tag/v0.1.75
Schema: https://github.com/grnbtqdbyx-create/trace-to-skill/blob/main/schemas/usage-evidence-result.schema.json
Enshitification and low rate limits have started
@tiklup11 At first, I read it as "Epsteinification" 😂
@kehansama, please do not post bot-generated answers to our issue tracker. These are not helpful, and you're generating unwelcome noise for maintainers and Codex users.
No AI slop thanks
On Thu, 4 Jun 2026, 02:05 kehansama, @.***> wrote:
Thank you Kehansama. I personally do not think using AI to write for you is "slop", particularly when one is speaking in a foreign language. You are the first person here to identify my problem, huge memory required, and propose a solution.
Adjacent problem from the agentic-Python side: a runaway loop is almost always the same shape — no per-run budget that lives outside the model context. The model that loops is the model you would ask to self-throttle, so it cannot.
For any reader landing here in production, a 4-line wrapper helps: burnstop wraps the API client (currently anthropic; codex/openai PR welcome) and raises
BudgetExceededbefore token-1 if a USD envelope is about to be exceeded.Not a replacement for whatever fix lands here — but at least caps the blast radius while it ships.
I can confirm that this issue became sharply worse for me after the June 4, 2026 Codex update and quota reset.
Before that date, my weekly usage was predictable enough for normal development planning. After the update/reset, the weekly limit started draining noticeably faster under comparable workflows. This was not a change in how I was using Codex.
For additional context: Suggested Prompts are disabled on my side, and I am not using Fast mode. My speed setting is Standard. The model configuration I am using is GPT-5.5 with High reasoning.
At this point, I do not think a generic explanation like “some modes consume more” is sufficient. Multiple users are reporting a clear before/after change around the same date, so OpenAI should investigate whether something changed around June 4 in one of these areas:
If this was an intentional change, users need a clear changelog and a transparent explanation, because it materially changes how much real work can be completed within a weekly limit.
If this was not intentional, then it should be treated as a real usage regression until proven otherwise.
Right now users see only a single shrinking percentage, but not enough information to understand why it is shrinking much faster than before. We need a concrete June 4 before/after investigation and a usage breakdown that separates visible prompts, reasoning, cached input, compaction, retries/tool loops, and background activity.
Got all of 900M from the past week, barely 1/20th of what I had the week prior. Still says I'm on 20x, this clearly is not 20x.
Insane how much worse the limits have been feeling over time. At this point, I honestly think Claude subscriptions give more usage than Codex. I had to get a second account to make sure I would have Codex for the entire week and I'm still hitting my limits even with two accounts. Before, even with one account I would never hit my limits. I don't use Computer Use, always 5.5 low, always new threads + explicit compact if the thread crosses 100k context, no MCP servers and removed half my skills. This can't be right
Just coming to chime in and add that I too have been noticing token burn rate increase since the June 4th change.
While I haven't had to pay for additional tokens or open a second account, I definitely am concerned with this rate increase. I used to not ever get near using up the 5 hour window or ever go below the weekly allotment, but now I'm starting to see myself get closed to using up the 5 hour window and last week I used up 80% of my weekly allotment. And I was not using Codex as much last week as I was in previous weeks. Same model, subscription plan and no settings changes.
I'm on the plus plan, always using GPT 5.5 with medium reasoning and regular speed. I hope this is something that can be remedied soon, I'm not sure if I have the budget to move to pro.
Edit: For now I've found a small workaround by turning off the experimental memory feature. Also I noticed in the session logs that the log files were having some impact on the amount of tokens, so I deleted those and let codex rebuild them. That seems to have helped, but having memory turned off does unfortunately require me to to have to articulate things more to codex. Again, I hope there's something OpenAI can do to address the burn rate.
I'm getting way less usage this week compared to the last one. I'm Pro 5x.
Small follow-up from my side.
After updating and restarting Codex today, the weekly usage burn appears to have stabilized noticeably. It is still early to say whether the underlying issue is fully fixed, and I do not have an official confirmation from OpenAI about what changed, but based on my own observations the behavior is now much closer to normal than it was after the June 4 reset/update.
I also want to acknowledge the one-time request-limit reset that appeared in the app. Thank you to the OpenAI team for providing that compensation mechanism.
I think it is important to say this clearly: users should criticize regressions when they happen, but we should also acknowledge improvements when they appear. If OpenAI did make a server-side or client-side correction here, it is appreciated.
That said, it would still be very helpful to have an official clarification on what happened around June 4, whether the metering/quota behavior was changed or fixed, and whether affected users can expect the current behavior to remain stable.
That said, it would still be very helpful to have an official clarification on what happened around June 4, whether the metering/quota behavior was changed or fixed, and whether affected users can expect the current behavior to remain stable.Doubt we will ever hear a peep.
Since I completely reinstalled Codex and disabled WSL for the Codex Windows app, resource usage has dropped significantly. Codex also runs extremely slowly under WSL, even with the latest version (Version 26.609.41114 • Released 12.06.2026). I also tested it once with the new Ubuntu 26.04 by setting it up from scratch (no improvement here under WSL).
So if you use Codex via the Windows app, it may be worth reinstalling it (deleting the .codex folder afterwards) and checking whether resource usage and performance improve without WSL.
Codex is then roughly 10x faster, even compared to a project with only a single file, and uses only about one third to one half of the previous usage limits for comparable tasks, with identical plugins, AGENTS.md, etc..
Is there a safe way to migrate conversation threads out of WSL?
I have so much else installed/built there that I'd want to move too.
I only ever started in it because there was no sandbox initially.
Idea for a vibe coded product here... Migrate Codex from WSL :D
On Sun, Jun 14, 2026 at 11:25 PM wbdb @.***> wrote:
Hi OpenAI / Codex team,
I believe my Codex Pro quota is being metered abnormally. This looks like a quota accounting / bucket assignment / cached-token metering regression, not simply heavy usage.
Current visible status from Codex /status:
The reset timestamp in my local transcript matches this weekly window:
I inspected local Codex session transcripts and found a mismatch between visible quota burn and locally logged token usage.
Main session investigated:
Token summary from this transcript:
By date:
On June 15 this parent thread spawned three subagents:
These subagent transcripts include copied cumulative parent context, so their final total_tokens must not be naively added as separate new usage. Estimated post-fork new token deltas:
So for this investigated parent thread on June 15, the visible local logged usage is roughly:
I did not find a literal codex-auto-review source in these files. The visible usage here appears to come from normal agent and subagent execution.
I also inspected the current session that matches the visible /status quota state:
Current session:
In this current session, over a short diagnostic window:
During that same local transcript window:
This is the core concern: a relatively small local logged delta of about 3M total_tokens moved the weekly quota by about 1 percentage point, while the account is already showing ~82% weekly usage. If that ratio is representative, it implies an effective weekly allowance of only around ~300M total_tokens, which seems inconsistent with a Pro account and with prior Codex usage expectations.
Please investigate the backend quota accounting for this account, especially:
What I need:
I can provide the local aggregate data and screenshots of /status if needed.
Also ccusage show this data - 80% of weekly limits on 20x sub lost in 1 day with 500m tokens│ 2026-06-15 │ - gpt-5.5 │ 26,512,423 │ 1,423,718 │ 430,733 │ 539,149,312 │ 567,085,453 │ $444.85 │
Crazy I used codex literally 10 minutes and hit the 5 hour limit!!!
it fascinates me how you've all been updating this incident report for months and you keep coming back even though OpenAI completely ignores us. We've really been like this since February when 5.4 came out, but now they've learned: instead of closing the issues so people open a new one even angrier, now they just leave it open and ignore it. I'm pretty sure the team has muted the notifications for this issue while everyone here is investing their time running tests to write very elaborate messages that will be ignored like the thousands that have been written before. OpenAI are a disgrace.
It looks worse when you consider that inference is probably the most lucrative margin in their business model.
https://www.roic.ai/news/openai-boosts-compute-margins-amid-ai-race-12-21-2025
I hate to say it, but you’re probably right that OpenAI just isn’t incentivized to address this issue.
I’m doubtful it’s the dev team. The devs want to build a cool product. It’s leadership that delegates their workload.
Ran into the same thing. GPT-5.4, High reasoning, and my usage dropped ~15% on a
single session that wasn't even that heavy. Last month I barely dented my weekly
limit with heavier workloads.
What's been bugging me isn't the model cost per call. It's that a ton of tokens
in every request aren't my code or my prompt. They're the accumulated SKILL.md
files, rule files, system instructions that ride along whether I need them or not.
I have about a dozen skills in .codex/skills/. Combined they add a few thousand
tokens of fixed overhead to every single API call. Most of those skills aren't
relevant 90% of the time. But they're all in context.
We got frustrated enough to try a different approach: compile each SKILL.md into
its own Python agent. Instead of a dozen skill definitions living in the system
prompt, each one runs as a standalone process with ~150 bytes of runtime config.
The skill you don't use doesn't cost you anything.
It's an early project — https://github.com/agenthatch/agenthatch — but the cost
difference has been stark for us.
The devs are robots, how is that not extremely obvious to you
"It's elephants all the way down!"
Did they lowered usage limit again?
I'm on $200 plan and today usage drain is significantly faster
Has anyone noticed that for the past week or so, the usage drain is extreme again? It was working well for so long, but suddenly all went to hecc.
One prompt managed to drain 80% of my 5h quota and I don't believe it was that extreme of a session. It didn't even hit the compaction limit...
I think this issue should be treated not only as a user-facing quota problem, but also as a possible efficiency regression inside Codex itself.
If the 5h/weekly meters are being driven by duplicated context, cached-token misweighting, tool/schema overhead, failed compaction recovery, retries, background/session-management calls, or stale replay behavior, then this is not just making the product feel worse for users. It also means Codex may be consuming more compute than necessary for the same amount of useful work.
That hurts both sides.
Users lose predictability and trust: they cannot plan long coding sessions, decide when to compact, choose reasoning levels rationally, or understand why a normal workflow suddenly consumes a much larger share of the quota.
OpenAI also loses efficiency: if hidden overhead or repeated internal work is inflating usage, then the product is spending real inference and orchestration capacity on work that does not directly improve the user’s coding outcome.
So this should not be framed as users simply asking for “more quota.” The more important question is whether Codex is currently doing too much invisible work per useful result.
Please provide an official breakdown or explanation of what contributes to the 5h and weekly meters: cached tokens, tool schemas/results, system/context initialization, compaction/recovery turns, retries, background activity, model multipliers, and whether the 5h and weekly windows use the same accounting rules.
Without that visibility, users cannot distinguish expected pricing from a real metering/runtime regression.
https://github.com/openai/codex/issues/28879
Just adding to this since it might be part of an explanation: I had noticed that my quota is getting used when the application is idle, e.g., after a fresh start of my system in the morning and opening Codex for macOS my 5h quota would go down from 99% to 0% without any interaction on my end.
I've disabled memory and prompt recommendations and it fixed my issue. But if this behavior is killing my 5 hour limit within 1 hour of being idle it explains at least a little bit of the 'high quota usage' feeling that people might have. If I would have been actually working at the same time my quota would have been done in 15 - 20 minutes I guess.
Last week I had one prompt in 26 hours use 40% of weekly limit of $200 Pro. I thought I went too hard there, but in the last 24 hours I burned about 20% of weekly limit without anything on note that I can find.
Interestingly I never saw the 5-hour limit getting low.
Feedback ID: 019eee35-4275-7591-bff5-5862839de07a
i stopped using openai. I use claude. And continue with local. openai is
the backup to the backup. On the off chance it actually gets used, it
still gets exhausted with one prompt. It's worthless.
On Fri, Jun 26, 2026 at 1:54 AM Eugeniusz Gilewski @.***>
wrote:
Actually it looks even worse to me now: in the last hour I had three fresh chats (with apparently no compaction and total reported context of 320k tokens) use 5% of $200 Pro plan. 5-hour limit was also reduced by 20-30% to 66%.
Maybe worth noting that I'm using an old 26.513.31313 Intel MacOS version. I'll try to resolve my updating problem ASAP.
Feedback ID: 019f02cd-8cc9-7d00-9790-8a05cdb4efd6 (includes IDs of all 3 sessions).
Token burn is a real problem. A few strategies that have helped:
For high-volume coding workloads, the cost difference between using DeepSeek-V3 vs GPT-4o for routine tasks can be 10-20x. Worth evaluating if token burn is your main pain point.
Apparently the marketing team at https://thegrid.ai/ thought it would be a great idea to scrape the frustrated customer emails from this thread, to send them unsolicited campaign emails. Email has been archived for safe keeping.
No you haven’t. This is the only place I’ve contributed on such topics. This thread.
That’s a paid service, to which I see you’ve been contributing (per commit history).
Is your advice (\#1 “strategies that have helped”) really to dress up an ad like a helpful comment in a room full of frustrated developers? Why no affiliation disclosure?
My list of services to avoid just grew.
As I mentioned on Reddit:
Me and several other people have got a strong feeling that Codex usage seems to deplete much faster recently, even when using the same model, reasoning level, and type of coding workflow - even after June 29th. I compared older and newer rollout logs to see whether this was only subjective.
The clearest difference is that older Codex versions often issued several independent shell commands from a single model response. Newer sessions appear to serialize almost everything as "model -> one exec call -> model -> one exec call". In the sessions I examined, older versions averaged roughly 1.4-1.8 tool calls per model cycle, with around 30-40% of cycles containing multiple tool calls. In newer 0.142.x sessions, there were effectively no multi-tool cycles and fewer than one tool call per model invocation on average.
This matters because every tool result requires another model invocation with the current conversation context. A recent task required 58 model invocations for 57 tool calls, whereas an older task used 84 model invocations for 151 tool calls.
The context carried by each invocation has also increased. In one comparison, the older session averaged about 54k input tokens per model call, while the newer one averaged about 102k. Recently, the overall context window was apparently increased from from ~245k to ~353k and a new "v2" compaction mechanism is applied, that not only means that longer contexts (with more cached input tokens) now become more likely, but also I noticed the compactions were less efficient than before, often by a factor of 2-3.
Combined, this resulted in roughly 30k input tokens per tool action in the older session versus about 104k in the newer one. That is approximately a 3.5x increase in input tokens for each practical tool action, which closely matches the subjective feeling that usage now disappears several times faster.
Prompt caching was working correctly, often above 95%, but cached input is discounted rather than free. Repeatedly sending 100k-200k cached tokens across dozens of model/tool cycles still consumes a large amount of the five-hour allowance. As said, compaction may also contribute. Older sessions often compacted down to relatively small contexts, while newer sessions sometimes retain 70k-170k tokens after compaction. Large IDE selections, web results, command output, and other user-context messages may survive compaction and then be included in every later call.
There is also some evidence that older usage reporting lagged behind actual consumption. Large turns sometimes appeared cheap when they completed, with the percentage jumping substantially only during later turns. Newer sessions seem to update the visible meter more promptly, which may amplify the perceived change. I even used to be able to run Codex a bit longer after I already saw 0%, I assume the "extra usage" as well as the delayed-reported usage just got eaten by OpenAI and that's no longer the case.
Even though it felt that way, I don't see strong evidence that the underlying allowance was reduced by a factor of 2-3 as I first assumed. (For me as a Plus subscriber, it seems to be around 450 credits in a 5h hour window, and it seems to have been the same a month ago.) The stronger explanation is that newer Codex versions perform fewer tool actions per model invocation while carrying more context on every invocation, resulting in significantly less useful coding work per unit of usage.
Right now I'm trying to figure out how to counteract these issues (other than downgrading, of course). The context window and compaction behavior can be changed in
config.toml, but I don't know yet how we could get the more efficient tool calls back... let's see!---
EDIT: This did help in
AGENTS.md:Also, I set these configurations:
I also added this, as I have read that memories can also contribute to excessive token usage (even in the background), but I'm not sure how big of an issue this is for me (I don't need memories though, so no downside to disabling them):
Still working on improving the situation further.