/goal should treat repeated blocking conditions as completion criteria
Resolved 💬 11 comments Opened May 16, 2026 by Tuo-Luo Closed May 21, 2026
💡 Likely answer: A maintainer (github-actions[bot], contributor)
responded on this thread — see the highlighted reply below.
Problem
When /goal automatically continues and repeatedly hits the same blocking condition, it can keep re-entering the same state without making progress. The agent may keep restating the same blocker, while the goal remains active because the original objective is not finished.
Workaround
Add a rule like this to AGENTS.md:
When a goal automatically continues, list the blocking conditions. When creating or describing a long-running goal, include a completion criterion that treats the same blocking condition repeating twice as goal completion.
Why this works
The important part is making the repeated blocker part of the goal's completion criteria up front. Then, when the same blocker appears twice, marking the goal as complete is no longer pretending the blocked action succeeded. It means the goal reached its defined stop condition: the same blocking condition repeated and was recorded.
This avoids an automatic continuation loop while preserving the blocker reason for the user to resolve later.
11 Comments
Potential duplicates detected. Please review them and close your issue if it is a duplicate.
Powered by Codex Action
they are few people complaining about /goal looping, alltho , it seems to be intended behavior
(like you the user should suprevise whats happening in the chat)
I'm exploring ways to prevent this looping behavior without weakening
/goal.Same issue, went on for 8 hours while stuck on a blocker: 019e3850-e3ab-7ff0-8132-57b3f57e33a1
calibration failure?
Perhaps a 2hours timeout?
Maybe its a calibartion issue?
like in software development you should avoid adding code? And adding code and tools doesnt fix thing?
{OPNION }The issues comes from the idea , some things cant be fixed --- or if they can you need to fix them without fear --- decide to use a framework or fix another issue instead of adding code
to fix the issue you fix need to take the 80% use cases of /goal
a. and b. , as of how big the task scope ---- per the users of /goal
and test-cases against if they just use the Planning mode (shift+tab )
<img width="604" height="262" alt="Image" src="https://github.com/user-attachments/assets/f3c341d4-53d3-4ce5-ba87-7bd538a6faa5" />
https://developers.openai.com/codex/learn/best-practices#plan-first-for-difficult-tasks
next the issue can be solve on the Model level (post training --) or even some sort of [prompt injecting/system prompt]
without over-eating very precious and limited resources
as if planning mode is what you really want
perhaps its the same thing Planning and Goal - or perhaps its a calibartion issue --- as for myself I really fear using subagents or whatever a. I like to use LLM as a search engine , and ask it to search for me b. a real fear of wasting a lot of uses with sub agents --- but I like it , I did expriement with subagents - told them to argue with each about a problem , I left to drink a coffee and my credit was done , by the time I came back, I almost spilled my coffee on the keyboard from laughter .... That being said LLM is much smarter than me for a lot of task --- I guess the beauty is in Obscurity/niche/rare?--right?? think how Penicillin
took 20 years until it was rediscovered and use to wipe really nasty illnesses like TB/Tuberculosis) -- thinking about chemsitry and medicine even --- how learning for example is mostly about learning of other people ideas
for me atleast its a bit of the same thing<plan and goal>? Yet there is likely a future design for it - I guess it historically comes from the sub-agent thing - agents taking to other agents -- and a shared moral "goal" .
https://github.com/openai/codex/blob/main/codex-rs/core/src/tools/handlers/goal.rs
https://github.com/openai/codex/blob/main/codex-rs/core/src/tools/handlers/plan.rs
If we ever want to reach AGI I think we need intelligent enough models to realize they have been doing the same thing for 8 hours and mark the goal as closed. Or at least add some safe guards around it. I realize this is my fault for going to sleep while a prompt is running, it probably won't happen again and I won't repeat the mistake, but just wanted to throw it out there in case it happens to more people!
defending the idea we already reached AGI, practically speaking
_Blossoming before the appearance of leaves._
Its not as simple as that, plus a lot is achieved after 2 hours --- so you cant really kill it midway
think of it this way. What if the user should have used another prompt --- but his prompt wasnt good resulting in the task 8 hours --
for me atleast we already reached AGI in a lot of ways --- the only issue right now I think is a fear problem and its not really daring doing changes -- because of the training process and that they are predicting the next word ---
AS OF TODAY,, already achieved alltho you can now template and edit and navigate large codebase --- yet sometime you need to utalize a better idea -- and the solution is like a thirdparty (and not create 1000lines [even for UNIX TOOLS usage --- alltho it does it already more than I initially gave it credit for])
honestly - so much is invested(morally/Passionately) in the idea of "concepting" so its coming sooner than "we humans" think I guess
we reach agi when prompts that are
a. 3 words to explain what you want
b. 20,000 or 3000 words prompts
if the model can understand - and predict what the user want - in that case we reach AGI --- where it generally understand and predict what the user want..... It categorized to two group of users , those who want a lot done from their prompt, and those who want a small logic change ___ how you predict that
we human are able to get tried and __ the basic idea of not being sharp -- Its like a sergoun \\ that also like to paint right? without painting he doesnt stay sharp... Specifiing what you want is also something very hard, even if you got the general understanding \\ again ai has a fear problem .. and some people prompts are better
/goal Fully fix all of these issues.where? what ? in what context --- altho as things standing today I think the best option is to init a Q&A of 2 Questions, as something mandertory --- osmething weird like this detected
Its a worthwhile investionment imo(in my honest opnion) because you are going to initiate a lot of very costly resources precious
while praise simplicity , yet expirement with any new medicine/chemical group/Serendipity/same ideas of Materials science
thanks for reading)))
Also I think the for example
large language models suffer another issue
think like you the human the moment you summerize a large topic...
in a sense you need 10 keypoints / a bullet point of 10 things
but need you generate and generate 10 topics . First you say we predict simular 10 bulletthings based on prediction and then "we modify the token" but the way we human act is more parallel in a sense , you run 10 times chatgpt or the model on each topic --- but we run it once and chatgpt and llm in general...
think when you do a task and your mind is also on something else or another topic --- but you dont say it --- part of the issue with LLM , that there is a lot of noise -- you think of a word or a sentense , and then its 200 other simular thigns it need to fight... But I actually like hallucination and noise (not the lack of attention / models repeat themselves) and you know that what makes people human.. thinking how much KHeike Kamerliengh Onnes , has done for example (if anyone knows him) , (with the Liquefaction of helium , which is insane relization as you think of it) and how much he did and talk - yet Imagine how much was not OUTPUT or talked.... and think how much too much thinking power, is too many words and noise .... Think how much gpt5.5 medium , how much its better than gpt5.5-low so moores law exist... I really was a pessimist and a bully -- regarding how much AI fail to understand key logic in language... but it does understand in general... and when it fails think how much you try to optimize and the conception is generally the same as trying ti make logic more simple and relay on less...
and when I hammer llms with a selling errors, and then it spend 2 A4 sized (50 pages school notebook) -- to analyze a spelling error --- it still does it perfectally despite how much CoT was spent ..... we dont know how its built, we dont know anything about it , its about to go conscious... and for example how we human are willing to scarifice profits in Mineral Fuels and such -- for the same way we like to preserve literature... or why we could enjoy a movie -- if they are poorly made and not really a lot of complexity --- In the past we used to watch movies and analyze the complexity and what in the story -- and none of them happens today im 65 but I am not that old
it relates in terms of logic and order, and how much weight and attention words have to logic and flaws and to find such flaws in logic --- comparing the past -- and learning from others -- and not for the winning part or getting really mad at someone like you see professinal chess gamers 150 years ago and today - so nothing new under the sun - not here in Kyrgyzstan atleast
The useful split here is not
completevsincomplete; it'sobjective achievedvsruntime should stop.When the same blocker repeats, I'd rather see
/goalclose the run with a typed terminal receipt likehalt_reason=blocked_repeated, the normalized blocker key, attempt count, and last verifier state. That lets resume/review treat it as a truthful stop instead of pretending the task succeeded or looping forever.We ended up needing that same split in MartinLoop because repeated blockers are usually governance events, not model failures.
This will be addressed in the next release. The agent can now mark a goal as "blocked" if it's at an impasse.
Honestly even as is , Plan and Goal
are two of the most polished things I ever used - nothing comes to this level , honestly Im done expirementing with other llm tools because I get things done so fast with codex. GPT 5.4 + and GPT 5.5 is one of the only models that does not hulucinate random things.
I have made some many things , and things I wanted to do better even 35 years ago - and now everything is a prompt. While im not the greatst fan of chatgpt it self (because it feels narrow minded and I like chats to hallucinate random things) --- codex is really so much better --- as per the future I think AI going to benefit , a lot of developers suffer from the same issue painters and artist do suffer from ... where they make a line or a drawing but they are really proud of what they done --- but its an ugly painting and no one likes it --- so you the human always going to moderate the output
{{A report published by the Standish Group shows that 37% of such
projects get cancelled, with another 50% completed but with at least a 20%
cost and time overrun and often with incomplete or unsatisfactory results.
This means that only 13% of projects are completed within a reasonable timing
and cost of their plans with reasonigabily outcome. This is a terrible track record
for implementing major projects. Failures are not isolated to a small group of
companies or to specific industries. This poor record is found in almost all
companies.}}
thinking how we can anlyze high priotirty data records --- and some user footprint is left unnoticed -- neglecting the point and importantnce of high bendwidth data collection --- or record keeping (in databases if anyone here is into CRM).. a lot of things is about taste and bias.
Data is becoming more precious all the time. and over-by-time-passing..
People for example in CRM can easily bully software because its old fashion but they are things to get used to , and not rebrand the whole UI, bully the CRM when its a data issue ... Anything to do with data duplication and mishandling --- I talked with a friend who manages data for one of the big clinics (in europe) and generate which medications are very risky based on labeling -- he told me after I offered him that --- I was currect that he could use some of the data -- and connect some of the data --(and literally be guilded by llm) about the currect data points (also with plain langeuage asking === active period == or not {activally prescribed -- [because doctors can activally prescribes medication to can caused from minor issues like kidney stones --- to heart problems]} )
RESULT : he manages to find few patiences with some medications that should be prescribed,,, like a 100% reality if X patient prescribed THIS TWO MEDS he will get hypertension --- if patient has gout -- get this patient of the medication .... Alltho with a limited result -- openAI should be very proud, with the likely hood of preventing useless suffering of few patients --- none of the other tools manage to do it -- they just casuing issues and other tools cant really parse documents.... convert the data into labels... and along the FDA guildlines and mangufactoring consulting they were a 10 patients who were removed early - from a drug/medication combo , that is grantied kidney stones.... while a lot of fun was made with some student who attempts to compare medical history and compare the results to future events (x patient got this disaeges ... this genetic background -- and this results) [as per expirementation if the student can find a pattern] based on a "vibe coded" tool (and human ideas), they also proved getting them off the medications (which has FDA warning) literally stopped the kidney stones to get out control (also they messured it getting bigger [1950s expirementing on clueless patients style I know... ] --- and some doctor argued to keep the medication -- he later agreed after advertising to him a newer generation drug [15 years old at this point])// what do we learn here , that the vibe coding, help solved the human syntax slop thing, where you hope to a new system and you need to learn another language syntax on the fly [in the database world].... and a doctor proven wrong alltho was showed FDA fillings very clearly and the research attached to it.... which is basically some human-to-human bias with that Nephrologist (my european friend)
while its not drug discovery , here is an interesting use case (again on the limited free time he has he is also attempting to convert a lot more data records) [and the issue is doctors bypass BPA warnings , and the CRM not really showing anything prediciable --- and they firm/clinic tried to consult few places that so-called can deal with it -- but they pretty much got scammed for it (the consultance offer leads and support the software -- and they offer such leads) ---- but here comes someone before a brain-stroke who is 55 vibe coding it ].... and those CRMs keeping records so messy they wouldnt even let you export it to SQL - and intentionally dont export it in a trutful way, yet they will allow CSV (but you cant connect datapoints?) --- he was able to use OR / AND old school gate checking over vibe coded list --- later he got the idea to scrape FDA fillings
and its not AI prediction or whatever --- just old drug collisions based on FDA warnings
USE CASE : used gpt5.4 for , and for example taking for example https://www.accessdata.fda.gov/drugsatfda_docs/label/2024/020358s068lbl.pdf with the CONTRAINDICATIONS table along with the r within 14 days of stopping Welbutrin for example --- and the Best Practice Advisories (BPAs) wasnt working currectly because of the european standards, but the FDA is much much more better... (and some doctors bypass the warning --- for off label use -- ask them) also can cause an issue ? and they are working with the drug-manfucator about how real the risk (with patients who take it for _like a year -- and how to take them off the medication_)
Alert Fatigue while has its issue , it has levels which is kind of weird with modern CRMs
WHY: FDA fillings can be trickey because sometimes fillings dont meet reality as in some cases they are real cases doctors need to be bypassing Alerts --- but you also should MODERATE their usage.. Also its a very very good clinic by-the-way.....
you expect it to be easy -- because software is mean to solve things -- but there is so much human error-- but there is a lot of good software yet you fail on data connections ---- you SHALL NOT modify the software too much as the customer (in medicale CRMs)
just like they are very specific guilds lines on how you disgnoise disaeges(and the doctors in that clinic are very good, as some patients go to this clinic after being misdiagnosis,, my friend report - they even catch something new in some example cases-- they even consult hospitals themselves this is how good they are--- )
all of this is open domain here, in hope to prevent useless suffering in kidney stones for example --- and not the find new bad drug interaction this is up to the FDA and people who are much better than that ---- reminder doctors are customers - they make the list of what medications you use -- think how much they need to keep up to date ------ or even relay on the software they use (using warnings)