this is just an example why CODEX Is better, codex would ask, what do you mean by refactour this codebase

Resolved 💬 8 comments Opened Apr 20, 2026 by Mahkhmood9 Closed Apr 20, 2026
💡 Likely answer: A maintainer (github-actions[bot], contributor) responded on this thread — see the highlighted reply below.

What version of the Codex App are you using (From “About Codex” dialog)?

i am using Codex , on a FreeBSD machine

What subscription do you have?

I love codex! it was really bad at 5.2 , but now even 5.4 mini IS ANYTHING I EVER WANTED

What platform is your computer?

_No response_

What issue are you seeing?

<img width="782" height="604" alt="Image" src="https://github.com/user-attachments/assets/1618c4b6-0389-4444-b0d9-a99daf1379bd" />

  1. "refactouring" means , better code practices , more readable functions,,, the user wanted to minify the codebase
  2. this is just an example why CODEX Is better, codex would ask, what do you mean by refactour this codebase
  3. Codex , is able to skim the codebase using rg , and it visits the AST , and crawl the syntax in a sense
  4. codex is able to search a file, within a very very very large code
  5. I given codex 5.4 medium , the task to add SOCKS5 support over the chromium codebase , it did it perfect Cl****de opus 4.7 didnt know how to even compile the code, because it didnt know to visit the "howto.md" and the documentation literally built into the large codebase
  6. Codex is so good it takes 1second to do a large query , it firsts think of the codebase, instead of rushing to code

-

  1. Codex doesnt rush , it ask the user what to do next, like "you asked me to refactour" , this word , could mean few things to other people "john thinks its a shoterning" , "Abdul thinks its legnthening the codebase " "and mukhamd thinks its a better code design"
  1. classic openai W
  2. codex is literally I ever wanted, it does git better than me ... it controls my packages better than me ,,, it does nginx better than me ,, it knows to deal with linux file premission better than me ,,,, and again I only use gpt5.4mini MNINI!.... from here whats left is for it to get cheaper, and to understand bad prompts (and improve prompt injections)
  3. codex is amazing , its better than anything else---- yes it cant analyze files , yes its sometimes forget ,,, yes the "memorizes" functions stop after 400 lines .... yes it repeat mistakes .....

_____ 13... I know regret,, for example after repeated trial and error , the CODEX rushed into coding,,, I asked it over and over to improve , and it kindda failed ,,, its bad with opencv,,,, it doesnt know how to analyze Images , or really use opencv
---- 14..... I think the moment AGI comes I think it can analyze with opencv or whatever,, it stills needs a lot of improvement ,,, but its really great, more than I really ever wanted ... 10 years ago I wanted to get like an automated chatbot , that speak like a human , and no such product existed,,, now we are the the stages ,,, it does hard tasks ,,, and even less slopy than a human
15.... GPT.5.4 still cant reason for example research papers , even if given a lot of time , even if you ask it to summerize .... or to analyze your dataset , and make a simple ML vision

a ranking issue or whatever
16.... the only option is now to shorten the time to execute a task, it doesnt matter if you score 85% when before you scored 50 ... agi is about SHORTCUTS , less words, more highways and speedup

What steps can reproduce the bug?

Long prompts trips it up

What is the expected behavior?

_No response_

Additional information

_No response_

View original on GitHub ↗

8 Comments

github-actions[bot] contributor · 3 months ago

Potential duplicates detected. Please review them and close your issue if it is a duplicate.

  • #17522
  • #18257

Powered by Codex Action

Mahkhmood9 · 3 months ago

the way Machine learning now feels is like

"spews a lot of slop,, in the end , I am going to figure out something that works and use it.... [keep 2000 lines of slop]"

etraut-openai contributor · 3 months ago

Thanks for the feedback!

Mahkhmood9 · 3 months ago
Thanks for the feedback

the whole Chain of Thoat thing is also problematic , because , its first great a 10bullet point list

you ask it do a complex task it takes it 10minutes , but it should have just this LINUX POSIX command instaed

Mahkhmood9 · 3 months ago

@etraut-openai

llm should also avoid prompts , and system prompts ... Its like it has in its database how to solve the issue ,,, but because of a single keyword it refuses to do

but its a very capiable tool , you just forgot to add the correct words

about 60% of the tokens and computer , is spent on what the user mean ---

that being said its very good , more than I ever wanted,

I didnt even expected to ever reach where we are at now

Mahkhmood9 · 3 months ago

@etraut-openai Its like medicine that works 200 years ago ,,, wrong about a lot of things but still reach a desired output

like understanding disaeges speard ,,, but stds and other deadly infectous desigies were very comon --- e g , they think it was a curse ,, or the idea of "bad air",,, or getting afraid

so its generates a lot of slop , does some trial and errors ,, and always succeed ---- but its tooo carefullllllllll and its not a risk taker

but on the other end , its very good

Mahkhmood9 · 3 months ago

@etraut-openai people messure how good a machine model by how good it can make a webpage,, which I guess is a lot of it,,, but I think this, is for debate

I think for llm , and agentic coding, that-be is the

should for example know how to conduct research ,

when I ask it

git clone efficientloftr

and I want you to heavily import code from there

i expect it to work , Or ask me ,
Do you want to import the code? andd use it this way?

bot CODEX is so clever it already adjust per users ,,,

if you behavior was like , oh I like this type of thing , I want you to import the code functions, or this era of why I like it so much

the output is slop,,,, as per small nounces in a ML artitecture ,

I expect it atleast copy paste and modify and make it work ,,,,,,,we are not there yet but it could be a issue with I dont know how to prompt correctly

llm has solved the issue of coding, now we should target the automated auto correct too , and keep improving

instead of being proud of our human generated slop code , keep the same base for 5 years

developers already able to transform whole codebases, without any issues --- surely sometimes the developer magically want it to improve,,, but he didnt know he just had to ask for contesnt

Люди измеряют качество модели машинного обучения тем, насколько хорошо она умеет верстать веб-страницы. Наверное, в этом есть доля правды, но, на мой взгляд, это вопрос спорный.

Я считаю, что для LLM и агентного программирования (agentic coding) приоритеты должны быть другими. Модель должна, например, понимать, как проводить исследование.

Когда я пишу:
git clone efficientloftr
и говорю, что хочу массово импортировать оттуда код —

Я ожидаю, что это просто заработает. Или что модель спросит: «Вы хотите импортировать код и использовать его вот так?»

Тот же CODEX настолько умен, что уже подстраивается под пользователя.

Если бы поведение модели было в духе: «О, мне нравится такой подход, я хочу импортировать эти функции или использовать этот стиль, потому что он мне близок»...

Но на выходе часто получается «слоп» (низкопробный контент) из-за мельчайших нюансов в архитектуре ML.

Я ожидаю, что модель как минимум сможет скопировать, вставить, модифицировать и заставить код работать. Мы еще до этого не дошли, но, возможно, проблема в том, что я не умею правильно составлять промпты.

LLM решили проблему написания кода, теперь нам нужно нацелиться на автоматическое автоисправление и продолжать совершенствоваться.

Вместо того чтобы гордиться нашим «слопом», написанным людьми, и хранить одну и ту же базу кода по пять лет.

Разработчики уже способны трансформировать целые кодовые базы без каких-либо проблем. Конечно, иногда разработчик магическим образом хочет что-то улучшить, но он просто не знал, что нужно было лишь попросить контекст.

Mahkhmood9 · 3 months ago

For example , as per research ,,, grok is the best for research ,,, but it misread everything ---

Its not to the same level as if you were to take a machine-learning-vision expert