Hacker News
Why Are Coding Agents So Dumb?
polyterative
|next
[-]
A lot can be improved, but this is already so much speed.
Rapzid
|next
|previous
[-]
Codex Astra can do a great-(ish) job as a project coordinator dispatching tasks to a pool of 6.1 Sol sub agents. You can even give it an explicit goal and ownership over ensuring the work is carried out efficiently.
However the OOTB harness(and prompt) configuration may not do this for you. You'll have to provide guidance over how you want it to operate through your prompt, a skill, or etc.
And I'll say even though it's really good at this.. Having even more layers than 2 can help; a single agent given too many responsibilities will start to become fixated on a number of them while neglected others. You can check in occasionally to "nudge" it or you might need to split out responsibilities more..
I will say it's crazy Codex doesn't have more built-in task and sub agent management features. I almost wish that it had some stock orchestration patterns that worked OOTB, and then you could opt-in to a leaner setup where you provide more of the instruction.
mstank
|next
|previous
[-]
I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.
ilamont
|next
|previous
[-]
This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.
Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?
TeMPOraL
|root
|parent
|previous
[-]
The descriptions are near-useless and tend to flip around, as model families are not released in sync anymore, that's true, but fortunately, thanks in a big way to subscription pricing, the choice is simple: start with the best model on offer, and when you run out of quota, downgrade to the next best (or briefly switch providers).
Neywiny
|next
|previous
[-]
arjie
|next
|previous
[-]
I let most agents work asynchronously and don't pay attention so I don't care that much about the sequential nature. But if it's a problem for you then fix your harness. This is a bit like saying "Why are shoes so shit? There's a stone in one and it just gets stuck there and your foot steps on it and it hurts". Take off the shoe, and shake out the rock. Put the shoe back on. You have the power.
pipes
|root
|parent
|next
[-]
Edit: I have access to codex, vscode, GitHub co pilot cli and all anthropic and openai models (excluding mythos).
namrog84
|root
|parent
|next
|previous
[-]
They even have split my decisions to human decisions. Proposed and approved work. They can iterate on approved work without me just fine.
And I only just started with agentic coding in last few weeks before that I was mostly a copy paste chat person.
Neywiny
|root
|parent
|previous
[-]
Your shoe analogy also breaks down because really the shoe is the issue, not the stone. And expecting everybody to make their own shoes is, well, I mean we just don't do it that way anymore for good reason. Let the cobblers make the shoes, and the runners wear them.
361994752
|next
|previous
[-]
kgeist
|next
|previous
[-]
I run an LLM server with Qwen 3.6 in the office, and OpenCode, which the OP mentioned, usually defaults to sequential TODO lists, and it works fine with our little LLM server with 3-4 parallel users. But I noticed that once in a while the LLM got overloaded with requests in the queue, and you couldn't do anything for 20-30 minutes. My investigation led me to an employee who used QwenCode. I tried it myself then, and indeed, it immediately launched something like 6 parallel subagents, where OpenCode would have sequential TODOs with the same model by default.
So in the end, I had to detect QwenCode on the server side and serialize all its parallel requests into a single request queue, because it made life miserable for other OpenCode users :)
brandtcormorant
|next
|previous
[-]
Have you tried telling models about your dream agent environment?
They can build it.
cbrake
|next
|previous
[-]
One thing that seems to help for me is to do the docs before plans (collaboratively edit with agent). Then I understand what this change is going to look like from the user's perspective before we start implementation. This seems to help keep things on track.
While I don't use this plugin a lot anymore, I think doc-driven development is one of the most effective ways to do development in any paradigm, I should probably refresh this plugin and use it more:
https://github.com/tmpdir-org/tmpdir-claude-code-marketplace...
wrs
|next
|previous
[-]
scotty79
|next
|previous
[-]
I have no idea what stops that person from just making it, with an agent of course.
imimayj1337
|next
|previous
[-]
Kuyawa
|next
|previous
[-]
I asked DeepSeek to translate a page to five languages and it opened five subagents each one working independently on the translation, once they all finished the main agent informed me of the job completion with a bell. Fantastic!
Sooo, which agent?
chrisjj
|next
|previous
[-]
As expected. A model's knowledge is what was it ingested a creation.t
Unfortunately what we get is worse - for the same reason. Model version thinks it is its previous version.