or, how to build a star out of mud

User marginalia[1] on lobste.rs compared the output of poorly-directed software development LLM agents to the "katamari" from Katamari Damacy: a video game in which you roll a clump of miscellaneous objects around, sticking everything you can find to the outside. It's an apt comparison; agentic development tends towards feature addition without attention to composition - they take the shortest path to accomplish the prompt. It has all the worst qualities of an underpaid and short-on-time human engineer. However, I think the analogy does a disservice to Katamari! The spirit clod is carefully crafted (algorithmically, sure, but the algorithm was carefully crafted) to appear thrown-together and unplanned, but the method by which they made the ball naturally round has enough complexity that they thought it warranted a patent. By contrast, LLMs don't have the sense required to determine how to properly add whatever new feature their prompter desires, it's just stuck on at random (modulo a probability distribution). Despite that, katamari architecture is a catchy-enough buzzword that I hope it "sticks" around!

Katamari architecture is here, but it's not a hopeless problem. Let's see if we can learn from prior work.

The obvious comparison is the Big Ball of Mud architecture. Foote and Yoder argue that several "forces [...] conspire" to create BBoMs - time, cost, experience, skill, visibility, complexity, and scale. Considering driving forces is a good way to understand a phenomenon: steelman Chesterton before slandering his fence. LLMs of course run roughshod over the metaphor, since they send fences zipping across vast expanses for no intelligible reason, moving them around at random as they sycophantically make bullshit replies to your incredulous questions. But I digress; can you build a katamari from the same impulses as mudballs?

  • Time: LLMs have plenty of time. Ostensibly they can work far faster than any human developer, 24/7 and over weekends and holidays. Time is not the issue.
  • Cost: Many companies have unlimited token budgets, whence tokenmaxing. If LLMs were any good at architecture, cost would be no object. The entire pitch of LLMs is that they're cheaper than humans to do the same work; cost isn't the issue.
  • Experience: Frontier LLMs are trained on (more-or-less) the sum total output of all software ever written, along with all books, blogs, and forum posts ever written. There is no architectural process that LLMs are unfamiliar with, even though the nature of the beast indicates that LLMs don't bring up anything besides the middle-of-the-road, lowest-common-denominator ideas without being, uh, prompted. Experience isn't the issue.
  • Skill: Debatable. LLMs exhibit inhumanly "spiky"[2] intelligence, similar to other automated systems. They're impossibly good at some tasks (underspecified search through a large corpus of text) and hilariously bad at others (any number of publicised LLM epic fails, e.g. strawberry syndrome, car-wash transportation, considering a vending machine metaphysically impossible). As anyone who has to review LLM output for a living knows, the mistakes they make are not the same mistakes a human would make, and they're much harder to spot. Lack of skill is probably contributing to katamari architecture.
  • Visibility[3]: LLMs excel at generating vast swaths of code as far as the eye can see - which is part of the problem. No one is going to perform a close reading of a +6,000/-400 sloc PR - it's exhausting, and it's not going to change anything. The more LLM-generated a codebase, the worse visibility becomes. Andrej Karpathy doesn't even read the code anymore. As Foote and Yoder put it, "[i]f the system works, and it can be shipped, who cares what it looks like on the inside?" That's the mantra of the modern vibecoder to a T. Just surrender your cognition, embrace the clod!
  • Complexity: Most software has quite little essential complexity, and at this scale Conway's law isn't relevant. Complexity may explain some BBoMs, but not the LLM-generated katamaris. I imagine a plurality of truly complex domains still have mostly hand-authored software, though it's a downward spiral at this point.
  • Change: Human effort is a natural brake on the pace of change. Automatic coding accelerates change; LLMs make implementing a new change as easy as requesting it (with some rather heavy asterisks). Unlike human change, though, LLM-effected change tends to agglomerate onto the existing architecture (such as it is) rather than cut through the heart. A heavily LLM-affected katamari codebase often has a well-designed (human-designed, typically) core, obscured by layers of stylized household objects. An agent may decide to redesign the entire system, but rarely does it have the wherewithal; a redesign or rewrite is rightly feared by software artisans as complex, painful, and interminable. LLMs going for a redesign on their own are doomed to failure.
  • Scale: Foote and Yoder's point (unless I've misunderstood - the original is a little unclear to me) is that otherwise-skilled designers struggle to find elegance when faced with massive projects. I don't find this applies to agents all that well, they have mediocre performance even in the small. They don't get any better at working at a large scale, though.

Context compacted


The main contributors from this list are skill, visibility, and change. LLMs are simply not very good at writing code relative to a skilled human, they're a whole new level of invisible code, and they axe the aerodynamic drag that keeps otherwise-muddy projects from collapsing into slop, ten 15,000 sloc PRs at a time.

The visibility one is getting stuck in my craw. It wasn't good enough that users couldn't see the horrible mud inside the application, now not even the developers are reading the code? we're so cooked chat

what is to be done?

LLMs are the accelerationist dream realized. Burn the coal, poison the well, kill the open web, hack the planet (starting with the rainforests). Weed out anyone with latent schizophrenic tendencies. Obsolete trust. Flood the internet with the unspeakable. Hope we come out the other side with infinite renewable energy, infinite compute, secure and performant software, abolished copyright and universal leisure. Cure all the diseases while we're at it. My anarchist utopia sure isn't going to happen as long as the billionaires hold the keys, though. If jyn has any idea what they're talking about, we'll all have Astra-class models on our phones in a few years. Seems unlikely on the face of it.

Hard though it may be to believe, I'm no doomer. Vibecoders are way off, but it's apparently possible to use these accursed orbs to build better software. I also don't claim to have a magic spell capable of bending the cthonic entities to our wills, but non-generative applications have the most promise. LLMs are evidently capable of finding legitimate bugs in carefully written code. Like fuzzing, it burns oodles of compute for this privilege; like fuzzing, hyperscalers are happy to send some alms free compute to keep their pet OSS debugged.

One idea I've seen floated is that of an arms race between slopmeisters and software maintainers: "you cannot use LLMs to contribute to this codebase, unless I can't tell that you used an LLM." I like this - it gives overburdened maintainers carte blanche to send any sloppy PRs to the shadow realm, for one, and in Eliza's words, "the act of rewriting the model-generated code to not Look Like That forces the person who generated it to actually read and understand it thoroughly."


Foote and Yoder eventually conclude that BBoMs are a cynically optimal choice - software changes so rapidly under our feet that "expedient, slash-and-burn, disposable programming is, in fact, a state-of-the-art strategy". I disagree with them, and I assert that katamari architecture is no better. BBoM and katamari architectures are coping mechanisms, wrought from the trauma of disillusioned practitioners being compelled to deliver endless vanity features while blowing past delivery's cosmetic deadlines. Software ate the world, and now it's eating us. We don't need damn-the-torpedos full-throttle break-neck speed, we need care and attention.

I want smaller software with less bugs made by people who are paid more to work less[4] and I'm not kidding!


  1. Author of the excellent smallweb search engine marginalia ↩

  2. Spikes are relative to "human standard", which means neurodivergence (especially autism and ADHD) is often classed as spiky. In no way am I describing LLM intelligence as akin to autistic intelligence - the spikes are not aligned. There's only one way to be normal, but a variety of ways to be abnormal. ↩

  3. A deeper visibility problem, that doesn't fit with the rest of this polemic - LLMs are more-or-less black boxes (notwithstanding some fascinating attempted brain surgery). We have no idea what's going on inside there. Explainable AI is as much a pipe dream as it was in SOPHIE's (recommended listening) era. ↩

  4. Worker-owned co-ops might help. ↩