The Automation of Understanding
AI coding tools make cargo cult programming scale to whole systems, and the real risk is not bad code but losing the understanding of why the code works.
· 10 min read
Note
A local LLM was used to rephrase this content. Reviewed by Yatharth Sood, check him out @ ysood.xyz.
This is about what happens when the machine that writes your code also writes the tests for that code, and you stop being the person who understands either of them.
Cargo Cult Programming at Scale #
There is an old idea in engineering called cargo cult programming. The term comes from Richard Feynman’s famous 1974 talk, Cargo Cult Science. Feynman described people who copied the visible parts of a scientific process while missing the thing that actually made it work. They built runways, control towers and antennas because they had seen those things around airports, and then they waited for the airplanes to arrive.1
Programming has its own version of this: a developer finds some code on Stack Overflow, does not fully understand it, but copies it because it works. Over time that habit scaled up through libraries, tutorials, GitHub repositories and blog posts, and we learned to assemble software from pieces written by other people. Now we have AI agents (agentic coders, or whatever cool term they use this week), and the scale of cargo cult programming is changing again.
This creates a strange situation, because the software can actually work. A developer can ask an LLM to build an API and receive a database schema, an authentication system, API routes, Docker configuration, tests, logging and deployment configuration, then run it, make a request, see the expected response and ship it. The system works. The developer may understand what the application does without understanding why it was designed that way, and they may not know which assumptions are hidden inside it, which parts are unnecessary, or what will break when one component changes.
That is where this starts becoming interesting. Bad code is easy to distrust, but good-looking code is much harder to distrust, and the better LLMs become, the easier it is to trust them. It is a natural human reaction: if the machine keeps giving good answers, we stop checking every answer. We already do this with other forms of automation. We do not manually calculate every number produced by a calculator, we do not check the mathematics behind every route suggested by a navigation app, and pilots do not manually control every part of a modern aircraft. Automation works because humans learn to trust it, and that trust creates a known human-factors problem.2 NASA has studied automation complacency for decades,3 and the pattern is consistent: when an automated system becomes highly reliable, operators pay less attention because they expect it to work, and when the automation eventually fails, they have worse awareness of what is happening and are slower to recover.
Software development is moving in the same direction. Imagine an LLM that gets things right 99.9% of the time. You ask it to implement something and it works, you ask it to fix a bug and it works, you ask it to refactor the database layer and it works, you ask it to write tests and they pass. After thousands of successful interactions, your brain starts building a simple rule: the AI knows what it is doing. At that point the AI stops feeling like autocomplete; it starts feeling like a colleague, then a senior colleague, and eventually an authority. This is closely related to automation bias, where humans place excessive trust in automated recommendations, and research on automated decision systems has repeatedly shown that people will follow automated recommendations even when those recommendations are wrong. The AI does not need to be perfect; it only needs to be good enough for us to stop looking.
Then comes the interesting failure. Suppose the AI makes a mistake, but the resulting application still works. The API responds, the database contains data, the tests pass, and the logs look normal. Then one unusual combination of events produces a subtle bug. The AI generated a perfectly reasonable solution; it just happened to be wrong.
Now the human has to debug it, and where do they start? They can inspect the code, but they never really understood why it was structured that way; they can inspect the architecture, but the architecture came from the AI; they can inspect the assumptions, but they do not know what those assumptions are or why they were made. So they ask the AI to explain the code, which means they ask the same system that created the problem. The AI looks at the code, produces an explanation that sounds convincing, and the developer accepts it. The original problem remains. This creates a strange debugging loop: the AI generates the system, the AI explains the system, and the human trusts the explanation. The human has become a supervisor without the knowledge required to supervise.
When the Tests Become Part of the Black Box #
There is another layer to this problem. Software has traditionally had an important escape hatch: tests. Tests are supposed to be an executable version of what humans actually care about, so the implementation can change completely, but the tests define what must remain true, and they give the human a way to control the machine logically instead of relying entirely on explanations. But tests are nothing but code with assertions, and if they are code, then LLMs can generate them too, which creates a much larger loop.
A product idea can be turned into a PRD by an LLM, the PRD can be turned into development context by another LLM interaction, that context can be used to generate a test suite, and another LLM can then generate the implementation that satisfies those tests. The human can review each step, they can approve the PRD, approve the tests and approve the code, but the amount of actual human reasoning inside the loop can become surprisingly small. The human is increasingly becoming the person who says “looks good,” sitting in a chair watching someone else do the work that used to be theirs, aka a cuck, a cuck programmer. This matters because tests are supposed to provide an independent constraint on the implementation, and if the same system generates both the thing being tested and the definition of what counts as correct, that boundary becomes much weaker. Imagine asking a student to write an exam and then allowing the same student to grade their own exam. The student might do everything correctly, and that is not the point. The point is that the grading system no longer provides much independent information about whether they did.
This is the deeper version of cargo cult programming. Let’s just call it cargo cuck programming.
Note
Reminder to drink water.
The Loss of a Mental Model and Understanding #
This is where cargo cult programming becomes much bigger than copying code from Stack Overflow. Traditional cargo cult programming usually meant copying a small piece of code without understanding it, but with LLMs the same behavior can happen at the architectural level: a developer can generate a database design without understanding database design, generate a distributed system without understanding distributed systems, generate a concurrency model without understanding concurrency, and generate a security system without understanding security. The developer can still be productive, and that is what makes this difficult.
“Oh but Nikhil, this was possible before as well.”
Yeah, I know. It’s just easier and much cheaper now, which promotes a faster FAFO development procedure, and that gives birth to yet another problem: optional understanding, optional skill.
Skills (not that skills, get your brain out of the AI gutter) are like muscles: if you stop using them, they weaken. Aviation has studied this problem for decades. As cockpits became more automated, pilots increasingly became supervisors of automated systems, reducing opportunities to practice the manual skills needed when something unexpected happens.4 The same thing can happen to programmers. If the machine always writes the code, the programmer gets less practice writing code; if the machine always explains the bug, the programmer gets less practice debugging; if the machine always designs the architecture, the programmer gets less practice designing systems.
The programmer starts looking more like a pilot. A pilot does not manually control every valve, calculate every number or navigate every second of a flight. The pilot supervises an enormous automated system, and most of the time that is exactly what we want. The problem appears when something unexpected happens, because then the pilot needs a mental model: they need to know what the aircraft is doing, what the automation is doing, how to recognize when the automation is wrong, and how to take control. Software engineers will need the same thing. An engineer who only knows how to ask an AI for code may be extremely productive while everything works, but an engineer who understands the system can still function when everything stops working, and that difference will become increasingly important.
The current research already gives us some interesting signals. DORA’s 2025 research found that AI adoption among software professionals had reached 90%, with many developers reporting productivity and quality improvements, and DORA also describes AI as an amplifier of existing organizational strengths and weaknesses.5 At the same time, the results are not universally positive. A 2025 METR randomized controlled trial found that experienced open-source developers using the AI tools available at the time took 19% longer on the tasks studied,6 and more interestingly, the developers believed the tools had made them faster. METR later reported that its newer experiment was producing an unreliable signal because developers increasingly refused to participate in a no-AI condition, which is interesting in itself: the tools had become deeply embedded in the workflow.
The lesson is not that AI coding tools are bad (for us); the lesson is that our perception of what the tools are doing can differ from what they are actually doing. The biggest risk is not bad AI-generated code, because tests, reviews, static analysis and security tools can catch a lot of bad code. The bigger problem is losing the understanding of the system itself. A codebase can survive ugly code, technical debt and bad abstractions, but it becomes much harder to maintain when nobody knows why it works. The code becomes a map drawn by someone who has never visited the territory. As AI makes it possible for one developer to generate what once required several developers, the amount of software can grow much faster than the amount of human understanding behind it. That is cargo cuck programming pro max, and this is worse than spaghetti code, trust me on this.
The answer is not to stop using AI, but to make sure humans can still take control when things go wrong. Engineers should still understand the code, the architecture and the assumptions behind the system, because AI can write the code without understanding the consequences, and when everything works, that understanding feels unnecessary, but when something breaks, it becomes the only thing that matters.
The better AI gets, the easier it becomes to forget that.
-
Richard Feynman, Cargo Cult Science, Caltech Magazine, 1974. – PDF (2nd Page, 5th paragraph on the left) ↩︎
-
NASA, Examination of Automation-Induced Complacency and Individual Difference Variates. ↩︎
-
Summary by /u/grauenwolf on r/programming; Note: I do not wish to give away my data for a mere PDF. Other source Aviator, AI Won’t Fix Broken Systems: Lessons from the 2025 DORA Report ↩︎
-
METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. ↩︎