However I think it's aggrandizing what human engineers actually do with remarks like "Engineers own tradeoffs." My experience is that certainly less than half of the employed software engineers don't actually give a real analysis to questions like:
"Given these constraints, this team, this business, this infrastructure, this budget, these risks, and the expected evolution of the product, what is the most appropriate way to implement X, today?"
Thus I think AI is more able to replace the average engineer more than this article admits, however the inadequacy of "average engineering" will be much more apparent now: codebases can become large/complex enough to be unwieldy in months now when it used to take 5 years [a timescale where accountability is effectively impossible].
These get overlooked so often. The way you build software if you’re at the helm vs the way you need to build it when dealing with a more/less capable team and business, especially if someone else will be doing the deployment and will need lots of consultations, is way different.
This is my favorite part of software engineering. It's not just a set of rules you can apply to get the right answer. You need to use your judgement to make a context-appropriate decision.
It's been ten months since good models started landing and threatening our current job descriptions.
Do you think this is where it stops? This is where it begins.
Machines will be good at managing complexity too. You can't draw a line and say improvement stops here, because everything we've seen so far flies in the face of that.
I shudder to think what these models will be capable of in 24 months.
> Do you think this is where it stops? This is where it begins.
No, this is pretty much where it stops.
The models are good enough for the average coding task, and the slop they produce often is in the category of what a bad or careless dev that's being contracted out might produce.
Yea, they'll get better, but not in next-level sort of way.
The limitation is not the models or intelligence, it's the human in the loop. We're still stuck on stupid human issues ranging from usability bugs, to figuring out what the product should be, to how we should program in the first place.
I know the models are capable of sorting out issues it gets stuck on because it's writing error handling in the wrong way, or just it doesn't have the right abstractions, because we can't settle on the right way to program. I still see people arguing about dyanmic vs. static typing.
And obviously, there is a next level, but that's real singularity, and we're all out of jobs.
I think the limitation is still context. Humans are better at judgment than AI because humans have more context to put together better decisions. when AI has all of the context, It beats the socks off me.. for better or worse sharing and organizing context is a notoriously difficult problem that won’t be solved overnight. This could easily be a 10 or 20 year problem because it depends on human organizational structures and investing in free flow of information and information storage tools. Unifying search tools like Glean help a lot here but they’re also expensive.
Lost me at the first assumption. People can argue about how useful AI is, but it's obviously not essential because we somehow managed to write code without it a few years ago. I would even say the code was better back then.
The two tasks of writing code and engineering software cannot be separated without damaging the integrity of the mental model of the engineer. Having architects who didn't interact with the code always produced map/territory mismatches.
The mental model was required when your brain was the only chance to reason about changes, answer cross-cutting questions (architecture), and develop a visceral feel for the project, because that's what you needed to write high quality software
It's hard to let that go, but you already had to in larger human organizations/collaborations where you might be assigned work on systems you never/seldom touch, or coming back to a project you haven't touched in a long time.
You don't need a mental model when you can automate the reasoning and the benchmarks that vet the reasoning. Your mental model is better spent pondering->reconsidering high level things like invariants, and then automating the the proof and implementation of those decisions.
Consider how you can just get Claude to start a workflow of 15 Fable agents to fan out over your system looking for correction/simplification/perf opportunities before spawn another wave of agents to vet the list of findings. How much time and energy and studying of the code would it have taken you to build and vet the same list?
This whole week I've been dealing with incidental complexity created by shortcuts taken and edge cases not handled in code written in the before times, both in mine and in others'. I realized at some point yesterday that these kinds of shortcuts would no longer be accepted with competent LLM use.
Have you looked at the machine instructions your compiler produces?
No? Why?
Because software languages are a pretty good abstraction.
To the extent that good abstractions are in place, you can avoid looking at code specifically.
Those don't perfectly well exist, so it takes a lot of self discipline and the right tools/methods, but invariably, AI will produce better systems.
That said, its very easy to produce slop, so well see much more of it.
But mostly, it will be AI from here on in, as a matter of productivity. There are some arguments on the margins but those will fade over the next few years.
'At minimum' - the 'power tools' are here to stay.
A compiler translating high-level code to machine code is a predictable and repeatable process.
An LLM translating a prompt to to high-level code has a much lower degree of predictability. To say an LLM prompt is a comparable abstraction is unfair, though I admit it's getting very close.
I don't get the use of "abstraction" in this context, I must admit.
Programming abstractions offer interfaces to functionality that are both simplified in use and restricted in capability. (e.g. any API or compiler.) I don't see how LLMs meet that definition.
It seems more like we're talking about offloading or delegation, here. And that's a valid business tactic, certainly, but it's not a software abstraction any more than a CTO is an abstraction of a tech lead, no?
> But yes, it has to fulfill some kind of contract defined by the absraction.
I don't follow. Is the contact here the design specification for the system? If so, again, I'd argue that's not an abstraction.
An abstraction could be a design requirement, expressed in some way.
That's definitely an abstraction.
IDLs are a form of abstraction, they're a requirement somewhat more formally described.
Remember UML? That was an attempt to go 1/2 layer above the code, that was an abstraction.
There were tons of tools like that.
APIs are an abstraction - maybe the best example. We write code to match exactly the behaviour defined by an APU - as long as it meets the requirement of that contract, then 'it's good'. And there could be many ways of doing that.
Higher level languages are still formal languages. I think there's a conceptual difference between moving from one formal language to another (machine instructions to asm or asm to C) and moving from a formal language to natural language. So yes, developing, looking at and understanding a formal description of your system has benefits for an engineer compared to handing off this step completely.
Sure but you only look at it when optimizing some performance critical code, usually the hot path. That is usually a tiny fraction of the codebase. I also use AI to generate large amounts of code but I only inspect the actual code when critical, delicate or architecturally important stuff is involved.
If you're doing high performance Java you're definitely looking at the compiler output. At least you should be.
It's arguably more important with Java than with the compiler output for something like C++, as C2 is much more unpredictable and dependent on runtime circumstances. You also want to be real certain that bounds and null checks are omitted as those come at a pretty big performance premium.
> Have you looked at the machine instructions your compiler produces?
> No? Why?
> Because software languages are a pretty good abstraction.
No, it’s because compilers produce deterministic output. I am so tired of this argument.
If I’m not concerned with the performance of my code, I can be 100% confident that that exact code will produce the correct assembly every time. That’s why I don’t read it. Not because I don’t care.
No, determinism is actually required. If compilers changed between producing decent assembly and crap assembly from run to run, we would be reading the generated assembly every single time.
LLMs are pretty much the exact opposite of a reliable compiler-like abstraction.
This is true to such an extent that I have to question the overall competence of anybody who makes the comparison. It's an enormous red flag.
I'd recommend reading Joel spolsky's leaky abstractions essay coz while it applies less and less 20 years later to things like kernel abstractions it explains very well why treating the LLM as a compiler sets you up for abject failures.
I don’t know, I could be wrong but I think one of the key aspects of system design which I don’t know if there is a lot of training data for us the “why” behind decisions. Separating out good design from bad isn’t always black and white and like the article mentions it’s about managing complexity and trade offs. It’s hard to capture in code/training data “we designed everything in a certain way but compromised in this one area because we were under time constraints and assumed we could fix it later”
A fun little exercise you can do is design a system and write some code and then ask LLM to explain why you wrote it that way. Results are varied and interesting but in my experience rarely capture the actual why behind decisions.
Software complexity grows superlinearly, if not exponentially as you add components.
There are things an engineer can do to flatten the curve - that is OP's complexity management idea - but complexity growth can never be linear as long as you are adding to the software.
Code generation has exposed that verification is the central problem of software engineering. And I think it always has been.
Defining what is "correct" can be hard enough, let alone building a system that lends itself to verification, let alone spending the time to verify. Releasing software and letting users find bugs is therefore a very efficient strategy, because it spreads the burden. But you have to ride the line between losing users and getting enough feedback to find and fix the bugs that matter.
As we confront whether AI might take our jobs, I take some comfort in the idea that the world might be too complex for even the largest, best trained AI we can imagine. At a certain point, you need to simulate the whole world (or some substantial portion of it) and the cost/benefit of trying to do all that with compute may not be worth it versus using the real world (that is, humans) as your verifier.
Sometimes these tradeoffs involve half a dozen over a few lines of code. And that’s where I’m hesitant to let an agent work. It’ll do fine with creating correct code. And you can somewhat constrain it to think about one other thing. But it loses track, ignores constraints, cheats, and do you layer complexity on top to prevent this? Or just look at a dozen lines of code to fix it?
Always reminds me of why OOP came about in the first place; it was a way to manage complexity and led to much better and grander software. Now nobody talks about OOP because abstractions are built into just about everything.
It was sold as a way to manage complexity but then made everything complex in a different way. For some problems, OOP makes sense, but for many I think it doesn’t. Unless you have a Rust like trait system that looks a lot like OOP but isn’t. That works well. Essentially don’t put state inside your classes, or you will be spending lots of refactoring time on moving variables up and down in the class hierarchy or throwing computer out of the window because a variable on second thought shouldn’t have been added near class Y.
The real wild shit I'm seeing and having trouble reconciling with continuing my career in this field is that there seems to be a majority contingent of C-suite out there that is absolutely obsessed with force-feeding their organizations AI.
As an software engineer, I will readily admit that LLMs have greatly increased my output - especially on the menial work.
But now we have leadership telling everyone to "use moar AI" on everything, everywhere. I literally have observed folks dropping into incident Slack chats saying things like, "hey all - i asked Claude about this issue and then i had it write a solution. here's the PR." This feels like the kind of thing that should be a fire-able offense, but instead they're getting shout-outs from the CEO.
Hell, the next time I go on vacation, I think I could put Claude Code on YOLO mode for 2 weeks and I'd probably come back to find I'd been promoted.
I do not know how this is going to end, but I have a feeling it's going to get way darker before it gets better.
Thanks, might be the best blog I've read in 2026. AI can of course do architecture as well but oh boy will you have a bad time when your application breaks and neither you or AI can fix it.
One question missing from the list, and it’s the first one I tend to ask… do we really need this? I’m not sure I’ve ever seen an agent pushing back on a request.
You can tweak them to do so. I personally tweak mine to act like a disappointed stack exchange veteran.
I personally recommend, but I understand many people do not want to be pushed back by something they see as little more than a servant.
This setup does work to also have agents argue with each other. That can be very interesting, though you have to set them up to be very skeptical. Otherwise they will tend to read another agents assertion as authoritative off the bat.
I am convinced much of the harness/prompt engineering we are doing now will also be automated away. Within 5 years the best practices for the most popular use cases will have been found, automated and fully baked in.
AI writing code is a force multiplier and amplifies an orgs existing practices. In other words, if you lack structure and are a fan of chaos engineering, then that gets way worse. On the other hand, if you already have god workflows and an overall structure, it'll help you get things done more quickly.
Its about creating a programmatic tool or game that solves something or creates joy to someone or a group. Thats it. The rest is just a way to make to above goal cheaper or easier to accomplish
This. Software hypothetically can permit unlimited complexity. There are only so many ways you can try to build a bridge such that it won't collapse. There are an infinite number of ways that a set of instructions can reach a desired state given enough time, processing and memory resources. This, combined with our relative naivete in how to design and build software that does not approach infinite complexity compared to other engineering disciplines means that the primary thing about building software is managing complexity.
Look at the evolution of any complicated human-made thing - for example jet engines or even better computer hardware. Managing complexity in physical things is different and harder I think - because you're putting a stake in the ground deeper and earlier, as opposed to a software which, as the phrase goes, is just a collection of bits being fed into a machine.
These software engineering analogies are getting tiresome.
People are shouting “yeah the hard part was never writing code, it was managing complexity” as a sort of last hurrah before AI engulfs them.
This is reality: not only can AI write code. It can manage complexity.
Prompt: read the article in this thread then execute its principles on my codebase. Write a harness and programmatic procedures that will trigger you to respond with the articles philosophy to code changes. Be vigilant and monitor every aspect constantly.
I would say for the above prompt, AI is about 60 to 70 percent as a good as a human now. A year ago it was 20 percent. The gap is closing.
No the AI cannot. Did you even read the article? The principles in the article are not rules that can be applied or handed to a prompt. They are questions, not answers. Questions that are impossible to answer and that have no right answer except by human judgement in a concrete context.
Even a human could not “manage complexity” if it’s not in the right context. This is not about AI vs. human capabilities.
It is revisionist to say that software engineering was never about writing code. It was, in fact, a huge component, and it also wasn't easy. Sure most code is glue but even the glue was tedious and the actual hard and novel parts still aren't really done that well by AI (yet).
It's less about writing code now but we're lying if we try to pretend it was a distraction and not a big part of the real work.
And every claim about what the job actually is or was all along has an implied (for now) at the end of it.
It’s not AI-generated. I’m the author. I wrote it in about an hour this morning, then spent another twenty minutes or so fixing a few mistakes. There are probably still a few more in there :)
I have no idea why people believe AI will not be good at all the other things. It's a general reasoning machine, it surely can reason on many things beside the actual code.
I've been hearing this "writing code is not what being an enginner is" mantra for years like some sort of gotcha. (It was prevalent even before AI, and I think people underestimated a lot how many people were simply incapable of writing code even given all the specs and design choices.)
Technically it is not a reasoning machine. If it was a reasoning machine then we would not see results like this:
> To systematically investigate the role of end-user semantics of derivational traces, we set up a controlled study where we train transformer models from scratch on formally verifiable reasoning traces and the solutions they lead to. We notice that, despite gains over the solution-only baseline, models trained on entirely correct traces can still produce invalid reasoning traces even when arriving at correct solutions. More interestingly, our experiments also show that models trained on corrupted traces, whose intermediate reasoning steps bear no relation to the problem they accompany, perform similarly to those trained on correct ones, and even generalize better on out-of-distribution tasks.
Because it’s not about being “good at things”. You missed the whole point of the article. It’s not about AI vs. human capabilities, even a human could not build software if they are not in the right context.
Here’s an example: back in the 2000s, everyone was afraid programming was going to get outsourced to India or other countries. It didn’t happen, Sillicon Valley continues to spend billions to import engineers to work in person even though it’s 10x cheaper to hire remote outsourcers in India who are just as skilled programmers. Why do they spend 10x to move physical bodies to the office? Because it’s impossible to write good software without being in the physical context of the problem domain and team.
Similarly you cannot outsource to AI, because it cannot have complete context. No matter how good AI is, the problem is not mechanically solvable.
However I think it's aggrandizing what human engineers actually do with remarks like "Engineers own tradeoffs." My experience is that certainly less than half of the employed software engineers don't actually give a real analysis to questions like:
"Given these constraints, this team, this business, this infrastructure, this budget, these risks, and the expected evolution of the product, what is the most appropriate way to implement X, today?"
Thus I think AI is more able to replace the average engineer more than this article admits, however the inadequacy of "average engineering" will be much more apparent now: codebases can become large/complex enough to be unwieldy in months now when it used to take 5 years [a timescale where accountability is effectively impossible].
These get overlooked so often. The way you build software if you’re at the helm vs the way you need to build it when dealing with a more/less capable team and business, especially if someone else will be doing the deployment and will need lots of consultations, is way different.
Do you think this is where it stops? This is where it begins.
Machines will be good at managing complexity too. You can't draw a line and say improvement stops here, because everything we've seen so far flies in the face of that.
I shudder to think what these models will be capable of in 24 months.
No, this is pretty much where it stops.
The models are good enough for the average coding task, and the slop they produce often is in the category of what a bad or careless dev that's being contracted out might produce.
Yea, they'll get better, but not in next-level sort of way.
The limitation is not the models or intelligence, it's the human in the loop. We're still stuck on stupid human issues ranging from usability bugs, to figuring out what the product should be, to how we should program in the first place.
I know the models are capable of sorting out issues it gets stuck on because it's writing error handling in the wrong way, or just it doesn't have the right abstractions, because we can't settle on the right way to program. I still see people arguing about dyanmic vs. static typing.
And obviously, there is a next level, but that's real singularity, and we're all out of jobs.
The two tasks of writing code and engineering software cannot be separated without damaging the integrity of the mental model of the engineer. Having architects who didn't interact with the code always produced map/territory mismatches.
It's hard to let that go, but you already had to in larger human organizations/collaborations where you might be assigned work on systems you never/seldom touch, or coming back to a project you haven't touched in a long time.
You don't need a mental model when you can automate the reasoning and the benchmarks that vet the reasoning. Your mental model is better spent pondering->reconsidering high level things like invariants, and then automating the the proof and implementation of those decisions.
Consider how you can just get Claude to start a workflow of 15 Fable agents to fan out over your system looking for correction/simplification/perf opportunities before spawn another wave of agents to vet the list of findings. How much time and energy and studying of the code would it have taken you to build and vet the same list?
Problem with the analogy is that the strain in software engineering is necessary for an in depth understanding of the code.
The question is whether that depth of knowledge is ultimately more helpful than the speed that we can build with AI.
I love to say that
No? Why?
Because software languages are a pretty good abstraction.
To the extent that good abstractions are in place, you can avoid looking at code specifically.
Those don't perfectly well exist, so it takes a lot of self discipline and the right tools/methods, but invariably, AI will produce better systems.
That said, its very easy to produce slop, so well see much more of it.
But mostly, it will be AI from here on in, as a matter of productivity. There are some arguments on the margins but those will fade over the next few years.
'At minimum' - the 'power tools' are here to stay.
An LLM translating a prompt to to high-level code has a much lower degree of predictability. To say an LLM prompt is a comparable abstraction is unfair, though I admit it's getting very close.
But yes, it has to fulfill some kind of contract defined by the absraction.
It's less a problem of the LLM, and more so how we use them, and the inherent tooling around it.
Programming abstractions offer interfaces to functionality that are both simplified in use and restricted in capability. (e.g. any API or compiler.) I don't see how LLMs meet that definition.
It seems more like we're talking about offloading or delegation, here. And that's a valid business tactic, certainly, but it's not a software abstraction any more than a CTO is an abstraction of a tech lead, no?
> But yes, it has to fulfill some kind of contract defined by the absraction.
I don't follow. Is the contact here the design specification for the system? If so, again, I'd argue that's not an abstraction.
That's definitely an abstraction.
IDLs are a form of abstraction, they're a requirement somewhat more formally described.
Remember UML? That was an attempt to go 1/2 layer above the code, that was an abstraction.
There were tons of tools like that.
APIs are an abstraction - maybe the best example. We write code to match exactly the behaviour defined by an APU - as long as it meets the requirement of that contract, then 'it's good'. And there could be many ways of doing that.
Looking at the compiler output is a totally valid concept, but it's definitely a niche case.
It's arguably more important with Java than with the compiler output for something like C++, as C2 is much more unpredictable and dependent on runtime circumstances. You also want to be real certain that bounds and null checks are omitted as those come at a pretty big performance premium.
> No? Why?
> Because software languages are a pretty good abstraction.
No, it’s because compilers produce deterministic output. I am so tired of this argument.
If I’m not concerned with the performance of my code, I can be 100% confident that that exact code will produce the correct assembly every time. That’s why I don’t read it. Not because I don’t care.
It's entirely the nature of the abstraction.
You want it to work as expected, it does not have to produce the same thing each time.
This is true to such an extent that I have to question the overall competence of anybody who makes the comparison. It's an enormous red flag.
I'd recommend reading Joel spolsky's leaky abstractions essay coz while it applies less and less 20 years later to things like kernel abstractions it explains very well why treating the LLM as a compiler sets you up for abject failures.
It's not as good at system design as writing code, yet. But it feels like it's better than most of my coworkers.
I think in a few months, system architecture will have its Claude Code moment, and humans will be outclassed.
A fun little exercise you can do is design a system and write some code and then ask LLM to explain why you wrote it that way. Results are varied and interesting but in my experience rarely capture the actual why behind decisions.
There are things an engineer can do to flatten the curve - that is OP's complexity management idea - but complexity growth can never be linear as long as you are adding to the software.
I made a model/theorem for this that I posted on X: https://x.com/i/status/2027771813346820349
Code generation has exposed that verification is the central problem of software engineering. And I think it always has been.
Defining what is "correct" can be hard enough, let alone building a system that lends itself to verification, let alone spending the time to verify. Releasing software and letting users find bugs is therefore a very efficient strategy, because it spreads the burden. But you have to ride the line between losing users and getting enough feedback to find and fix the bugs that matter.
As we confront whether AI might take our jobs, I take some comfort in the idea that the world might be too complex for even the largest, best trained AI we can imagine. At a certain point, you need to simulate the whole world (or some substantial portion of it) and the cost/benefit of trying to do all that with compute may not be worth it versus using the real world (that is, humans) as your verifier.
As an software engineer, I will readily admit that LLMs have greatly increased my output - especially on the menial work.
But now we have leadership telling everyone to "use moar AI" on everything, everywhere. I literally have observed folks dropping into incident Slack chats saying things like, "hey all - i asked Claude about this issue and then i had it write a solution. here's the PR." This feels like the kind of thing that should be a fire-able offense, but instead they're getting shout-outs from the CEO.
Hell, the next time I go on vacation, I think I could put Claude Code on YOLO mode for 2 weeks and I'd probably come back to find I'd been promoted.
I do not know how this is going to end, but I have a feeling it's going to get way darker before it gets better.
I wish everyone in our industry read 'No Silver Bullet, & Grug-brained Developer'.
a lot of complexity - is about what can we do now, with what we have.
I personally recommend, but I understand many people do not want to be pushed back by something they see as little more than a servant.
This setup does work to also have agents argue with each other. That can be very interesting, though you have to set them up to be very skeptical. Otherwise they will tend to read another agents assertion as authoritative off the bat.
I am convinced much of the harness/prompt engineering we are doing now will also be automated away. Within 5 years the best practices for the most popular use cases will have been found, automated and fully baked in.
How often do engineers get a say in product direction?
Every one keeps saying that AI isnt moving the needle on the bottom line.
Well duh, code doesn't move the bottom line, features do, products do.
If you're building all the wrong things faster, all your doing is performing a speed run to a legacy code base.
The site just goes into a reload loop on iOS?
People are shouting “yeah the hard part was never writing code, it was managing complexity” as a sort of last hurrah before AI engulfs them.
This is reality: not only can AI write code. It can manage complexity.
Prompt: read the article in this thread then execute its principles on my codebase. Write a harness and programmatic procedures that will trigger you to respond with the articles philosophy to code changes. Be vigilant and monitor every aspect constantly.
I would say for the above prompt, AI is about 60 to 70 percent as a good as a human now. A year ago it was 20 percent. The gap is closing.
Even a human could not “manage complexity” if it’s not in the right context. This is not about AI vs. human capabilities.
It's less about writing code now but we're lying if we try to pretend it was a distraction and not a big part of the real work.
And every claim about what the job actually is or was all along has an implied (for now) at the end of it.
Which is why the productivity of people of people has never been correlated with typing speed.
In other jobs productivity is correlated with typing speed and in those jobs a typing speed like 60wpm is part of the job requirements.
I've been hearing this "writing code is not what being an enginner is" mantra for years like some sort of gotcha. (It was prevalent even before AI, and I think people underestimated a lot how many people were simply incapable of writing code even given all the specs and design choices.)
> To systematically investigate the role of end-user semantics of derivational traces, we set up a controlled study where we train transformer models from scratch on formally verifiable reasoning traces and the solutions they lead to. We notice that, despite gains over the solution-only baseline, models trained on entirely correct traces can still produce invalid reasoning traces even when arriving at correct solutions. More interestingly, our experiments also show that models trained on corrupted traces, whose intermediate reasoning steps bear no relation to the problem they accompany, perform similarly to those trained on correct ones, and even generalize better on out-of-distribution tasks.
https://arxiv.org/abs/2505.13775
Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
Here’s an example: back in the 2000s, everyone was afraid programming was going to get outsourced to India or other countries. It didn’t happen, Sillicon Valley continues to spend billions to import engineers to work in person even though it’s 10x cheaper to hire remote outsourcers in India who are just as skilled programmers. Why do they spend 10x to move physical bodies to the office? Because it’s impossible to write good software without being in the physical context of the problem domain and team.
Similarly you cannot outsource to AI, because it cannot have complete context. No matter how good AI is, the problem is not mechanically solvable.