Practical Lessons From Six Months Using AI All Over The SDLC
First Published:
Updated:
Note: this piece, like everything I post here on this blog, is 100% my words. I think AI generated code is fantastic, I think AI as writing coach is fantastic, but you will never read AI-generated text on this site.
It's been six months since the February inflection point, Opus 4.6's release, which made LLM-generated web- and data-stack code viable to put into production with minimal human review. This changed a lot of things, and has revealed some interesting new ways of working. Here's a list of what I've learned so far, which I suspect is mostly things you have learned as well - assuming you're reading this as a fellow technologist - but hopefully contains a few novel ideas as well:
An LLM generating code doesn't replace the need for software engineers in many codebases. This seems obvious to me, but it also seems apparent that it's not obvious to everyone. It will in a few years when people observe that a structural engineer doesn't usually swing a hammer all day, nor do electrical engineers weld circuit boards.
Complementary point: Not every codebase has ever needed an engineer, some just need a coder/programmer, and now those projects are easier than ever for anyone to standup. I'm thinking prototypes, but also simple internal tools inside of tech orgs.
In the meantime, if you're worried about there not being jobs anymore for software makers, this is a decent history (I suspect AI generated from the writing, but it does lay out a good map of the past).
The gap between the skeptics and the advocates is getting bigger. As often is the case in matters cultural, Charity Majors has a great post on this. Might as well just quote it, because I'm not going to say it any better:
The enthusiasts are not wrong. We are starting to see real, non-imaginary, discontinuous leaps in capabilities from teams that lean in hard to working with AI. And this does not feel like a normal technology cycle where you can wait for the dust to settle; teams that sit this out while competitors are hustling could be out of business before the dust settles. That’s a real, existential threat.
The skeptics are also not wrong. When you ship code faster than engineers can read it, in domains where nobody has full context, you are making withdrawals from a trust account that took years to build. Reliability degrades, institutional knowledge evaporates. You end up with systems nobody understands, products burbling into incoherence, and on-call rotations that grind people up and spit them out. That is ALSO a real existential threat.
This much change in this little time is really, really hard. You don't need to be a "doomer" or a decelerationist to feel anxious or worried. I'm generally fairly optimistic, but not pollyannaish, and I'll admit there are a few potential paths I could see us going down that would not be good for huge swaths of people.
So what to do about it?
A few thoughts:
- Empathy. I think about the distinction between sympathy and empathy a lot. Another post I should probably write. As social media continues its mission to make every problem my problem (and your problem), I feel more and more guarded every year when it comes to sympathy. Empathy, on the other hand, where you understand and share the others feelings? Always room here. So I try to start every conversation about AI with just a little acknowledgement of, wow, this is a really hard year. Fun - amazingly fun. And... hard.
- Understand and enumerate. What are the changes happening? One of them is that different work styles challenge us, our brains and our energies differently. Another is that expectations are changing, have changed. Another is that power dynamics have changed. Not all of that is evenly or fairly distributed.
- Nobody knows where this goes. We're all just making our best guesses.
AI will challenge your first principles. The truisms and maxims you've learned and developed over the years, and maybe the decades if you've been in the profession that long? Some of them still hold, some of them need revising or refactoring, and some might need to be thrown away. Determining which is which will take time, and this will cause...
Chaotic-Good and Chaotic-Neutral oriented peers will see and sense your indecision / reflection, and push the boundaries. As Kevin Kelly says, "...large systems must tread a path between the ossification of order and the destruction of chaos", and AI might be the biggest anxiety-driven change catalyst introduced to tech workplaces in my career. More significant that the GFC, more of a paradigm shift than Cloud.
Some of the quick scans you used to make no longer work. I used to be able to look at a ticket - Jira, Trello, whatever - and given even a modicum of knowledge about the product, immediately identify a ticket / story / increment that needed fleshing out.
AI is really good at generating lots of words that sound authoritative. Now the ticket that needs fleshing out - or rather needs a good flaying - looks at first glance like the ticket that's thorough and concrete. Which leads me to...
Suddenly, everything has a concrete solution. Where humans waffle, AI will confidently state a direction, built on whatever evidence is in its context window. Which is sometimes the full picture, but rarely the full full picture which humans usually reason from.
This is good and bad. My current shop is in a lean-and-mean phase, so there's not a lot of Resume Driven Development going on, but AI is really good at spotting it anyway. No, you don't need an eigth persistence technology, the seven you already use cover the Venn diagram of optimal use cases.
Some of the added output capacity must go into care & feeding. If you're getting a 3x increase in velocity, for example, but it all goes into features or worse into feature complexity? You're going to be in a bad spot in a few quarters as aspects of the system that AI doesn't accelerate as quickly start to stretch.
Skills and loops for everything. All the way to the factory model, or just towards it? More flowing through the system will result in more features, even if you try to redirect output as advised above. Bare LLMs without skills and loops are great at 80/20ing. As mentioned above, the eye tests you might be using need recalibrating.
Agents checking agents' work can be a powerful pattern. I could write an essay on that alone, and probably should. The caveat of course is that the frontier model providers seem to be building this in as rapidly as we can built it out. Opus 5 anecdotally uses 2 or 3 times as many tokens as Opus 4.8 on some coding tasks. What's it doing with these tokens? A lot of it seems to be spinning subagents to verify its work.
So the harness around the harness will change and adapt as they're subsumed by the providers, and likely new businesses will emerge around providing different levels of rawness / Bring Your Own Harness (BYOH?). Edit from September 2026: like Jev.
Your non-engineering coworkers who want to build software are going to have to learn all the lessons of software engineering. For an example in just one layer, get ready for everyone to learn why we build component libraries and restrictive style guides, as a hundred front-end PRs merged into your repos with inline styles everywhere creates a new layer of tech debt and fragility.
The general problem is an old one. You don't need an "engineer" to get to the first 80/20 of almost any web software app. You absolutely need one, and often many of them, to squeeze the 5th 9 out of a 99.99% uptime system. Modeled correctly, most people would identify a specific line that would be crossed along the way. Nobody ever models it though... often you don't know what you don't know about a thing that might start innocently enough...
Hard problems remain hard problems. Constraining complexity. Testing at the top of the pyramid. Testing in production. Prioritizing the right quality work while avoiding temptations to gold plate. I've yet to see, much less be convinced, that AI in the SDLC make any of these canonical problems shift from hard to easy.
Code review and functionality review / verification reviews are the new bottlenecks. Here's a vomit of general thoughts, YMMV of course as a lot of this is contextual to the impact of downtime. If downtime is a non-event - because you're pre-release, or your audience uses your software only at specific times - then fire away. If you're in a mature org, the art is carving out pieces that can be treated as if they're in pre-release mode.
At the beginning of my career, I might ship to prod 20 times a day. That's not an exaggeration. I was building a combined external experience for several thousand customers plus internal tooling for about ~20 users across sales and operational teams. If someone found a bug, I expected myself to have it fixed within a few minutes and have it in production within a few seconds after that. That might sound absurd given that most CI runs don't even have the bootstrapping for their tests stood up in less than a minute, but I can assure you it was entirely possible in 2004 given a developer in perfect context with the entire codebase - me. This has been impossible in collaborative environments - until now, with agents + context.
There's a lot of incentive to perform, to be a first mover, and to generally brand yourself as AI-fluent, AI-native, AI-whatever. I suppose Im doing so as well with posts like this. Every shop above a certain size - 50 engineers, to throw out an entirely arbitrary ballpark - needs someone who pretty much plays with newness full time. One thing you can do if you're an engineer in a shop like this is position yourself right behind them - watch what they do without being tempted to dive in too soon, then take what's working well from their experiments.
The information coming at you is 10x. Some of it is of dubious truth. Some of it is out and out bullshit, generally a result of partial or just plain wrong context, but occasionally with elements of malice or ignoble aims as well.
This has strange ramifications, one of them being that anything is blamable on AI. Where previously someone fabricating a deliverable on a performance review might be grounds for sanction if not termination, now they can say "well, AI embellished that, I didn't notice it". How long this grace period of sorts will last is anyone's guess.
I don't see enough written about the second and third-order effects here. Maybe another post I should write.
Using AI to build AI / agentic features is the best way to learn. I hate the advice of moving around frequently, generally believing that people need a year or two to really hit their stride on a functional team, but sometimes it's the right one - if you're still only working on CRUD apps, I suggest finding a place in your org and, if needed, a new org where you can work on agentic features.
Agent-as-assist vs agent-as-sku. I have to admit, I'm a product maker and software developer first, with my business brain tailing behind. So it took me a while to see that while my first agentic features were and are working well, they weren't being lauded as well as another set of features with less net adoption. I thought for a while it was just how I was telling the story, but a coworker helped me see the difference: my agents were helping users do a thing they were already doing faster/easier, whereas my peer's agentic software was an upsell sku. Ah!
Everything around me is being built like a garden being sewn - a series of rapid 80/20s that arrives towards an intended solution with sometimes surprising rapidity. I'll write more about that later this year, it's a really important switch and accounts for things to hear like we're still hiring software engineers, it's just that we're looking for different things now.
A side note here: cost to iterate and maintain as a second order effect. Everything needs a reset, because the software appears in these 80/20 iterations, and nothing is popping out of the oven as fully baked as it used to when human hands were writing the code.
On Probabalistic vs Deterministic and the shift towards the former that we're seeing. I think about this a lot, and am tempted to say I'll write about it later this year, but everything I want to say is already said really well here.
AI usage resembles petitionary prayer, where one of if not the primary benefit is to force you to decide and clarify what you want so you can ask for it clearly, succinctly and without ambiguity. As I said earlier in the year, this might be my most profound thought about AI, and one I haven't seen too many places. It deserves its own post.
AI usage seems to encourage pushing back, something many in the workplace forgot how to do in the last decade. AI usage makes it more natural to live with and react to an entity who, though the ideas are good, you disagree with. And we're pattern following monkeys, so now we're slightly more likely to do the same with a human.
Note from September 2026: This is nested inside a profound idea a coworker shared with me that I'm still wrestling with: that learning to prompt well might also be learning to socialize well.
Just do the rewrite. Hot Garbage Architecture FTW (still one of my all time favorite talks). The big advantage of the rewrite is that you can write your context docs from first principles, fresh, accurate. This helps with the build but it really helps with future maintenance - that is, until someone else decides to throw out what you've built and start again again.
Questions I've been repeatedly pondering this year, this 2026 of so much change. Keep in mind my day job is leading a team of software engineers, a role that's joyously gotten a lot more hands on this year.
- What is the right system to use to rapidly build high quality software with LLMs in the SDLC given the conditions and constraints of my shop?
- What kind of engineer becomes more valuable when writing the code becomes the cheapest part of the SDLC?
- For those who aren't motivated by the ability to build things, what replaces the joy of puzzle solving and crafting a perfect module when the code itself is written by a machine?
- What becomed possible when a very small team or even one person can quickly conjure significant fractions of the capabilities that used to require a large org?
- More generally: What becomes scarce when intelligence becomes abundant, and what are the most valuable things for a human to be doing in this pre-AGI / highly-capable LLM era?