Not to brag, but I’ve got eyeballs. As such, I have been reading countless articles about agentic development this past year. This did not spark joy.
I’ve been reading about how there are 5, or 8, or 10 stages or levels of agentic engineering.
I’ve been reading about dark software factories with promises of churning out amazing quantities of code in record times. About how people run 10 parallel agents to do everything.
I’ve been reading about giving Fable a ten word sentence and it refactored the entire project into 🚀 BLAZING ⚡ FAST 🏃♀️ RUSTY 🦀 CLAW 🤖.
Blegh.
The genre has a “what proposing to my girlfriend taught me about B2B sales” vibe.
But…the story is appealing. As an architect in a startup it’s appealing because we’ve got a backlog the size of a small galaxy, and we should be spending our time doing what we’re best at, and less time doing literally anything else.
But I didn’t know how to go from “we write text in a textbox lol” to “we ship end to end features without human involvement”. If you’re like me, then you can feel the world is heading there, and you’re not sure how to take the first step. It feels so very daunting!
But! As engineers, we know that all complex problems are actually smaller, simpler problems that need to be arranged and solved in succession. Before we build Slack in an hour and replace Salesforce with a prompt, let’s set a much smaller and attainable goal:
How can we automate a pull request that a human can merge in 5 minutes?
…and build up from there.
Let me share some of my humble journey from 0 automated changes to >200 responsibly merged PRs. You’ll see lots of implementation details, you’ll see the difficulties, and you’ll see the internal arguing
More importantly, you’ll see how you too can start the first step in this journey. If you haven’t done so yet, I urge you to follow along with me to see just how attainable these goals are.
!! AI use disclosure
During evening walks with my adorable dog Teddy, I recorded long voice clips which I instructed an LLM to turn into a rough outline. Abhorred by the result, my human hands clicked on the keyboard to completely replace the outline, out of spite. I then used an LLM as an alpha reader.
Start with a PR
We want:
- Pieces of code shipped to prod
- A human reviewer who determines whether the code is good. I’m assuming that you, the reader, are either technically proficient to determine what is good, or have someone who can
- The code that’ll be shipped improves the product/codebase in some way
- A cadence that we can manage
The temptation is to start big to show that we’re doing interesting bombastic things, like “look how we migrated out entire legacy system”.
We resisted the urge. This was our first time actually doing this kind of thing, and we did this while working on our actual jobs, so it can’t take too long. We wanted to do the opposite of a bombastic change: We want to show accrued value over time, so that we can iteratively build up our skills. Many small changes over time = big changes.
So, what we ended up was with a task that can be:
- Easily reviewed in under 5 minutes
- Has a clear marker that it’s AI generated
- Has a small blast radius
- Can be shared with other engineers
I realised that if I do this enough times, I can show that hey, we can have agents ship code! So what should our agent do?
Our first target: Dead code
I don’t recall how, but we decided on trying out dead code removal. It was probably in one of the seven million articles on how to do agentic engineering.
Dead code removal is the opposite of what a human wants to do. It’s boring and not glamorous. It’s just…there, and it takes up space, and it’s annoying, and somebody should do something about it!
I fired up my agent of choice and sent:
1. Clone/fetch the latest code from the main branch
2. Scan for ONE instance of dead code (unused functions, classes, imports, variables). Be thorough but conservative - only remove code that is demonstrably unreferenced. Use grep, AST analysis, or static analysis to verify nothing references the code.
3. If dead code is found:
1. Remove all identified dead code
2. Ensure all tests pass: `uv run pytest`
3. Run linting and formatting, fixing any violations: `uv run ruff format` and `uv run ruff check`.
4. Create a Linear ticket with the `ai-agent` label. In the ticket body, state your goals and explain HOW you found the dead code. Do NOT specify what should be done next.
5. Create a conventional commit with title format `chore($PROJECT): remove dead code in [area]`. In the commit body, detail WHY each part is dead code. Include which commit(s)/PR(s) originally introduced it and which made it effectively dead. Note if it was dead on arrival.
6. Create a pull request to main with the commit subject and body. Add the `ai-agent` PR label.
7. Extract the PR URL and send a message to Slack channel #agent-prs: "Daily dead code cleanup: [PR_URL]"
4. If no dead code is found or no changes are needed, do nothing and do not create tickets or PRs.
Be conservative and focused - only address one area of dead code per run. Include explanatory comments in the PR for non-obvious removals.
It’s pretty basic, but for a beginner, also non-trivial. To pull this off, we need:
- Access to the codebase
- Creating PRs
- Write-access to Linear & Slack
Notably, we don’t assume the agent can figure out how the project works, we give it specific instructions for some project conventions. This has benefits of runtime and cost, but honestly, it also kinda “forced” a workflow that we wanted to add.
Some people might scuff at how obvious it is, because they’re already at floating jar head level. To me, it required to wire up the agent properly - the first time is always an adventure. It’s ok if you’re having trouble doing this, it can take a while to figure out.
We hit Enter, wait a few minutes, and we have a PR!


Do you FEEL the robot uprising coming!?
This is perfect: The change is super small and can be validated in literally under 5 minutes. The agent even told us what it did, so we could validate it ourselves. We had our first merge!
Give yourself a pat on the back.
To make it repeatable, I threw it on an automation solution. In our case it was Cursor Cloud, though it honestly doesn’t matter which one we used - you can use ChatGPT Schedules, Claude Routines, OpenHands, whatever you want, they’re all the same. The important bit is that they run an llm in a similar way to a local coding agent, and you can give them access to the shared resources (like github, linear, slack).
Set it to run every morning. For the following week, I woke up, and saw an unread message in this channel:

I opened it, read the PR description & change, saw the green CI checks, and did some code searches. Satisfied, I said thanks to the removed code, approved, merged, and moved on. On Saturday morning I did it from bed.
After a week, we had 7 merged PRs. After two weeks, we had 14. Below is a graph of automatically opened PRs (blue) and how many were closed without merging (red) over the span of a few weeks.
.png)
Nice.
MORE!
If you’re following at home, take time to pat yourself on the back and celebrate for a moment. You’ve done the first step! Even if you stop right now, you’ve already done something.
You can go and show the rest of the company what you have so far, which is the start of a reliable way to ship automatic changes which take minimal attention. You’ve shipped 14 PRs and removed something like >200 lines of code. Transhumanism, here we come!
This is where you find your allies: Who’s impressed by what you’ve done and wants to join?
I gathered the impressed peeps, and ideated on what’s next. The guidelines were the same: Automations that could be wholly owned by one person, merge-able without review or follow-up (because who has time to argue with robots).
Thinking about it more deeply, we see that when we talk about more difficult workflows, we had two directions we could go in:
- Problem difficulty: The scope of the problem escalates from “we’re going to delete 10 lines of code” to “we’re going to work across a module”
- Operations difficulty: Reasoning about the problem doesn’t just require code, it also requires something else that’s not trivially available - like live logs, or a database field, or a human-in-the-loop
We ended up going up the problem difficulty route, and picked structured logging.
Medium project: Structured logging
Our problem was another boring mechanical problem that nobody wanted to properly tackle. We have a structured logging library that behaves like so:
logging.info("Did a thing", count=len(whatever))But because of old code, we still had callsites that didn’t take advantage, and did calls like:
logging.info(f"Did a thing over {len(whatever)} items")Told you it was boring. We decided that we’re going to encode this migration as an agent skill in the same repo as the source code (you’ll see why a little later), and to track management as a markdown file.
For task management, we committed a file that looks something like this, but with the table filled in:
# Structured-logging migration
Tracker for converting every workspace project to fully structured logging.
## Per-project table
Pick the top row whose `Status` is empty; skip `done` / `failed`.
| Project | Status | Ticket |
|---------+--------+--------|
| foo | | ... |
| bar | | ... |
| baz | | ... |
## Migration reports
_(One entry per migrated project, newest first. See the src-layout tracker's
reports for the expected shape: what was converted, judgment calls, tests
touched, validation run, difficulties.)_
We scope the PR per some unit-of-work (UoW). For us it’s a pyproject, for you it can be a subdirectory. Whatever, as long as the UoW is manageable to review. This is a poor man’s project management. We could have encoded this in our ticketing system, but we didn’t, for absolutely no reason at all other than we just saw this pattern and wanted to experiment.
As for the skill, we wondered how we would do this migration, and then specified the steps in way too much detail. I’m including it fully to give you a sense that it’s possible to write down in a reasonable time frame. It took us about an hour with the help of a local coding agent and some healthy arguing.
We wrote the following into .claude/skills/python-structured-logging-migration/SKILL.md - it’s a code-dump, feel free to skim AND to steal:
---
name: python-structured-logging-migration
description: Use when converting a Python project in this monorepo to structured logging.
---
# Python structured-logging migration
This skill converts one workspace project at a time to fully
structured logging, then **switches the ruff `G` rule on for that
project** so it can never regress. The repo-wide goal: every project
clean, every per-project `G` exemption removed, `G` guarding the whole
tree.
## The conversion spec (structured / idiomatic)
| Pattern | Bad | Good |
|---|---|---|
| Bare variable | `f"Processing {tenant_id}"` | `logger.info("Processing", tenant_id=str(tenant_id))` |
| Attribute / subscript | `f"{row.vuln_id}"` / `f"{d['id']}"` | `vuln_id=row.vuln_id` / `id=d["id"]` (readable kwarg name) |
[ ... cut for brevity ... ]
## The migration, step by step
### 1. Identify the next project
Read `docs/logging/structured-logging-migration.md`. It holds the entire state of this effort. Pick the next project from the table by its `Status` column — empty Status = not started; skip rows marked `done` or `failed`.
Don't second-guess the table. **Do not `git log`, `ls`, or grep around for work — USE THE TABLE.** One project per run.
### 2. Turn the guard ON for this project
Delete this project's `["G"]` line from `[tool.ruff.lint.per-file-ignores]` in the root `pyproject.toml`. This is the step that makes the migration permanent — **and it must happen first**, because while the exemption is in place ruff suppresses `G` for this project, so the worklist command in the next step would report nothing.
### 3. List the violations
```bash
uv run ruff check --select G --no-cache --output-format=concise <project-path>
```
This is your exact worklist for the project (the exemption was removed in step 2, so violations now surface). Read each offending line (and its continuation lines for multi-line calls) before converting.
### 4. Convert
Apply the conversion spec above to every flagged call. Edit only log calls (and any test that asserts on log *message text* — see step 7). Match the surrounding code's style.
### 5. Contextualize shared flow fields
While converting, look for fields that repeat across many calls in the same flow — `tenant_id`, `project_id`, `integration_id`, a batch/run id. The shared logger's `contextualize()` (context-manager form) carries them into **every** log call inside the block, so you set them once instead of passing them on each call.
Keep the context **lean** — only fields genuinely relevant to *all* calls in the flow; don't dump a pile of per-call fields into it. This is a judgment step: lift the obvious flow-wide ids when a conversion makes them repetitive, but don't restructure the code chasing it. `str()`-coerce ids here too.
### 6. Verify ruff
```bash
uv run ruff check --select G <project-path> # must be 0
uv run ruff format <project-path>
uv run ruff check <project-path> # whole-tree-clean for this project
```
If `ruff check` reports only import-order churn, `ruff check --fix <changed files>` then re-run format + check.
### 7. Run tests
```bash
# from the project directory
uv run pytest
```
Watch for tests that assert on **log message text** — those break when the static message changes, and must be updated to assert the new message / a structured field (not silently skipped). Behavior is otherwise unchanged, so most suites pass untouched.
### 8. Update state and write a report
In `docs/logging/structured-logging-migration.md`: set the project's `Status` to `done` (or `failed`), and add a short migration report under **Migration reports** covering: what was converted (rough counts per `G` code), any judgment calls (tricky format-spec / exception cases, contextualize lifts), tests touched, validation run, and difficulties.
### 9. Introspect and improve
Edit **this SKILL.md** with any precise lesson learned this run (a conversion edge case, a test-assertion gotcha, a project-specific quirk) so the next run is smoother. Keep it terse.
### 10. Open a PR
Stage narrowly (the converted files + the pyproject exemption removal + the doc update — nothing else). Conventional-commit subject:
```text
refactor(<project>): structured logging
```
Body: embed the step-8 report (½ screen — one paragraph + counts + validation; detail belongs in the commit, not a wall of text). If a tracking ticket exists for this project, add `Resolves <TICKET-ID>.` as the opening line. Add the `ai-agent` label. Open via the GitHub MCP.
### 11. Babysit the PR
Watch the PR checks (all of it) to green. The key signal: the `G` rule now runs unexempted against this project, so a quality-check failure here means a missed conversion — fix it, don't re-add the exemption.
### 12. Slack message
Extract the PR URL and post to `#agent-prs`: "Structured-logging migration for `<project>`: <PR_URL>".
## Troubleshooting
- **One project per run.** If a fix forces edits in another project, FAIL EARLY and record why in the tracker.
- **Don't re-add an exemption to make CI pass.** A red quality-check after migration means a real missed/incorrect conversion. Fix the conversion.ok, so it’s a lot more complex than the dead code removal. There are some highlights to be had here:
This skill is highly specific. Agents are scary because you’re not sure what they would do. But we knew what we would do, so we codified it to a high degree to reduce the number of possibilities. Less possibilities = less review time = more merges.
There is very little left to taste. I omitted some boring implementation details, but in the replacements table we’ve documented exactly what we want done: What patterns constitute bad code, what’s good code, how to look for it, how to verify, etc.
Like, check this line out, oh my god:
Don't second-guess the table. Do notgit log,ls, orgreparound for work — USE THE TABLE. One project per run.
This line wasn’t there at first. We added the shouting because during the first runs, the agent went around the repo looking for work, when clearly we wrote down where to find it and in what order. We've since added “prompt engineering” to our CV.
We’re also hand-holding the agent through a PR. We wanted to have a clear marker (the ai-agent label) so we can track how many are merged/closed and how long they take. We also want the agent to loop and respond to CI checks - asking it to do so turned out to be enough.
One of my favourites here is self-improvement. Check out step 9:
Edit **this SKILL.md** with any precise lesson learned this run (a conversion edge case, a test-assertion gotcha, a project-specific quirk) so the next run is smoother. Keep it terse.This means that subsequent runs are much smoother. Pretty much every run documented another edge case or shortened future runs for agents. This paid off very nicely. An example of an improvement that was made (feel free to skim):
+- A per-project uv run pytest can fail on a test that imports a sibling workspace project not declared as a dependency (ModuleNotFoundError: No module named '<sibling_pkg>'). Per-project sync only installs that project + its declared deps, so the undeclared sibling is pruned — but CI runs the full workspace, so the import resolves there. This is not your diff (and not fixed by adding the sibling as a dependency — that's unrelated scope): confirm the project's own CI test job is green on main HEAD (gh api …/commits/<main_sha>/check-runs), then reproduce CI's env with a root uv sync --all-packages and re-run via uv run --no-sync pytest from the project dir (a plain uv run pytest re-prunes the siblings).Whoops! This was a great insight that pointed us to a problem we hadn’t realised we had. Good catch. Needs to be more terse, but that’s ok, it can be followed-up.
But most importantly - this skill wasn’t perfect. If I wrote it now it would’ve been better. But we were ok with this: It’s easy to argue about how we should write the skill, should we embed the table in the skill itself vs. an external file, should we use an external task manager, should we outsource some elements to a script, should this be claude-specific vs. using an agent package manager, wait how come we don’t have loop engineering here, aaaaaagghhh so many choices!
Just…start. Do a thing. Worse is better, perfect is the enemy of good enough.
This produced larger PRs:

But some PRs were still small:

The cadence was similar: Wake up, review, merge.
What I like best about this one is how additive it was. We didn’t do anything really fancy, we just did what we’ve done previously with a few small twists outlined above.
Boring work, completed by something else, that can be reviewed quickly. Nice!
MORE!!!
The beast of automation must feed. Example automations include (ideate with your own peeps and figure out what makes sense for you):
- Changing subproject file layout
- Deprecating and removing unused feature flags
- Aligning internal documentation to their code
- Auto-fixing flaky CI tests
- Moving from one database driver to another (typeorm → drizzle, psycopg2 → psycopg3, whatever)
We created more. From a few months ago:

We’ve also made a mistake.
Bottleneck is you
We started with one daily PR that took us 5 minutes to merge. Then it was two daily PRs that took 10 minutes total. As we added more automations, it led to more PRs.

This was ok until we hit a bottleneck. One of the automations produced large PRs that took around an hour to review. Another automation didn’t produce satisfactory results on the first-shot, requiring a review round.
Both meant that the human attention remained the limiting factor. If our goal was to introduce a dark factory, I would have delved into how we solved it with automated reviews and - but this is not one of our goals at the moment. Not all bottlenecks need to be eliminated, but instead planned around.
The most obvious “solution” is defining ownership. If I own three of these automations, it’s just part of my day to make sure they’re done. If I can’t do them, then I delegate, or change the automation to reduce its cadence, or drop and move on.
Other failures
These “let’s run a prompt on a schedule” solutions are, shockingly, still made in software. This means that they require maintenance and gentle applications of lubricant.
In the first automation solution we used, some runs just…failed. The environments didn’t boot up, so they didn’t do their thing, and got stuck in some limbo. It needed someone to go into the interface and click “cancel”. We didn’t notice for a couple of days because, well, it wasn’t important enough - see ownership above. Because these solutions are still kinda new, they’re not that good when it comes to observability and error reporting.
Another failure we’ve seen is around integrations. One of our automations needed to look something up in Linear, and its Linear connection expired, so it started writing confusing nonsense results. Again, the solution was simple: Notice it’s wrong, go to the UI, fix the problem, but it wasn’t that obvious.
It’s software all the way down - the bad parts as well.
Debugging these flows is kind of…weird. It’s both easier and harder than debugging software. You know the endless debate about the difference between “programming” and “scripting”?
This kind of feels similar. If you ship a prompt or agent to production, you do an eval, right? Would you connect an automation to an eval framework? Probably not, right? It feels excessive! Then that means you don’t know if you’re improving or degrading the automation - just like checking CI scripts.
The next step
Following along at home? At this point, you’ve got at least 3 PRs being created automatically. They’re simple PRs, or they started complex and got whittled down because nobody could care enough to attend to the complex ones. Your graph may look something like this (true story; blue = merged, red = closed):

Once you’ve got this running repeatedly, decide how you want to increase complexity.
One direction is problem complexity. You’ve written automations to create small PRs (like dead code removal), move up to creating medium-sized PRs (like migrating APIs). Alternatively, use them to correlate multiple data sources aside from source code (like given an exception and stack trace, investigate to find possible RCAs). This requires figuring out how to make agents write shippable code, which is a difficult problem.
Another direction is operational complexity. Accessing source code is something every cloud-agent platform can do out of the box - but how about accessing your log storage? If your organization isn’t AI-enabled enough yet, figuring out how to give access to agents is hard. What identity do they hold, what permissions do they have? If it’s self-hosted Grafana instance, how should the agents even access them? Do you maybe prefer having self-hosted runners, rather than on laptops or vendors? What does it mean about your security posture? All great questions.
Yet another direction is political complexity. You’re good mates with one team, but less familiar with another team who you know isn’t the most AI-friendly. Figuring out how to sell your solution is hard! But at this point you have at least a few successes: Publish them hard, find people to talk your achievements up, then go forth and conquer.
This may require being the annoying AI guy. I’m sorry. But this is like any other bottom-up organizational change: If you don’t believe it will work, and you don’t show up with results, you’re not going to get traction.
Coda
Look, maybe humanity is doomed in 5 years and we’ll all be paperclips. Maybe because of compute capture the LLM companies will turn multi-planetary and enact UBI, rendering this entire article pointless because we’ll be lounging in the moons of Titan. I don’t know.
What I do know is that until we high-five robot dogs, we can continue doing what engineering organizations have done previously: Identify problems, break them down into solutions, and apply them. Coding agents are one form of solution to a set of problems.
It’s daunting to think about falling behind while all the rest of the world is doing dark factories when you haven’t even done one automated PR. That’s where we started, taking this abstract far-fetched notion and making it more concrete.
The goal is not to solve all of the problems. The goal is to solve the next problem faster. To find a worthwhile problem that makes people even remotely excited, even if only because you’re doing something they haven’t had time to do.
Your goal is to start with one PR, and continue from there, one PR at a time.
Don’t fall into the hype. Take a deep breath. You can do this.


