All posts
Fullstack

Nothing merges unread

· 6 min read

I hand Claude Code whole features, and the only rule I hold without exception is that no diff reaches a branch I'd merge until I've read it, whoever wrote it.

The question I get about working with an agent is always some version of where the line is, as though there's a category of work I've kept for myself and a category I've handed over. I hand Claude Code whole features, at Akera and on my own products, so the line people expect to find is somewhere in the middle of that.

There isn't. The thing I hold without exception isn't a kind of work at all. Nothing merges until I've read it, whoever wrote it.

What I hand over

Mechanical work

Scaffolding, CRUD, refactors, renames, fixtures, migrations, translation files. This is work where the answer is already decided and the only cost is the typing, and it's also where I'm least reliable. A UI change that has to land in three language files, in the right nesting, with the generated types regenerated afterwards, is exactly the task I'd leave one file short of done while believing it was finished.

Investigation

Finding where something lives in a monorepo I didn't write all of. Reproducing a bug from a report. Explaining why something broke, which is usually three files away from where it surfaced. An agent that can read the whole repository at once is better at this than I am, and the output is a claim I can check against the code rather than a change I have to trust.

Writing

Commit messages, documentation, pull request descriptions. I also hand it my own diffs to review before anyone else sees them, which catches the dull things I stop seeing after an hour in the same file.

The line is not a kind of work

Both of those lists are about who does the first draft. Neither says anything about what gets merged.

An agent diff and a diff I wrote at the end of a long day fail in different ways. Mine tends to be narrow and slightly wrong about a thing I'd stopped questioning. The agent's tends to be broad and confidently finished, which is harder to read honestly, because a diff that looks complete invites you to skim it. Both get read.

Keeping that practical is mostly about size. The rules file I keep in Munitron tells the agent to stage files by name rather than stage everything, and not to commit at all until I say so. That one instruction is the difference between a diff I can hold in my head and one I'd be scrolling through at speed.

Four ways I check

Running the thing

Clicking through the feature on something resembling the real device finds what reading doesn't. In Roudhati, two flows are scripted so they get clicked every time, on a phone viewport in Arabic:

test.describe('recording a payment', () => {
  test('reaches the fee grid from home and records a payment end to end', async ({ page }) => {
    await signIn(page);

    // Tap one: the primary action on the home screen.
    await page.getByRole('link', { name: 'تسجيل خلاص' }).first().click();
    await expect(page).toHaveURL(/\/fees/);

    // …

    // The sheet must state what the money will cover before it is committed.
    await expect(page.getByText('سيغطي هذا المبلغ')).toBeVisible();
    await expect(page.getByRole('button', { name: 'تأكيد' })).toBeEnabled();
    // …

That flow caught a real defect, and the commit message says what it was:

The payment flow caught a real bug. The confirmation sheet only fetched its
preview on the amount field's blur event, so opening a cell and pressing تأكيد
straight away recorded money without the director ever seeing which months it
would cover — precisely what section 7 of the brief says must not happen. The
preview now loads when the sheet opens and debounces on every change, with the
in-flight request aborted by the next keystroke.

Reading that code would not have found it. The handler was there, the request was correct, the preview rendered. The bug was in which event woke it up: tap a cell, tap confirm, and the amount field is never focused, so it never blurs. You find that by tapping.

Tests

The rules file names the tests that have to exist, and they're the places where being wrong costs money or trust rather than looking untidy:

## Tests that are required, not optional

- Money math: allocation, carry-forward, partial payments, balances
  (`apps/api/src/modules/fees/allocation.test.ts`)
- Tenant isolation: a user of tenant A cannot read tenant B's children
- Arabic search normalisation
- Phone validation
- Playwright, two flows only: record a payment, mark attendance
  (`apps/web/e2e/`, run with `pnpm test:e2e`)

<!-- … -->

Where a rule is structural, the test has to cross all of the structure, or it only proves the top layer agrees with itself:

/**
 * The test the brief asks for: a user of tenant A cannot read tenant B's data.
 *
 * It runs against the real HTTP surface with real tokens, because the thing
 * being tested is the whole chain — guard, interceptor, client extension and
 * the RLS policies underneath — not any one layer's unit behaviour.
 * // …
 */

The rules file doing the checking

The non-negotiables in that file end with a sentence that's there for review, not for writing: violating one is a bug, not a style preference. When a diff uses a physical margin in a right-to-left app, or puts a float anywhere near an amount, that isn't feedback I weigh against how much work the change was. It goes back.

Two of those rules are enforced by the lint config, so they come back as errors before I see them at all. The rest I check by reading, which works because the list is six items long and I know it by heart.

The committed rules file in Manarway, which a teammate wrote, ends on the same idea from the other direction: run the typechecks after changing code, separate real errors from warnings that were already there, and say plainly when something couldn't be verified locally instead of leaving it implied. An agent that reports a gap is far more useful than one that reports success.

Reading the diff

Then the diff itself, in the order a reviewer would read it, with the other three checks treated as evidence rather than conclusions. Tests passing means the tests passed. It doesn't mean the change is the one I asked for, or that it's the smallest version of itself, or that the thing it touched on the way through is still true.

The parts it didn't say it did

The summary at the end of a run is a summary. Something chose what went in it, and what gets left out isn't random: it's the parts that were incidental to the task, which is also where the surprises live.

So the paragraph tells me where to start, and then I go looking for what isn't in it. A file in the diff that the summary never mentions. A test whose expectation changed rather than whose code changed. A type that got wider. A suppression comment. A default that moved. A migration that ran against my local database while I was reading. None of these are dishonest, and all of them are easy to carry along in a change that was mostly about something else.

Landing it in a shared repo

In the Akera repos the process does some of this for me. Work goes onto a branch and into a pull request, and someone else merges it, so every diff has a second reader by the time it reaches main. CI installs with a frozen lockfile and then runs the dependency check, lint, typecheck, tests and a build on every pull request, plus the affected targets for the projects a branch actually touched.

A teammate added a per-commit file limit on top of that, with a job that labels and comments on any pull request whose commits go over it. Its own comment gives the reason: smaller commits keep reviews faster and safer. That's the same constraint as my local rule about staging files by name, written as a policy instead of a habit. My commits in those repos are one concern each with a scope on the front, down to things like chore(auth): drop dead currentLanguage font ternary, which is a size I can read and someone else can read after me.

The agent's own configuration stays out of the repository. In Munitron, both CLAUDE.md and .claude are listed in .gitignore, so the rules I give the agent there are mine and don't speak for anyone else. The committed one in Manarway is the team's, written for agents and engineers in the same breath, and keeping those two apart matters.

What I'd change

Roudhati has no CI. There's a pnpm db:check-drift script, added with the migration fix that made it necessary, and the commit message that introduced it says outright that it belongs in CI. It still isn't there, because there's no workflow in the repository at all.

So three of my four checks currently depend on me remembering to run them on my own machine, which is a weak place to put a rule after writing a whole file about not relying on anyone's memory. A workflow running typecheck, lint, tests and the drift check on every push would settle the mechanical three for good and leave me with the one that was always going to be mine.

Was this any good?

NextThe rules file I write before any code