All posts
Fullstack

Writing code with a model, and shipping one

· 6 min read

An agent writing code for me and a model whose output lands in front of a user are two different jobs, and the second one is mostly ordinary engineering around the call.

Writing code with a model, and shipping a feature whose output reaches someone who never asked for a model, sound like the same skill. They aren't. I've done both, and only one of them is a way of working. The first is a way of working. The second is a product decision with a bill attached, and almost all of the code it needs is the sort of code you'd write around any flaky, expensive, non-deterministic third party.

This post is the evidence for that, from four places where I've done it: Qrder's AI photo studio, Munitron's lesson generator, Manarway's narration, and the assistant on this site. Three of those have posts of their own on this blog, so I'm not re-explaining how they work. I'm pointing at the parts that have nothing to do with the model.

The provider is a detail

Qrder's photo studio generates food photography from a menu item's name and description. Behind it is a provider interface with four implementations, picked by one environment variable, falling back to a mock that draws gradient plates when no keys are set.

What that buys is a model I can replace in an afternoon. What it doesn't buy is any protection from the provider being slow, which needs its own code:

const controller = new AbortController()
const timeout = setTimeout(() => controller.abort(), GENERATION_TIMEOUT_MS)

try {
  const images = await provider.generate({
    prompt: built.prompt,
    negativePrompt: built.negativePrompt,
    variations,
    aspectRatio: built.aspectRatio,
    signal: controller.signal,
  })

  // …
} catch (err) {
  if (controller.signal.aborted) {
    throw new ImageGenerationError(
      'The image provider took too long to respond. Please try again.',
      { provider: provider.id, status: 504, retryable: true, cause: err },
    )
  }
  throw err
} finally {
  clearTimeout(timeout)
}

Two minutes by default, from an environment variable. Image generation in that studio routinely runs twenty to sixty seconds, so the usual read timeouts are no help: the request has to be allowed to take a minute and still not be allowed to take forever. The abort signal goes down into whichever provider is configured, so cancelling is the provider's problem rather than a promise left running.

A failure a café owner can act on

The thing a model's failure modes have in common is that they're all meaningless to the person looking at the screen. So they get translated:

/** Never leak provider internals (or key hints) to the browser. */
function providerMessage(err: ImageGenerationError): string {
  if (err.status === 503) {
    return 'AI photo generation is not configured yet. Add a provider API key to continue.'
  }
  if (err.status === 429) {
    return 'The image provider is rate limiting us right now. Please try again shortly.'
  }
  if (err.status === 504) {
    return 'The image provider took too long to respond. Please try again.'
  }
  return 'The image provider could not generate photos right now. Please try again.'
}

Four messages, and the distinction between them is whether waiting will help. "Not configured" means stop and go and add a key. Rate limited and timed out both mean wait. The default means it's broken and we don't know why. A provider error string pasted through to the browser would have told a café owner none of that, and might have told them the shape of my API key.

That function lives in a route handler of just under four hundred lines. One of those lines calls the model.

The model's output is untrusted input

Munitron generates a short chat-style lesson from a phishing scenario: five or six turns, each with a trainer message and a few answers to pick from. The generated turns are then edited by an admin and eventually shown to an employee, which means they have to have a shape.

So there's a schema, and it runs before anything else touches the reply:

const turnResponseSchema = z.object({
  id: z.string(),
  label: z.string(),
  isCorrect: z.boolean().nullable(),
});

const trainingTurnSchema = z.object({
  id: z.string(),
  messages: z.array(turnMessageSchema),
  responses: z.array(turnResponseSchema),
});

export const trainingTurnsOutputSchema = z.object({
  turns: z.array(trainingTurnSchema),
});

A shape, and a repair

isCorrect being nullable is the interesting part. The prompt asks for exactly one correct answer per turn, and a prompt is a request. A model that returns null where a boolean was asked for has still returned something usable, so the schema allows it and the service fixes it afterwards by marking the intended answer correct. A teammate later narrowed that condition and added a shuffle so the correct answer isn't always in the same position.

Validating first also means the failure is one error rather than four screens of undefined. An empty turns array is thrown explicitly, because it parses fine and is useless.

When the model is one step in a pipeline

Manarway's narration gives every word of a story a timestamp. Almost none of that is a model. The timings come from a text-to-speech alignment folded into words, the paragraphs are matched to those words by position, and the offsets between clips are arithmetic. One small piece is a guess, and it's the one that can't be computed: which voice a named speaker should get.

async function detectSpeakerGenders(names: string[]): Promise<Record<string, 'male' | 'female'>> {
  try {
    // …
    const jsonMatch = text.match(/\{[^}]+\}/);
    if (!jsonMatch) return {};
    return JSON.parse(jsonMatch[0]);
  } catch {
    Logger.warn('Gender detection failed, defaulting to male voice');
    return {};
  }
}

The call in the middle, since it moved onto a shared agent, is a teammate's. The regex and the fallback are mine. Both say the same thing about how much I trust the reply: the JSON is dug out of the text with a pattern rather than parsed from it, and every failure path returns an empty object, which the caller reads as "use the default voice". The narration still generates. A teacher gets one voice instead of two and nobody sees a stack trace.

That's the test I'd apply to any model call in a pipeline. If it fails, does the feature fail, or does it get slightly worse? Word translations in the same service work the same way: the failure is logged and the transcript is built without them.

Somebody pays per call

Using an agent costs me a subscription. Shipping a feature costs per request, on traffic I don't control, and that lands in the code.

Two limits, one of them durable

Qrder's studio has two limits rather than one. An in-process burst limiter stops a stuck button hammering the provider within a few seconds, and a Postgres function enforces the real monthly quota per café, atomically, because the burst limiter lives in one instance's memory and the quota can't. The quota is consumed before the provider is called and refunded if the call fails, the upload fails, or the row can't be written. A café that got no photos has not spent a generation.

Tokens, because there's no cache

The assistant on this site pays in tokens instead:

/**
 * Groq allows this account 8,000 tokens a minute, and the notes above already
 * cost ~2.9k per message. A whole post would push a follow-up question over
 * the limit, so a page contributes an outline instead: capped here.
 */
const PAGE_BUDGET = 6000;
const PARAGRAPH_BUDGET = 420;

Because that account has no prompt caching, the grounding notes go out in full with every message, so their length is a per-message cost. That's why a post sent as context becomes its headings and the first paragraph under each, with code blocks stripped, rather than the post itself. The feature exists in the shape it does because of a number in a pricing page.

What the prompt is actually for

The prompts in all of this are short and dull, and they're mostly rules about what not to do. Munitron's generator has a language rule marked highest priority, so an Arabic scenario doesn't come back with English headings in it. The assistant here is told that its notes are everything it can state as fact, that it must never invent a project or a date, and that it must never agree to a rate or a start date on my behalf.

None of that is clever, and none of it is reliable on its own. A prompt is a request, and the code around it is what holds. The Zod schema is why a malformed lesson can't reach an admin. The quota is why a model can't spend money that isn't there. The prompt only narrows what gets asked for.

What I'd change

Qrder's quota refund is careful and its burst-limit refund isn't. When a generation fails, the route calls releaseRateLimit, which does bucket.hits.pop(): it removes the newest timestamp in that café's bucket, not necessarily the one this request added. Two staff members generating at once, one failing, and the failing request hands back the other's slot. The fix is to have checkRateLimit return the timestamp it recorded and release that exact value, which is a few lines. It's a small bug, and I like it as an example, because it has nothing to do with AI at all.

The second one is a gap in the validation rather than a bug in it. Munitron's schema checks that a generated lesson has turns, messages and answers. For a turn that shows the phishing page with its red flags outlined, the message content is a JSON string, and the schema only checks that it is a string. Nothing parses it, and nothing checks that the text a highlight is anchored to appears in the page it points at. An anchor that matches nothing is exactly the failure an employee would notice, and it's the one thing the validation doesn't look at. I'd parse that content in the schema and drop any annotation whose anchor text isn't present in the artifact, at generation time, before an admin ever sees it.

Was this any good?

NextA day of handing work to an agent