Practical writing on shipping AI systems, engineering craft, and the choices that shape modern software teams. From our team to yours.
Beyond eyeballing outputs — the small, boring metrics that catch regressions before your users do.
Most RAG problems are retrieval problems in disguise — how to spot them and what to fix first.
Streaming, citations, refusals, and repair — the new primitives every AI product interface needs.
Batching, quantization, and routing — where the real savings hide once you get past the sticker price.
How to make partial output feel intentional, not glitchy — the small state machines behind a good streaming UI.
The one habit that keeps front-end, back-end, and mobile teams from stepping on each other for a whole quarter.
A short set of questions that tells you which one your problem actually needs — and when you need both.
Latency budgets, drift detection, and rollback plans — an operations checklist tuned for LLM-backed services.
One or two pieces per month. No spam.