Category archive
Serving a model is a scheduling problem with a budget attached.
Latency, throughput and cost in a language model system are set by the parts around the weights: the scheduler, the cache, the gateway, and the memory budget every request competes for. These articles measure those parts and treat the model itself as the one input that is already fixed.
Articles in Inference Systems
Transformer Inference, From HTTP Request to the Next Token
A systems-level tour of prefill, decode, batching, memory movement, and the scheduling decisions that determine real-world LLM latency.
KV Cache Engineering for LLM Inference
How to size, allocate, reuse, compress, and evict the state that makes autoregressive decoding practical.
Semantic Caching Without Serving the Wrong Answer Faster
A production design for similarity keys, freshness, authorization, invalidation, and the economics that determine when semantic reuse is safe.
Speculative Decoding Without Hand-Waving
Draft models, acceptance math, tree proposals, and the operational details that decide whether speculation makes serving faster.
The LLM Gateway Is a Policy Engine, Not a Proxy
Designing the routing, budgets, identity, resilience, and evidence layer between products and a changing model portfolio.
Fine-Tuning with LoRA in Production
When fine-tuning is justified, how low-rank adapters work, and what it takes to evaluate, serve, and update them safely.
LLM Cost Engineering: From Token Prices to Unit Economics
Modeling full request cost, eliminating waste, forecasting margins, and optimizing without disguising quality regressions.
Apply the category
Put this against a real system.
If a decision in Inference Systems is in front of you right now, the fastest version of this is a call: bring the architecture, the failure you are seeing, and the constraint you cannot move.
Discuss the systemIn the age of AI
The advantage was never the model. It's knowing what to build with it — and having a team that can actually ship it.
That's the part I help with: finding where AI genuinely makes your business faster, deciding what's worth building, and standing behind it once it's live.
Four offices, one very full passport
Every dot on this map is a conversation I still remember.
- Where I've spoken
- Office







































