Field notes · 30 AI engineering articles
AI Engineering Articles
From the review desk
Production AI systems: inference and serving, retrieval, agents, evaluation and AI security, multimodal and edge pipelines, and AI-native product architecture.
Transformer Inference, From HTTP Request to the Next Token
A systems-level tour of prefill, decode, batching, memory movement, and the scheduling decisions that determine real-world LLM latency.
All articles
KV Cache Engineering for LLM Inference
How to size, allocate, reuse, compress, and evict the state that makes autoregressive decoding practical.
Semantic Caching Without Serving the Wrong Answer Faster
A production design for similarity keys, freshness, authorization, invalidation, and the economics that determine when semantic reuse is safe.
Speculative Decoding Without Hand-Waving
Draft models, acceptance math, tree proposals, and the operational details that decide whether speculation makes serving faster.
Synthetic Evaluation Data That Finds Real Failures
How to generate, filter, diversify, and maintain synthetic cases without turning an evaluation suite into a mirror of its generator.
The LLM Gateway Is a Policy Engine, Not a Proxy
Designing the routing, budgets, identity, resilience, and evidence layer between products and a changing model portfolio.
Privacy-Preserving AI Is a Dataflow Architecture
Purpose limitation, minimization, isolation, retention, redaction, and verifiable deletion across retrieval, models, tools, traces, and feedback loops.
Incident Response for Systems That Can Be Wrong Fluently
A response playbook for quality regressions, prompt attacks, retrieval contamination, runaway agents, cost spikes, and provider failures.
Production RAG Is an Evidence Pipeline, Not a Vector Search
Designing retrieval-augmented generation around ingestion quality, query planning, evidence assembly, and verifiable answers.
Hybrid Lexical and Vector Search
A practical design for candidate generation, score normalization, rank fusion, metadata filters, and retrieval evaluation.
Rerankers: The Quality Layer Between Search and Generation
How cross-encoders, late interaction, listwise ranking, and calibration turn noisy candidates into usable evidence.
Engineering the Agent Tool Loop
A concrete architecture for tool selection, state transitions, permissions, budgets, and recovery in production agents.
Durable Execution for Agents That Outlive a Web Request
Persisting checkpoints, replaying deterministically, handling human approval, and surviving crashes without duplicating side effects.
In the age of AI
The advantage was never the model. It's knowing what to build with it — and having a team that can actually ship it.
That's the part I help with: finding where AI genuinely makes your business faster, deciding what's worth building, and standing behind it once it's live.
Four offices, one very full passport
Every dot on this map is a conversation I still remember.
- Where I've spoken
- Office







































