Serving, Cost & Latency
LLM performance engineering: token economics, caching strategies, model routing, batching, self-hosting vs API trade-offs, and keeping p95 latency and unit costs under control.
This module hasn't shipped yet. Modules land in order, roughly one per week; the newsletter below announces each one. The track overview shows what's live today.
Closest live reading while you wait: