Serving, Cost & Latency

LLM performance engineering: token economics, caching strategies, model routing, batching, self-hosting vs API trade-offs, and keeping p95 latency and unit costs under control.

This module hasn't shipped yet. Lessons land in order; check the track overview for what's live today.