LLMs, efficiently
Memory, compute, and bandwidth are the only three budgets there are. Every technique for training and serving large language models — checkpointing, flash attention, GQA, quantization, ZeRO, tensor and pipeline parallelism, speculative decoding — spends one to buy back another. This is that ledger, derived rather than listed.