Reserving 128k Context Costs Nothing. Filling It Costs Up to 70 %.
Two RTX 3090s gave me opposite answers about what long context costs. Both were right: reserving a 128k KV cache is nearly free, while actually filling it costs 25 to 70 percent of throughput — with nothing offloaded to the CPU. Most context benchmarks, including my own earlier one, measure the first and sound like the second.