close

Linden Li on X: "Important, very much needed update to inference performance benchmarks from @SemiAnalysis_ ! Fixed token-in, token-out is a very different inference workload from what we observe in practice (both for RL training and production inference). The inference optimization story (engine, hardware, etc) is very different for these workloads, and the benchmarking north star should reflect that."

Important, very much needed update to inference performance benchmarks from @SemiAnalysis_ ! Fixed token-in, token-out is a very different inference workload from what we observe in practice (both for RL training and production inference). The inference optimization story (engine, hardware, etc) is very different for these workloads, and the benchmarking north star should reflect that.
3:24 AM · Aug 25, 20267.6KViews