Skip to content
r/MachineLearning · Communities

Cloud-vLLM Benchmark Differences [R]

Does anyone know of any evidence/forum/paper analyzing benchmark result differences between cloud inference platforms (togetherai) and running models locally with vLLM under greedy decoding? submitted by /u/No_Cardiologist7609 [link] [comments]