Skip to content
arXiv cs.LG · Papers

Measuring and Reducing WebGPU Dispatch Overhead for LLM Inference

arXiv:2608.08730v2 Announce Type: replace Abstract: Large Language Models are deployed to multiple types of environments, from internet browsers to edge devices, and WebGPU serves as a modern cross-platform standard. The engines for browser-based LLM inference have proliferated, yet the overhead of WebGPU per-operation