Skip to content
r/LocalLLaMA · Communities

ModelExpress: Distributing Model Artifacts at the Speed of Light – NVIDIA Technical Blog

We cut DeepSeek-V4 Pro startup from 8 minutes to under 2 minutes by moving weights over the fastest path to GPU memory with GPU-to-GPU RDMA. This was achieved using NVIDIA ModelExpress (MX), the weight distribution and cache management service in NVIDIA Dynamo, and this same approach speeds up both inference and RL pos