Skip to content
X · @togethercompute · X / Twitter

We shipped a lot in this update to Dedicated Model Inference. One part worth understanding is the resource model underneath it. Requests are allocated…

We shipped a lot in this update to Dedicated Model Inference. One part worth understanding is the resource model underneath it.Requests are allocated across deployments by capacity, computed per ready replica, not by fixed percentages. Scaling changes a deployment's share automatically, and replicas that are not ready