Skip to content
X · @togethercompute · X / Twitter

Re 1/ New deep dive: Configuring Dedicated Model Inference. Three primitives: endpoints, deployments, configs. Capacity-aware routing distributes traf…

Re 1/ New deep dive: Configuring Dedicated Model Inference.Three primitives: endpoints, deployments, configs. Capacity-aware routing distributes traffic by capacity (weight × ready_replicas), not fixed percentages, so a deployment's share tracks its replica count automatically.We tested it on a live endpoint with tagge