X · @togethercompute
· X / Twitter
Re 1/ New deep dive: Configuring Dedicated Model Inference. Three primitives: endpoints, deployments, configs. Capacity-aware routing distributes traf…
Re 1/ New deep dive: Configuring Dedicated Model Inference.Three primitives: endpoints, deployments, configs. Capacity-aware routing distributes traffic by capacity (weight × ready_replicas), not fixed percentages, so a deployment's share tracks its replica count automatically.We tested it on a live endpoint with tagge