Skip to content
X · @huggingface · X / Twitter

RT NVIDIA AI: We took a 30B model and split it in two to write tokens in parallel instead of one at a time. Introducing Nemotron-Labs-TwoTower: a diff…

RT NVIDIA AIWe took a 30B model and split it in two to write tokens in parallel instead of one at a time.Introducing Nemotron-Labs-TwoTower: a diffusion language model from NVIDIA Research adapted from Nemotron-3-Nano-30B-A3B. Here’s how it works: one half holds the context, the other writes the tokens, with both reusi