r/LocalLLaMA
· Communities
inclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8
Went public in the last few minutes, both repos ungated. Ling-3.0-flash, BF16, 24 shards, ~255GB Ling-3.0-flash-fp8, official FP8, ~128GB 127.5B total, they quote 5.1B active. What jumped out at me in config.json is 512 experts with 8 active per token, which is a lot finer-grained than most of what gets posted here. Ar