Introducing Shieldstral. | Mistral AI
submitted by /u/tengo_harambe [link] [comments]
submitted by /u/tengo_harambe [link] [comments]
Why nobody is talking about this? Seems pretty significant to the community submitted by /u/MuzafferMahi [link] [comments]
I‘m doing a research proposal at my company about running local LLMs to replace daily coding models. Qwen 3.6 27B (or 3.8 potentially) is widely seen…
Just wanted to share my agentic coding benchmark run of DSv4F 0731 at both High and Low reasoning efforts (not Max)... I ran a 109-question subset…
A new llama.cpp PR (#26563) adds a heatmap that tracks which MoE experts are used most often. Instead of keeping every expert on the GPU or…
https://x.com/xiong_hui_chen/status/2084695353346117760?s=20 submitted by /u/Capital-Remove-6150 [link] [comments]
Released today, with emphasis on agentic capabilities. I really like their models for simple, high volume tasks ("summarize these gazillion documents") and their 8b-a1b was my…
Link: https://cursor.com/blog/mixture-of-kittens GitHub: https://github.com/cursor/mixture-of-kittens Seems like a neat way to squeeze more performance out of MoE. I'm sure everyone has a favorite MoE model they'd like…
https://preview.redd.it/522fsdwvtdhh1.png?width=1200&format=png&auto=webp&s=6a6cf7a467514167a8193029dbd20fb3a9ba4f6c It ranks lower than both Sonnet 4.6 and Luna. I'd wager Luna costs in the same ballpark as DS4F considering Luna’s token efficiency. DeepSeek being…
Weights went up today so this is downloadable now, MIT, ~128GB for the official FP8. I ran these on the API before that landed, so treat…