Well.. it’s a step up from nonstop bot spam I guess
submitted by /u/ForsookComparison [link] [comments]
submitted by /u/ForsookComparison [link] [comments]
The quadratic complexity of attention poses a critical bottleneck for long-context processing, spurring interest in hybrid attention designs. Most open-source hybrid models adopt a layer-wise strategy.…
First of all, a huge thank you to the r/LocalLLaMA community and the 3090 club. This benchmark started from your shared recipes... These are my findings…
GLM 5.2 has been getting a lot of hype in the past two weeks and for the good reasons. MIT weights, 1M context, ~$1/$4.20 per M…
Meta reportedly had hundreds of contractors pose as minors and send suicide, sex, and drug-related prompts to chatbots from OpenAI, Google, and Character.AI. In a single…
submitted by /u/paf1138 [link] [comments]
https://x.com/Chinazhidx/status/2071877413685109071 TODAY: #Huawei open-sources OpenPangu-2.0-Flash #OpenPangu 2.0 includes two 512K-context models: • Flash: 92B total,6B active—Weights+inference code+training ops released • Pro: 505B total,18B active—flagship model, coming…
Looking forward to compare with Antirez's DS4 imamtrix https://huggingface.co/bartowski/DeepSeek-V4-Flash-GGUF submitted by /u/challis88ocarina [link] [comments]
They are just MTP-only GGUF subsets of Qwen3.5/3.6 Middle/Large (27B and above) models (to accelerate token generation of Qwen-based models without MTP tensors). But I hope…
https://huggingface.co/nvidia/Qwen3.6-27B-NVFP4 submitted by /u/vanbukin [link] [comments]