The lab named after humanity is the only one that will not share weights with humanity
submitted by /u/InternationalGap3698 [link] [comments]
submitted by /u/InternationalGap3698 [link] [comments]
I finished the speed leg of my spec-decode benchmarking for Qwen3.6-27B, main algorithms across quants. Overall: the heavier the quant, the more spec-decode buys you (10…
Forked SGLang, wrote TeilLang FlashAttention for V100, used open-source marlin-v100, ungated flashinfer for sm70, made Dflash work for Qwen3.5/3.6 models, added Laguna S2.1 support, tried to…
submitted by /u/SignificantLegs [link] [comments]
Chinese chipmaker CXMT surged by almost 500% on its first day of trading, bringing its total market capitalization to approximately RMB 3.28 trillion and making…
Hey r/LocalLLaMA, I’m a retired engineer with a background in distributed computing, currently running a 1-person startup. Like many people here, I’d love to experiment with…
I now have llama.cpp running pretty well for my needs, but the inability to quickly set/swap models and system prompts isn't ideal. Lm-studio let's you save…
I've been seeing a lot of news about the latest gemma 4 and qwen 3.6 being really good and the current go-to models but those are…
Test Prompts: 1.1. Algorithm & Logic (10 pts): "Write a function in Python that finds the contiguous subarray with the largest sum (Kadane's algorithm). Include time…
https://preview.redd.it/k97l56d8ypfh1.png?width=606&format=png&auto=webp&s=2a2e2156bea56b25f5709e8f1df2bf82525eb089 https://x.com/alexandr_wang/status/2081501627836661928?s=20 submitted by /u/External_Mood4719 [link] [comments]