I compared GGUF quants of Qwen3.6 27B to NVFP4, AWQ, AutoRound, and FP8
There's an interactive chart and some extra data in the blog post if you're interested. There are plenty of KL-divergence benchmarks for GGUF models, but most…
There's an interactive chart and some extra data in the blog post if you're interested. There are plenty of KL-divergence benchmarks for GGUF models, but most…
Article URL: https://www.vectorware.com/blog/simd-on-gpu/ Comments URL: https://news.ycombinator.com/item?id=49247477 Points: 15 # Comments: 6
Article URL: https://alphatheta.com/en/information/important-notice-security-vulnerability-in-pro-dj-link/ Comments URL: https://news.ycombinator.com/item?id=49247461 Points: 3 # Comments: 0
That's insane yo! I haven't had a chance yet to test it on my 4090 at home but it sounds so promising. And read here that…
Article URL: https://www.anthropic.com/research/riemann-zeta Comments URL: https://news.ycombinator.com/item?id=49247070 Points: 11 # Comments: 1
Article URL: https://galratner.substack.com/p/zillows-ceo-just-fired-500-people Comments URL: https://news.ycombinator.com/item?id=49246910 Points: 8 # Comments: 2
Hey HN,Henry from Cactus here!We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes,…
(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views,…
Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables,…
Looks like the Ling team open weighted a much smaller version of the Ling-3.0-flash they open weighted a few days ago. It's 8B params with 1.3B…
arXiv : https://arxiv.org/abs/2608.00146 Full Paper : https://arxiv.org/pdf/2608.00146 Tweet : https://xcancel.com/googlegemma/status/2086849199052845451#m FYI both (llama.cpp) PRs ( 24423 & 24427 ) went to Draft mode. I'm still waiting…
I've created mobile-harness which is an open-source agent harness giving agents one unified API and a consistent control path across local iOS and Android devices. You…
This work was done by Purvi Chaurasia with Daniel Tan and Chloe Li as part of the SPAR Program for Spring 2026.All code related to the…
Hi HN, we’re Eren, Berat and Kaan. We’re building Stoa (https://www.stoaexchange.com), a marketplace for new and used GPUs and AI servers.GPUs are the collateral in the…
Related: Which programming languages are most token-efficient? - https://news.ycombinator.com/item?id=46582728 - Jan 2026 (91 comments) Comments URL: https://news.ycombinator.com/item?id=49245936 Points: 23 # Comments: 14
Having a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I…
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the…
This article is a summary of an original study by Compassion in Machine Learning (CaML): Brazilek, J., Chaudhary, M., Lu, Z., & Tidmarsh, M. (2026). Coercion…
Article URL: https://thehistoricalinsights.page/2026/06/why-addresses-have-numbers.html Comments URL: https://news.ycombinator.com/item?id=49245646 Points: 11 # Comments: 5
Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement from…
Article URL: https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii Comments URL: https://news.ycombinator.com/item?id=49245491 Points: 24 # Comments: 5
My daily driver is Qwen3-235b-a22b-instruct-2507-Q4_K_M.gguf and it has been for a long time. I get around 75 t/s prompt processing and starting lower context ~5.5 t/s…
Article URL: https://finance.yahoo.com/healthcare/articles/harvard-study-links-glp-1-123000637.html Comments URL: https://news.ycombinator.com/item?id=49245487 Points: 65 # Comments: 44
It looks like Demis Hassabis is stepping away from Google DeepMind. In honor of his rise and presumed fall, I wrote an essay on the powers…