12GB VRAM gang, what’s our plan?
Seems like we're limited to qwen finetuned MoEs for now. Looking at the current landscape - focus seems to be on dense models (muse glimmer 30b,…
Seems like we're limited to qwen finetuned MoEs for now. Looking at the current landscape - focus seems to be on dense models (muse glimmer 30b,…
Available b10356 onwards. Overview ROCm 7.14 is the first production release using TheRock build system. It can be installed using multi-arch deliverables from wheels, debs, rpms,…
Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse…
Once upon a time, John Wentworth and I thought we had a proof of a very useful looking theorem. We did not have that proof. An…
Article URL: https://correctiv.org/en/europe/2026/04/21/half-of-europes-towns-and-villages-have-fewer-residents-than-60-years-ago/ Comments URL: https://news.ycombinator.com/item?id=49253813 Points: 26 # Comments: 22
Article URL: https://www.cnbc.com/2026/08/10/openai-wraps-7-billion-share-sale-ahead-of-potential-ipo-.html Comments URL: https://news.ycombinator.com/item?id=49253785 Points: 8 # Comments: 2
Article URL: https://www.nytimes.com/2026/08/11/world/europe/france-ban-unsolicited-telemarketing-calls.html Comments URL: https://news.ycombinator.com/item?id=49253772 Points: 13 # Comments: 1
Article URL: https://manish.sh/writings/models/inside-deepseek-reverse-engineering-an-ai-assistant-by-interviewing-itself Comments URL: https://news.ycombinator.com/item?id=49253738 Points: 4 # Comments: 0
Article URL: https://github.com/activeing123/mcptoon Comments URL: https://news.ycombinator.com/item?id=49253721 Points: 15 # Comments: 2
In this post, we find that when teacher models are prompted to imitate one another, students learn the imitated model's detectable writing signature but their direct…
Confirmed by the official Qwen account. submitted by /u/Bestlife73 [link] [comments]
submitted by /u/fallingdowndizzyvr [link] [comments]
TL;DR: We’ll probably be getting more AI safety incidents, so amplify the ones that would justify or highlight the urgency of your preferred policy solutions even…
Article URL: https://www.neowin.net/news/windows-11-admins-unhappy-as-microsoft-found-installing-unexpected-new-onedrive-photos-app/ Comments URL: https://news.ycombinator.com/item?id=49253329 Points: 26 # Comments: 7
Article URL: https://blog.mozilla.org/security/2026/08/10/updated-gpg-key-for-signing-firefox-and-thunderbird-releases/ Comments URL: https://news.ycombinator.com/item?id=49253250 Points: 13 # Comments: 3
Article URL: https://whatever.scalzi.com/2026/08/10/why-my-father-is-wrong-a-defense-of-guitar-hero/ Comments URL: https://news.ycombinator.com/item?id=49253176 Points: 16 # Comments: 2
I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes.…
A few things right off the bat: it reasons very efficiently. Like Grok 4.5 levels of efficient thinking it quantizes very well. My first few tests…
TL;DRMachine unlearning is a proposed technique for removing harmful knowledge from AI models. However, recent work has shown that most current unlearning methods are not robust…
“Diversification internationally," Canadian Prime Minister Mark Carney remarks, “is not just economic prudence — it is the material foundation for honest foreign policy.” Carney’s address to…
A few days ago, I came across a Reddit thread about anomalous responses produced by Anthropic’s newly released model, Claude Opus 5. The trick, apparently, was…
Recently, I was improving a small LLM-powered classifier and noticed a few continuously failing test cases. As many would, I asked my AI code assistant to…
In 2015 I made a little webapp that would use the NextBus API to show predictions for the MBTA: A few years later they moved from…
This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all…
Community-driven discussion of AI. r/LocalLLaMA (open-weight model talk, quantization, llama.cpp tuning), Hacker News AI-tagged front-page items, the Alignment Forum and LessWrong AI section. The Reddit and HN streams are noisy but capture early signal that the formal press picks up days later.
6541 stories indexed in this category. See also all models and the full source catalogue.