20GB VRAM + 64GB DDR5 – Qwen3.6 35B A3B still the best choice?
Hey everyone. Trying to figure out the best setup for my hardware. I've got a 64GB DDR5 RAM laptop and A 20GB 7900xt egpu. Usecase is…
Hey everyone. Trying to figure out the best setup for my hardware. I've got a 64GB DDR5 RAM laptop and A 20GB 7900xt egpu. Usecase is…
Maybe the most pathetic genre of China cope, this "Soli Invicto" Hajnaltard doesn't even understand what he says. Chyna stole from the West, China exploited its…
oh yeah baby!If they just slash prices by fiat… they're probably still in the greenbut that'll motivate them to optimize the architecture for cheaper cacheand Grok…
This post introduces a tool: an Epistemic Audit for Existential Risks from AI. It is a structured way to map, organize and track your beliefs across…
SDK regeneration (#810) * [fern-generated] Update SDK Generated by Fern CLI Version: unknown Generators: - fernapi/fern-python-sdk: 4.42.0 * [fern-replay] Applied customizations Patches with unresolved conflicts (1):…
> This is not about whether children can access social media.> It is about when social media can access our children.Genius wypipo word magick! Uno reverso!…
We are pleased to share our latest research, now published in Nature Communications: “Smart Cellular Bricks: Physical Modules That Recognize Their Own Shape and Repair Themselves.”Blog:…
RT Sakana AIWe are pleased to share our latest research, now published in Nature Communications: “Smart Cellular Bricks: Physical Modules That Recognize Their Own Shape and…
Google Research's SensorFM is a foundation model trained on more than a trillion minutes of wearable data from five million Fitbit and Pixel Watch users. It…
chat : fix reasoning leak with force-opened bare templates (#24674) chat : fix reasoning leak with force-opened bare templates The reasoning start tag inferred from prior…
Realistic and diverse traffic simulation is essential to autonomous driving development. Yet prevailing benchmarks predominantly reward realism, and recent methods have optimized accordingly, leaving diversity underexplored.…
I wanted to see if an LLM could run inside Godot without llama.cpp, Python, a server, or a GDExtension. It works. This Godot 4.7 project runs…
Waze is getting an AI makeover. Google is integrating its flagship AI assistant, Gemini, into the driving app with the goal of letting users personalize their…
Python has become the default language for data-intensive work, AI applications and...
Agentic LLMs keep failing the same way because they lack specific, reusable capabilities. Stanford's TRACE diagnoses those gaps from an agent's own trajectories, synthesizes one verifiable…
Article URL: https://raymyers.org/post/zed-creator-calls-spade-a-spade/ Comments URL: https://news.ycombinator.com/item?id=48889637 Points: 19 # Comments: 3
sycl: add fused top-k MoE (#25217) sycl: add fused top-k MoE sycl: address review: GGML_SYCL_ENABLE_FUSION env, move fusion dispatch to topk-moe sycl: print GGML_SYCL_ENABLE_FUSION at startup…
Anthropic’s Jacobian Lens work introduced a way to inspect verbalizable representations inside language models. Follow-up experiments suggested that entropy in this internal “workspace” might help identify…
Article URL: https://www.disruptionbanking.com/2026/07/13/inside-berkshires-397-billion-bet-against-an-overheated-market/ Comments URL: https://news.ycombinator.com/item?id=48889429 Points: 16 # Comments: 0
The difference between Elon Musk and his Chinese competitors is that they do 1 to 100 and he does *0.5 to 100*. He just repeatedly finds…
Article URL: https://shkspr.mobi/blog/2026/07/another-ridiculous-interrail-holiday-6379km-and-13-countries-over-7-weeks/ Comments URL: https://news.ycombinator.com/item?id=48889350 Points: 12 # Comments: 0
The startup, PrismML, said it has shrunk down Qwen 3.6, an open-source large language model developed by Chinese internet giant Alibaba, to run on an iPhone…
Anthropic is keeping Claude Fable 5 in its subscription plans through July 19, 2026. The model was supposed to switch to pay-per-use today. Subscribers can use…
When JetBrains introduced Mellum 2, they advertised its latency being as low as Qwen2.5-Coder 7B. This was achieved via MTP. In the GGUFs they've published, I…