I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)
I kept answering the same question for friends ("I've got a 16GB MacBook / a 3060, what can I actually run?") and got tired of guessing,…
I kept answering the same question for friends ("I've got a 16GB MacBook / a 3060, what can I actually run?") and got tired of guessing,…
It is important to document requirements, capture key decisions, and record design goals. Best practices: Requirements doc stored in repo. Store plan files created by LLM…
I've seen projects which stream tool use status and subagent generation, and represented it with a nice little visual based on the tool being used, etc.…
submitted by /u/tarruda [link] [comments]
My attempt at running Qwen3.5 122B on my 5090 (32GB VRAM) + 64GB RAM is really bleak. I'm getting a speed that starts at 6 tps…
Spyware-like code in Claude Code that covertly targets Chinese users. submitted by /u/zakadit [link] [comments]
I have been working on Hister, a self hosted search engine that automatically indexes pages you visit, local files, and documentation, then keeps them searchable with…
Some in this sub have tested GLM5.2 on 4x DGX Sparks (or Ascend GX10) with 400-500 tok/s prompt processing and ~15 tok/s output at 128k context.…
1. Introduction openPangu-2.0-Flash is an MoE model trained on Ascend. The model has 92B total parameters and 6B activated parameters. Its context length is 512k. The…
Been lurking here a while, this sub is basically why LokalBot exists. It's a Mac app that records + summarizes your meetings, autocompletes your typing in…