A community project to build open source models
This is probably a dumb question, but I am going to ask it anyway. In light of open source models getting so popular, the chance any…
This is probably a dumb question, but I am going to ask it anyway. In light of open source models getting so popular, the chance any…
I previously built a telegram bot that chatted with people. But it was never put in production. Besides that I often find that even the cheap…
Introduction I watched 2.49 GB of state restore from disk in 1.23 seconds — and then get thrown away. llama-server's slot save/restore promises exactly what long-context…
Been lurking here for a while and using local models for client work. Kept hitting the same wall, the people who would benefit most from local…
Follow-up to my earlier posts: Should I sell my Mac Studio? https://www.reddit.com/r/MacStudio/s/GK7QP8Lg87 Kimi benchmark: https://www.reddit.com/r/LocalLLaMA/s/ujBsYLYmpd Short version: my Mac Studio was sitting mostly idle, and from…
I wanted to test the official results vs fp8 + fp8 kv. basic sglang setup on H200. If anyone wants one of the official tests do…
Hi, r/LocalLLaMA! SupraLabs just released a model called: SupraLabs/reasoning-summarizer-800m-pre-gguf It is a thought trace summarizer. Basically, the user/dev can send a reasoning + tool calls (if…
Working on a new benchmark suite that tests how well agents can design physical objects in Box3D. tasks include designing cars, trucks, cranes, etc. submitted by…
A friend has been kind to let me test his new device. 125 ddr5, AMD Ryzen 395. What models you guys recommend for coding that I…
This began as me attempting to run the ONNX version of LivePortrait (https://github.com/KlingAIResearch/LivePortrait) in Chrome with WebGPU. It took 30 seconds to generate a single frame.…