Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots.
Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables,…
Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables,…
Looks like the Ling team open weighted a much smaller version of the Ling-3.0-flash they open weighted a few days ago. It's 8B params with 1.3B…
arXiv : https://arxiv.org/abs/2608.00146 Full Paper : https://arxiv.org/pdf/2608.00146 Tweet : https://xcancel.com/googlegemma/status/2086849199052845451#m FYI both (llama.cpp) PRs ( 24423 & 24427 ) went to Draft mode. I'm still waiting…
I've created mobile-harness which is an open-source agent harness giving agents one unified API and a consistent control path across local iOS and Android devices. You…
Having a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I…
My daily driver is Qwen3-235b-a22b-instruct-2507-Q4_K_M.gguf and it has been for a long time. I get around 75 t/s prompt processing and starting lower context ~5.5 t/s…
My daily driver is Qwen3-235b-a22b-instruct-2507-Q4_K_M.gguf and it has been for a long time. I get around 75 t/s prompt processing and starting lower context ~5.5 t/s…
Hello~ We just shipped Ante 0.2, and the part I think this community will care about most is offline mode. We wanted local to be a…
Google had hosted this Hackathon months ago- just checked that they are ready with the results and will release the results soon. Then saw that there…
submitted by /u/EmPips [link] [comments]