Bessent says U.S. could sanction China over AI model ‘theft’
submitted by /u/fallingdowndizzyvr [link] [comments]
submitted by /u/fallingdowndizzyvr [link] [comments]
Been running DeepSeek-V4-Flash for an offline batch job (cleaning a big pile of short text records, so lots of small prompts rather than chat). Single B300,…
Kind of unexpected. Happy for Gemma-4/Google, big win for us, LocalLLMers. Yet Qwen3.6 still does better in Hermes than Gemma-4 somehow. We need Gemma-4.1 fine-tuned on…
Model Size Terminal-Bench 2.1 SWE-bench Multilingual SWE-Bench Pro (Public Dataset) DeepSWE SWE Atlas (Codebase QnA) Toolathlon Verified Laguna S 2.1 118B-A8B 70.2% 78.5% 59.4% 40.4% 46.2%…
HF: https://huggingface.co/poolside/Laguna-S-2.1 GGUFs available for use with llama.cpp custom fork: https://huggingface.co/poolside/Laguna-S-2.1-GGUF Posted on X: https://x.com/poolsideai/status/2079613777343848465?s=20 submitted by /u/Lowkey_LokiSN [link] [comments]
https://huggingface.co/Nanbeige/Nanbeige4.2-3B Nanbeige4.2-3B is a compact agentic model built on Nanbeige4.2-3B-Base, designed to combine strong agentic behavior with broad reasoning and alignment capabilities. Its Looped Transformer architecture…
Four things it does: Local text editing: mark a region (or just say it in natural language) and replace specific text. Fixes typos, swaps numbers, changes…
No weights yet. I feel sad for them, that training run cost maybe 100s of thousands of dollars and they didn't even beat GPT-OSS 20B in…
Most of us already know the main important mainline local LLMs to save like Qwen3.6 27b, Gemma4 31b, GLM5.2, etc. And then for diffusion models, Z-Image…
pi 0.81.0 now has integrated support for llama.cpp (llama-server router). https://pi.dev/docs/latest/llama-cpp This seems to be able to replace the huggingface/pi-llama extension and/or manually managing models in…