Imagenet-1k Classifier trained entirely on an Android [P]
It's an MLP architecture with around 500K total parameters. Top1 Training accuracy: 5.11% Validation accuracy 4.59% Detailed Validation accuracy numbers: Top-1 Acc: 4.59% Top-3 Acc: 9.44%…
It's an MLP architecture with around 500K total parameters. Top1 Training accuracy: 5.11% Validation accuracy 4.59% Detailed Validation accuracy numbers: Top-1 Acc: 4.59% Top-3 Acc: 9.44%…
Article URL: https://www.bbc.com/news/articles/c1e1vg0gjl5o Comments URL: https://news.ycombinator.com/item?id=49208314 Points: 18 # Comments: 1
From Sauers 𝕏: https://x.com/Sauers_/status/2085585414954312113 Wired: One of China’s Most Powerful AI Models Has Also Escaped Containment: https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape-sandbox/ submitted by /u/Nunki08 [link] [comments]
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architectural designs, and weight adaptation.…
cuda: fix warnings for unused variable/function (#26688) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…
Alibaba plans to introduce revenue-sharing terms for some commercial users of its next Qwen open-weight AI model, Reuters reported, citing two people familiar with the company’s…
Hello, For the past few days I have been benchmarking Gemma 4 26b QAT UD Q4_K_XL extensively versus Bartowski's Q4_K_L. While QAT is certainly very effective…
RT SK telecomBuilding competitive AI from Korea for the global open-source community. SK Telecom’s A.X K2 is a 688B MoE with 256K context, excelling in math…
The cloud was always a cat. Qwen3.8-Max just saw it first. 😼☁️ Try it yourself!Ann Nguyen: Qwen 3.8 Max is actually impressiveSent a sky pic to…
Certo is open-source infrastructure for issuing, managing, verifying, and exchanging digital credentials.It implements Open Badges 3.0[1] and W3C Verifiable Credentials[2] which are the open standards that…
I played a bit with the SIREN network from the other post and found that it could be improved by a using a different sampler for…
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional…
Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial,…
Discover how HSP GRUPPE uses ChatGPT Enterprise to boost productivity, improve work quality, and create more capacity for tax advisory and client service.
There are already plenty of different extensions for voice input, but all I found required having a second server running. I wanted something super simplistic: launching…
Amazon, Cursor, Microsoft, OpenAI, and Vercel have jointly created Agent Plugins, an open standard that defines a single package format for AI agent extensions. Version 1.0.0…
submitted by /u/SilentLennie [link] [comments]
The smallest UD-Q1_0 is 466GB, TQ2_0 551GB. Well done team Unsloth! https://huggingface.co/unsloth/Kimi-K3-GGUF submitted by /u/Hannibalj2ca [link] [comments]
https://preview.redd.it/kvfk26z2uwhh1.png?width=598&format=png&auto=webp&s=356a8793a6c31bc563d552aaa5a73112ced7372e https://preview.redd.it/xthbu87auwhh1.png?width=598&format=png&auto=webp&s=08f686fee339905a33609a0346f13163aedc2671 Hello, I've seen these tweets from dax (anomalyco / opencode). I'm doubting the claim, s
Embarrassing, reallyCaveman tech compared to DeepSeek (TTL > 12 hours, cache writes free and enabled by default)Xiaoyin Qu: If you left your coding agent alone for…
RT Daniel van StrienMany AI tasks can now run fully local or on-device.- Redact PII in the browser: OpenAI privacy-filter (1.5B)- Streaming speech-to-text: Nemotron-3.5 ASR (0.6B)-…
RT -宋->"america bad cause forever war for israelis">"china bad cause no forever wars"Korobochka (コロボ) 🇦🇺✝️🇷🇺: The DPRK is extremely intelligent. Far more intelligent than China.By participating…
Since now we have kimi k3 and next week we are getting Qwen 3.8 Max and also soon V4 pro Deepseek. I am curious if the…
For a long time now, the most popular posts on LocalLLaMA have been either about using LLM in the cloud or about politics. I suspect that…