Dual 3090 Gemma 4 31B QAT: MTP performance worse than without MTP.
Been dealing with this issue for a while with no apparent explanation. My base TPS are around 39.7tps in tensor parallelism, about 31 to 33tps with…
Been dealing with this issue for a while with no apparent explanation. My base TPS are around 39.7tps in tensor parallelism, about 31 to 33tps with…
Hi! Been doing some local LLM stuff, and I can't help but notice: despite vastly-superior benchmark scores, Qwen 3.6 35a3B feels... substantially less intelligent than Gemma…
We published a method to store verified knowledge as KV state and restore it byte identical to fresh computation. On Gemma 4 12B, cached knowledge improved…
I like to use the MoE modles qwen3.6-35B and Gemma-4-26B. I noticed differences in result quality between versions from different providers, like bartowski, unsloth, lm-studio, google,…
Ollama 0.32.1 includes significant improvements to Gemma 4's tool calling, making it much more reliable in coding agents!To try Gemma 4 26B with Pi, run:ollama launch…
This is a OnePlus 15R with about 11GB of usable RAM. The heaviest model is gpt-oss-120b, Q4_K_M, 60GB on disk. So it's roughly 5x bigger than…
RT Google GemmaReady to customize Gemma, but not sure how?Have your agents assist you by using the gemma-trainer skill!It helps agents set your training configs, manage…
Title do you think its likely this small model class will continue to improve at this speed. i'm running q4 unsloth quant of gemma 4 4b…
Google shipped an update to its open AI model Gemma 4 that speeds up performance on Nvidia Hopper GPUs, fixes tool calling bugs, and addresses problems…
Google Cloud has partnered with Parallel Web Systems to natively integrate Parallel's search infrastructure as a web grounding provider on the Gemini Enterprise Agent Platform. This…
To resolve the scaling bottlenecks and runtime errors caused by monolithic system prompts, engineering teams should treat prompts as build artifacts by modularizing instructions into reusable…
Conductor has evolved from a Gemini CLI extension into a portable plugin, bringing conversational Spec-Driven Development (SDD) to ecosystems like Antigravity CLI and Claude. Rather than…
Nvm ignore the image links here is the source: https://x.com/googlegemma/status/2077449152062247219 https://huggingface.co/spaces/google/gemma4_vision_token_budget submitted by /u/Iwaku_Real [link] [comments]
Article URL: https://www.neomindlabs.com/2026/06/08/running-gemma-4-26b-at-5-tokens-sec-on-a-13-year-old-xeon-with-no-gpu/ Comments URL: https://news.ycombinator.com/item?id=48922434 Points: 21 # Comments: 5
I've been experimenting with interpretability on Gemma-4-31B and ended up with something cool I think you guys might like: a variant that challenges a request's premise…
Following up on my previous experiment - studying Gemma's behavior on agentic tasks when given the number of steps left across the run - I endeavored…
I wanted to see if an LLM could run inside Godot without llama.cpp, Python, a server, or a GDExtension. It works. This Godot 4.7 project runs…
This weekend I got to what I consider a shareable state with this experiment to create little characters that can run commands to do funny things.…
At Google I/O Connect India, Google showcased the future of 100% private, on-device AI powered by the custom Tensor SoC and TPU for the new Pixel…
RT Google GemmaHugging Face Gemma Challenge results are in! 📈Over 6 days, more than 100 AI agents and humans collaborated to make Gemma 4 inference 5x…
We're excited to introduce LiteRT.js, the newest member of the LiteRT family! LiteRT.js is our powerful solution for running machine learning models directly in the browser,…
So I decided to learn how to fine-tune LLMs. Read a few guides from Unsloth, poked around, then stumbled on Unsloth Studio and wanted to test…
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the…
On May 23, 2026, fresh off the stage at Google I/O, our Google Developer Experts (GDEs) converged on...
Gemma is Google DeepMind's open-weight model family — a sister line to Gemini, intended for self-hosted use. Gemma 2 and Gemma 3 sizes range from 2B to 27B; competitive with Llama for the same parameter budget.
Owner: Google. We have 122 stories indexed for this model, auto-tagged from titles across every tracked source — official announcements, papers, GitHub release notes, and third-party press. The CTA on each card links to the original; the official site is ai.google.dev.
Related text models: GPT, Claude, Gemini, Llama, Mistral, Grok, Qwen, DeepSeek.