Turbo-fieldfare: Open-source engine running Gemma 4 26B in 2 GB RAM on Apple Silicon
Its a custom Swift/Metal inference engine that runs Gemma 4 26B-A4B-IT on M-series Macs with very low RAM. It uses ~2GB instead of ~14 GB. The…
Its a custom Swift/Metal inference engine that runs Gemma 4 26B-A4B-IT on M-series Macs with very low RAM. It uses ~2GB instead of ~14 GB. The…
Google's open-source TPU microbenchmark suite provides developers with granular performance metrics across Network, Compute, HBM, Host Transfer, and Attention components to validate real-world hardware capabilities. By…
This research was done as my capstone project during ARBOx4.Epistemic Status: I'm relatively sure the results I obtained and my interpretations are correct. I'm unsure if…
Hi HN,I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called…
What are your experiences with Gemma 4's QAT versions compared to their regular ones? So far I have mostly heard about regressions, but if you have…
I really love this model, I have been using the q4_k_l by Bartowski (I have heard QAT is quite the downgrade in some aspects) and it…
I’ve been doing some benchmarking on my mini-PC setup (AMD Ryzen 7 6800H) to see how it handles the new Gemma 4 and Qwen 3.6 MoE…
https://x.com/i/status/2081398564345802934 submitted by /u/jacek2023 [link] [comments]
Now that Google signed the letter, would be a good time for them to release the 100B Gemma 4 model, show they really mean it :)
RT Christopher Nguyen ⽗The Gemma series is an amazing set of highly performant open-weight models. They have proven extremely effective in industrial settings where site-deployed agents…
Hi everyone, Before I begin, I should mention that the system I'm showcasing was developed by the team at Noema, which I founded. I wanted to…
arXiv:2607.20522v1 Announce Type: new Abstract: This paper tests whether holonomy concentrates on active sparse-autoencoder (SAE) feature planes in Gemma 2 2B, a concrete operationalization of the…
This second installment explores how Ray’s higher-level libraries—Serve, Data, and Train—abstract the complexities of running AI workloads on Google's TPU slices. Ray Serve uses a simple…
Gemma 4 was updated (mostly chat templates) and I took it for a test. On a local llama.cpp server running on M5 Pro with 48GB, 26B…
RT Google GemmaGemma 4 just crossed 300 million downloads.Thank you to the developers, researchers, and open-source community building with us. Your work and feedback drive this…
The viability of orbital data centers hosting the largest and most capable large language models (LLMs) remains hotly contested. But enormous deployments that require thousands of…
Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty…
Hey HN, Henry & Roman here from Cactus.A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast.…
Tunix is Google’s new JAX-native post-training library designed to eliminate TPU idling bottlenecks when training multi-turn, tool-using LLM reasoning agents. It maximizes hardware throughput by combining…
RT Google GemmaVoice AI without the wait! ⏱️Thanks to Hugging Face and Cerebras, developers can now use the Gemma 4 31B model as the brain for…
Please tell me I'm doing something wrong. My config: [*] flash-attn = on jinja = true fit = true offline = true mmproj-offload = false mmap…
A single 24GB GPU is the practical floor for serious local inference. This guide compares six open-weight models that fit one card at Q4_K_M. It covers…
Ray 2.55 introduces official, first-class support for Google Cloud TPUs, enabling developers to run distributed Python workloads on Google's accelerators using the familiar Ray task-and-actor APIs.…
Gemma 4 26b is really good model and with recent update it's even better than before but why google did not add audio capability to this…
Gemma is Google DeepMind's open-weight model family — a sister line to Gemini, intended for self-hosted use. Gemma 2 and Gemma 3 sizes range from 2B to 27B; competitive with Llama for the same parameter budget.
Owner: Google. We have 122 stories indexed for this model, auto-tagged from titles across every tracked source — official announcements, papers, GitHub release notes, and third-party press. The CTA on each card links to the original; the official site is ai.google.dev.
Related text models: GPT, Claude, Gemini, Llama, Mistral, Grok, Qwen, DeepSeek.