Gemma 4: Google’s Most Capable Open‑Source Model
Google has launched Gemma 4, a new family of open‑source, Apache 2.0‑licensed model built from Gemini‑3‑level research, optimized for advanced reasoning, agentic workflows, and on‑device AI across phones, laptops, and servers.
Gemma 4 is Google’s fourth‑generation open model family, described as “byte‑for‑byte the most capable open models” available today. It is built on the same core research and infrastructure as Gemini 3, but released with open weights and a permissive Apache 2.0 license, targeting both open‑source developers and commercial builders. Since the first Gemma release, developers have downloaded the prior models over 400 million times and created more than 100,000 variants, which Google calls the “Gemmaverse.”
Key model variants and sizes
Google ships Gemma 4 in four main flavors:
E2B (Effective 2B) – an ultra‑efficient on‑device model with about 2.3 billion active parameters from around 5.1 billion total, optimized for phones and IoT.
E4B (Effective 4B) – a slightly larger edge model with roughly 4.5 billion active parameters from 8 billion total, still designed for low‑RAM, battery‑sensitive hardware.
26B Mixture‑of‑Experts (MoE) – a 26‑billion‑parameter model that activates only about 3.8 billion parameters per token, prioritizing low latency and fast token generation.
31B Dense – a 31‑billion‑parameter dense model maximizing raw quality and serving as a strong foundation for fine‑tuning, running unquantized on a single 80 GB NVIDIA H100 GPU.
All four families are positioned to handle complex logic and agentic workflows, not just simple chat.
Performance and “intelligence‑per‑parameter”
On the Arena‑AI chat‑evaluation leaderboard as of early April 2026, Gemma‑4’s 31B variant ranks as the #3 open model worldwide, while the 26B MoE sits at #6. Google claims these models can outperform proprietary models up to about 20× their parameter count on certain benchmarks, which is why it emphasizes “intelligence‑per‑parameter” and lower hardware overhead. The 26B MoE trades some absolute quality for very high tokens‑per‑second, making it attractive for latency‑sensitive applications.
Target hardware and deployment options
To run on “your hardware,” Gemma 4 weights are optimized for:
Edge devices such as Android phones, Raspberry Pi boards, and NVIDIA Jetson Orin Nano.
Consumer‑grade GPUs (e.g., 4‑bit‑quantized versions on 24 GB GPUs like RTX 4090 or RX 7900 XTX).
Workstation and data‑center GPUs (bfloat16 weights on an 80 GB H100) and TPUs such as Google’s Trillium and Ironwood.
Developers can serve these models locally via tools like vLLM, llama.cpp, Ollama, MLX, LM Studio, and MaxText, or in the cloud via Google Cloud’s Vertex AI, Cloud Run, GKE, TPU‑serving, and Sovereign‑Cloud options. Google also highlights forward‑compatibility with Gemini Nano 4 on Android, via the AICore developer preview and related Android SDKs.
Key technical capabilities
Advanced reasoning & agents: Gemma 4 supports multi‑step planning, math, and instruction‑following benchmarks, with native function‑calling, structured JSON output, and system instructions for building autonomous agents.
Code generation: The models generate high‑quality offline code, enabling local‑first AI coding assistants without a constant cloud connection.
Multimodality: All variants accept text and images; E2B and E4B add native audio input for speech recognition and understanding, while larger models also handle video and charts.
Context length: Edge models offer 128K‑token contexts, while 26B and 31B variants scale up to 256K‑token windows, enabling repository‑level analysis and long‑form document processing in a single prompt.
Languages: The models are trained natively on more than 140 languages, targeting global, inclusive applications.
Licensing, safety, and ecosystem
Google moved Gemma from older, more restrictive licenses to the Apache 2.0 license, removing many former commercial and usage restrictions. This encourages open‑source fine‑tuning, self‑hosting, and deployment in regulated or sovereign environments, while still applying the same underlying security, safety, and infrastructure protocols as its proprietary Gemini models.
To drive adoption, Google offers:
Quickstart access in Google AI Studio (31B and 26B MoE) and Google AI Edge Gallery (E2B, E4B).
Day‑one integrations with frameworks such as Hugging Face Transformers, vLLM, llama.cpp, Ollama, Baseten, LiteRT‑LM, NVIDIA NIM, ROCm, and Keras.
A “Gemma 4 Good Challenge” on Kaggle to incentivize AI projects that create positive social impact.
Why this is important
Gemma 4 represents a major step toward democratizing frontier‑class AI, giving developers commercial freedom plus high‑performance, multimodal reasoning that previously required large cloud‑only proprietary models. By optimizing for everything from low‑end Android phones to 80 GB H100‑class systems, Google positions Gemma 4 as a single model family that can span edge, laptop, and server stacks, which simplifies ecosystem development and deployment. The Apache 2.0 license, combined with strong on‑device capabilities and agent‑ready features, also makes it a compelling foundation for open‑source alternatives to closed‑API LLM ecosystems, especially in Europe and other privacy‑sensitive regions.