Agents / Google DeepMind
Gemma 4 · 12B
A compact multimodal model that understands text, images, audio, and video for local assistants.
Not yet testedRepository opened May 23, 2026
LOCALCLOUDOPEN WEIGHTS
This is a source-based starting point. We haven’t independently tested this model yet. Source-based estimates are separate from our test results. Measured speed and our verdict will appear after a reproducible test.
At a glance
- Parameters
- 12B
- Architecture
- Dense transformer
- Context length
- 256,000
- License
- Apache 2.0
- Recommended VRAM
- Not yet tested
- Minimum tested VRAM
- Not yet tested
- Disk space
- Not yet tested
- Software
- Transformers
HARDWARE ESTIMATE
What will it need?
Our test notes
Not yet tested. We’ll publish the hardware, quantization, context, speed, peak memory, and load time together. A parameter count alone is not a hardware requirement.
Our verdict
Not yet tested. Check the publisher’s documentation for capabilities and limitations; we’ll add an independent verdict after testing.
Explore its uses
AGENTSMULTIMODALCODINGRAG
Keep exploring
Agents
Z.aiGLM-5.3-Flash
A multimodal model for coding and agent tasks, with 18B active parameters from a 320B total model.
MODEL SIZE320B total · 18B active
Agents
Moonshot AIKimi K3
An open-weight multimodal model for long-running coding, reasoning, and knowledge work.
MODEL SIZE2.8T total · 104B active
Agents
Google DeepMindGemma 4 · E4B
A smaller on-device Gemma model for reasoning, coding, and text, image, and audio understanding.
MODEL SIZE8B total