Agents / Google DeepMind

Gemma 4 · 12B

A compact multimodal model that understands text, images, audio, and video for local assistants.

Not yet testedRepository opened May 23, 2026
LOCALCLOUDOPEN WEIGHTS
This is a source-based starting point. We haven’t independently tested this model yet. Source-based estimates are separate from our test results. Measured speed and our verdict will appear after a reproducible test.

At a glance

Parameters
12B
Architecture
Dense transformer
Context length
256,000
License
Apache 2.0
Recommended VRAM
Not yet tested
Minimum tested VRAM
Not yet tested
Disk space
Not yet tested
Software
Transformers
Read the official model documentation
HARDWARE ESTIMATE

What will it need?

Our test notes

Not yet tested. We’ll publish the hardware, quantization, context, speed, peak memory, and load time together. A parameter count alone is not a hardware requirement.

Our verdict

Not yet tested. Check the publisher’s documentation for capabilities and limitations; we’ll add an independent verdict after testing.

Explore its uses

AGENTSMULTIMODALCODINGRAG

Keep exploring

Gemma 4 · 12B: hardware, VRAM & setup | YouRunAI