A multimodal model for coding and agent tasks, with 18B active parameters from a 320B total model.
An open-weight multimodal model for long-running coding, reasoning, and knowledge work.
A compact multimodal model that understands text, images, audio, and video for local assistants.
The smallest Gemma 4 variant, built for local assistants with text, image, and audio input.
A smaller on-device Gemma model for reasoning, coding, and text, image, and audio understanding.
A compact vision-language model for reasoning, coding, agents, and image understanding.
A compact multimodal model for text and image understanding; a building block for private assistants.