Agents / Z.ai
GLM-5.3-Flash
A multimodal model for coding and agent tasks, with 18B active parameters from a 320B total model.
Not yet testedRepository opened Aug 25, 2026
LOCALCLOUDOPEN WEIGHTS
This is a source-based starting point. We haven’t independently tested this model yet. Source-based estimates are separate from our test results. Measured speed and our verdict will appear after a reproducible test.
At a glance
- Parameters
- 320B total · 18B active
- Architecture
- Mixture of experts
- Context length
- See publisher documentation
- License
- MIT
- Recommended VRAM
- Not yet tested
- Minimum tested VRAM
- Not yet tested
- Disk space
- Not yet tested
- Software
- vLLM, SGLang, Transformers
HARDWARE ESTIMATE
What will it need?
Our test notes
Not yet tested. We’ll publish the hardware, quantization, context, speed, peak memory, and load time together. A parameter count alone is not a hardware requirement.
Our verdict
Not yet tested. Check the publisher’s documentation for capabilities and limitations; we’ll add an independent verdict after testing.
Explore its uses
AGENTSCODINGMULTIMODALREASONING
Keep exploring
Agents
Moonshot AIKimi K3
An open-weight multimodal model for long-running coding, reasoning, and knowledge work.
MODEL SIZE2.8T total · 104B active
Agents
Google DeepMindGemma 4 · 12B
A compact multimodal model that understands text, images, audio, and video for local assistants.
MODEL SIZE12B
Agents
Google DeepMindGemma 4 · E4B
A smaller on-device Gemma model for reasoning, coding, and text, image, and audio understanding.
MODEL SIZE8B total