Text / Qwen

Qwen3.8-Flash-Next

An experimental open-weight multimodal model with sparse attention and 262K native context.

Not yet testedRepository opened Aug 24, 2026
LOCALCLOUDOPEN WEIGHTS
This is a source-based starting point. We haven’t independently tested this model yet. Source-based estimates are separate from our test results. Measured speed and our verdict will appear after a reproducible test.

At a glance

Parameters
125B language · 6B active
Architecture
Hybrid sparse attention
Context length
262,144 native
License
Qwen Community License 1.0
Recommended VRAM
Not yet tested
Minimum tested VRAM
Not yet tested
Disk space
Not yet tested
Software
Transformers, vLLM, SGLang
Read the official model documentation
HARDWARE ESTIMATE

What will it need?

Our test notes

Not yet tested. We’ll publish the hardware, quantization, context, speed, peak memory, and load time together. A parameter count alone is not a hardware requirement.

Our verdict

Not yet tested. Check the publisher’s documentation for capabilities and limitations; we’ll add an independent verdict after testing.

Explore its uses

REASONINGMULTIMODALLONG CONTEXT

Keep exploring