Qwen3.5 · 9B
An open-weight vision-language model with a compact Ollama package suitable for a private assistant or coding help.
At a glance
- Parameters
- 9B
- Architecture
- Hybrid attention
- Context length
- 262,144 native
- License
- Apache 2.0
- Disk space
- 6.6 GB
- Software
- Ollama, MLX, Transformers, vLLM
Will it run on your machine?
Save your machine to see a personalized rating and its reasoning.
Add my hardwareBest for
- Private chat and coding help on a single computer.
- Image understanding when the chosen runtime and memory support it.
Tradeoffs
- Vision input and very long context need separate memory validation.
Ways to run it
Ollama · Windows, macOS, Linux
Ollama lists a 6.6 GB package. Text-only memory estimates do not cover vision input.
ollama run qwen3.5:9bExplore its uses
Get it running
Set up a private AI chat assistant
Install a local runtime, run Qwen3.5 9B, confirm responses, and know when to choose the smaller 4B package.
View workflowAsk your own documents with local AI
Pair Qwen3.5 in Ollama with Open WebUI’s document retrieval, then check that answers cite your uploaded file.
View workflowKeep exploring
Bonsai 2 · 27B
A compressed 27B-class reasoning model with publisher-provided GGUF packs and a dedicated llama.cpp fork for CUDA, Metal, and CPU.
DeepSeek V4.1 Flash
A multimodal reasoning model with a compressed key-value cache, published as open weights.
Qwen3.8-Flash-Next
An experimental open-weight multimodal model with sparse attention and 262K native context.