Text / Prism ML

Bonsai 2 · 27B

A compressed 27B-class reasoning model with publisher-provided GGUF packs and a dedicated llama.cpp fork for CUDA, Metal, and CPU.

Verified sourceReleased Sep 17, 2026Source checked 9/23/2026
LOCALRENTED GPUOPEN WEIGHTS

At a glance

Parameters
27.36B
Architecture
Ternary hybrid attention
Context length
262,144
License
Apache 2.0
Disk space
7.21 GB
Software
PrismML llama.cpp fork
View the model source

Sources and files

Publisher model card and package measurements Publisher runtime and pinned releases
MY HARDWARE

Will it run on your machine?

Save your machine to see a personalized rating and its reasoning.

Add my hardware

Best for

  • Private reasoning and coding on hardware that cannot hold conventional 27B FP16 weights.

Tradeoffs

  • Requires Prism ML’s modified llama.cpp runtime; stock builds do not support these packed weights.
  • The optional vision tower increases memory use for image input.

Ways to run it

PrismML llama.cpp fork · Windows, macOS, Linux

Use the publisher fork or its pinned binary. Stock llama.cpp cannot run the ternary packs correctly.

hf download prism-ml/Ternary-Bonsai-2-27B-gguf Ternary-Bonsai-2-27B-PQ2_0.gguf --local-dir .
Official instructions

Explore its uses

REASONINGCODINGCOMPRESSED WEIGHTS

Get it running

Run a compressed private reasoning assistant

Use Bonsai 2’s dedicated low-bit runtime, download its PQ2_0 package, and verify a local response without accidentally using stock llama.cpp.

View workflow

Keep exploring

Bonsai 2 · 27B: hardware, VRAM & setup | YouRunAI