
Ternary Bonsai 2 27B: A 5.9GB Model for 16GB Laptops
Ternary Bonsai 2 27B squeezes Qwen3.8-27B into 5.9GB with 1.76-bit weights and keeps 98.2% of its benchmark score. Here's how to run it locally.
PrismML has made a 27-billion-parameter vision-language model small enough to fit on a cheap thumb drive. Ternary Bonsai 2 27B, released on September 17, 2026, compresses Alibaba's Qwen3.8-27B into a 5.9GB file by storing almost every weight as one of three values: minus one, zero or plus one. According to PrismML, the result keeps 98.2% of the full model's aggregate benchmark performance while being more than nine times smaller.
- Size: 5.9GB at about 1.76 effective bits per weight, more than nine times smaller than the full-precision base model
- Quality: 83.9 on PrismML's 20-benchmark suite, against 85.4 for Qwen3.8-27B
- Context and modalities: 262K-token context window, with text and image input
- License: Apache 2.0, with GGUF and MLX builds on Hugging Face
How Does Ternary Quantization Work?
Most local LLM users know 4-bit quantization, where each weight is squeezed from 16 bits into four. Ternary quantization goes much further. Each weight becomes -1, 0 or +1, and small groups of weights share an FP16 scaling factor that restores their magnitude. PrismML's 1.76-bit figure is the effective average once those scales are counted.
The payoff is not just disk space. Ternary weights turn many multiplications into simple additions and skips, which is why PrismML also reports energy gains: 0.714 mWh per token on an RTX 4090, which the company says is 40% more efficient than a full-precision 8B model.
Where the Accuracy Goes
PrismML published category-level scores against the Qwen3.8-27B baseline, and the losses are small and uneven:
- Math: 96.57 versus 97.06
- Coding: 81.58 versus 82.17
- Reasoning: 83.95 versus 86.66
- Vision: 78.59 versus 81.64
Math and coding barely move. Reasoning and vision give up about three points each, which is an honest trade for a model a ninth of the size.
What Hardware Do You Need to Run Ternary Bonsai 2 27B?
This is where the news gets exciting for self-hosters. MarkTechPost reports a 16GB laptop as the practical minimum for local deployment. PrismML lists decode speeds at batch size one of about 143 tokens per second on an RTX 5090 and 46.8 tokens per second on an Apple M5 Max, and MarkTechPost adds 96.7 tokens per second on an RTX 4090 and 27.7 on an M5 Pro.
Formats cover both major local stacks:
- GGUF: a 5.93GB PTQ1_0 file and a 7.25GB PQ2_0 file
- MLX: a 2-bit pack for Apple Silicon Macs
- Browser: a WebGPU demo for trying the model without installing anything
One practical note: MarkTechPost reports that the GGUF files need PrismML's own fork of llama.cpp, because stock llama.cpp cannot yet load the new formats. Mac users can take the MLX route instead, and our step-by-step MLX guide walks through the setup.
Why This Open-Weight Model Matters
PrismML's first release showed that a 1-bit LLM could run on a smartphone. The second generation applies that idea to a genuinely capable multimodal model with a long context window, built on the same base we covered in our look at Qwen3.8-27B running locally. A 27B model usually needs around 16GB of memory even at 4-bit; at 5.9GB, this one leaves plenty of room on mainstream laptops for the context window and your other apps.
The Apache 2.0 license also means startups can build commercial products on it without negotiating terms. Expect other labs to take note, because ternary compression that keeps 98% of quality is a compelling path for on-device AI. More open-weight releases live in our AI section.
Sources: PrismML — September 17, 2026; MarkTechPost — September 18, 2026; Hugging Face model page — September 2026.
More AI Stories

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

Meta Muse Gadgets SDK: Build Your Own AI Agent Hardware
Meta's open-source Muse Gadgets SDKs bring its Muse agent to Raspberry Pi and ESP32 builds, and 5,000 free Home Link dongles are going to subscribers.

AstaBrief 8B: Ai2's Open Model for Cited Science Reports
AstaBrief 8B is Ai2's Apache 2.0 open model that writes cited research reports in 51 seconds, 3.5x faster than its Claude pipeline. Here is how it works.
