PrismML releases Bonsai 2 27B, compressing Qwen3.8 to a 5.9GB footprint
PrismML has launched Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B that fits into 5.9 gigabytes of memory. The startup aims to bring reasoning capabilities to consumer hardware without significant performance degradation.

Key takeaways · 3
- 01
PrismML compressed the Qwen3.8 27B model to 5.9GB, retaining 98.2% of benchmark performance.
- 02
The model achieves up to 143 tokens per second on an RTX 5090 GPU.
- 03
Running the model currently requires PrismML's proprietary llama.cpp software fork.
Compression and performance
PrismML released Bonsai 2 27B on September 17, 2026. [2][3] The startup, led by CEO Babak Hassibi and backed by a $22.25 million seed round from investors including Khosla Ventures, compressed Alibaba's open source Qwen3.8 27B model down to 5.9 gigabytes. [2] This represents a 9x to 10x reduction in memory compared to the original model. [2] The company reports that the compressed model retains 98.2% of the full-precision version's aggregate benchmark scores. [1][3] PrismML states the model can achieve 143 tokens per second on an RTX 5090 graphics card, though running it currently requires a proprietary fork of the llama.cpp software. [3] The AI lab is betting that reasoning models can be made small enough to fit on standard personal computers and smartphones. [2]
What it means
The release of Bonsai 2 27B illustrates a continuing push to run capable reasoning models locally on consumer hardware. By matching Alibaba Qwen3.8's performance closely at a fraction of the size, PrismML provides a direct alternative to cloud-reliant inference. This development compares favorably to previous compression efforts by minimizing benchmark degradation, offering a viable path for on-device processing. What the sources don't address: whether independent real-world testing will consistently validate the company's high throughput claims across diverse consumer hardware setups.
The ability to run 27B-class models on local hardware without significant quality loss dramatically lowers inference costs and privacy risks. Developers can increasingly build offline, on-device AI applications that previously required cloud infrastructure.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
18 September 2026
PrismML releases Bonsai 2 27B, compressing Qwen3.8 to a 5.9GB footprint
18 September 2026
Event created from source cluster.
Sources
- PrismML releases Bonsai 2 27B, which compresses Alibaba's Qwen3.8 27B to 5.9 GB, small enough for smartphones, while retaining 98.2% of Qwen's benchmark scores (Julie Bort/TechCrunch)Techmeme
- PrismML hopes its tiny LLM will change how we all use AI | TechCrunchtechcrunch.com
- prismml-bonsai-2-27b-near-lossless-compression-2026explainx.ai