Skip to main content

PrismML releases Bonsai 2 27B, compressing Qwen3.8 to a 5.9GB footprint

18 SEPTEMBER 2026·2 MIN READ·3 SOURCES·Independently corroborated

PrismML has launched Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B that fits into 5.9 gigabytes of memory. The startup aims to bring reasoning capabilities to consumer hardware without significant performance degradation.

PrismML releases Bonsai 2 27B, compressing Qwen3.8 to a 5.9GB footprint

Key takeaways · 3

  • 01

    PrismML compressed the Qwen3.8 27B model to 5.9GB, retaining 98.2% of benchmark performance.

  • 02

    The model achieves up to 143 tokens per second on an RTX 5090 GPU.

  • 03

    Running the model currently requires PrismML's proprietary llama.cpp software fork.

Compression and performance

PrismML released Bonsai 2 27B on September 17, 2026. [2][3] The startup, led by CEO Babak Hassibi and backed by a $22.25 million seed round from investors including Khosla Ventures, compressed Alibaba's open source Qwen3.8 27B model down to 5.9 gigabytes. [2] This represents a 9x to 10x reduction in memory compared to the original model. [2] The company reports that the compressed model retains 98.2% of the full-precision version's aggregate benchmark scores. [1][3] PrismML states the model can achieve 143 tokens per second on an RTX 5090 graphics card, though running it currently requires a proprietary fork of the llama.cpp software. [3] The AI lab is betting that reasoning models can be made small enough to fit on standard personal computers and smartphones. [2]

What it means

The release of Bonsai 2 27B illustrates a continuing push to run capable reasoning models locally on consumer hardware. By matching Alibaba Qwen3.8's performance closely at a fraction of the size, PrismML provides a direct alternative to cloud-reliant inference. This development compares favorably to previous compression efforts by minimizing benchmark degradation, offering a viable path for on-device processing. What the sources don't address: whether independent real-world testing will consistently validate the company's high throughput claims across diverse consumer hardware setups.

The ability to run 27B-class models on local hardware without significant quality loss dramatically lowers inference costs and privacy risks. Developers can increasingly build offline, on-device AI applications that previously required cloud infrastructure.

Why it matters
Daily session

Put this to work — one session a day, built for your industry.

Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.

Start free

How this developed

  1. 18 September 2026

    PrismML releases Bonsai 2 27B, compressing Qwen3.8 to a 5.9GB footprint

  2. 18 September 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.