PrismML Releases Bonsai 2 27B: A Highly Compressed LLM with Near-Lossless Performance
PrismML, a startup founded by Caltech researchers, released Ternary Bonsai 2 27B on September 17, a highly compressed large language model (LLM) that reduces Alibaba's Qwen3.8 27B model to 5.9 GB while retaining 98.2% of its original performance. For broader context, explore our AI News.
Significant Compression and Performance Retention
The core innovation of Bonsai 2 27B lies in its compression technique. It reduces the Alibaba Qwen3.8 27B model from its original size to just 5.9 GB. This substantial reduction is achieved through the use of ternary weights, which are limited to values of -1, 0, or +1, combined with FP16 group-wise scaling at an effective 1.76 bits per weight. This method allows for a dramatic decrease in memory requirements without a proportional loss in capability. For broader context, explore our Top 100 AI Tools.
Compared to its predecessor, the first Bonsai model, which retained 95% of the original model's performance, Bonsai 2 27B improves this to 98.2%. This near-lossless performance retention is a key factor in its potential utility, as it suggests that the compressed model can perform complex tasks almost as effectively as its much larger counterpart.
Technical Specifications and Capabilities
Bonsai 2 27B offers several notable features:
- Context Window: It supports a 262K-token context window, enabling the processing of extensive inputs and generating detailed outputs.
- Multimodal Input: The model is designed to handle multimodal inputs, expanding its potential applications beyond text-only tasks.
- Throughput: Performance benchmarks indicate high efficiency. On an NVIDIA RTX 5090, the model can achieve throughputs of up to 143 tokens per second. For Apple M5 Max devices, it reaches 46.8 tokens per second.
- Licensing: Bonsai 2 27B is released under the Apache 2.0 license, promoting its use and integration into various projects.
The model's efficiency is further highlighted by its energy consumption, reportedly 0.714 mWh per token, making it 40% more efficient than a full-precision 8B model.
PrismML's Background and Impact
PrismML was founded by researchers from Caltech, including CEO Babak Hassibi. The company has secured a $22.25 million seed funding round from investors such as Khosla Ventures, Cerberus Capital, and Caltech, indicating significant backing for its approach to AI model compression.
The success of the first Bonsai model, which accumulated over 11 million downloads, demonstrates a clear demand for efficient and compact LLMs. Bonsai 2 27B builds on this foundation, offering enhanced performance within an even smaller footprint. This could have implications for deploying advanced AI on edge devices, in resource-constrained environments, or for applications requiring rapid inference.
Conclusion
The release of PrismML's Ternary Bonsai 2 27B marks a notable advancement in the field of efficient large language models. By achieving near-lossless performance with a significantly reduced memory footprint, the model addresses key challenges related to AI deployment and accessibility. Its technical specifications, including a large context window, multimodal capabilities, and high throughput on various hardware, position it as a relevant development for developers and organizations seeking to use powerful AI more efficiently.
Sources
- website/content/blog/bonsai-27b-rtx-pro-6000-dgx-spark.md at main · kubesimplify/website · GitHub
- prism-ml/Ternary-Bonsai-2-27B-mlx-2bit · Hugging Face
- prism-ml/Ternary-Bonsai-2-27B-gguf · Hugging Face
- PrismML hopes its tiny LLM will change how we all use AI | TechCrunch
- PrismML — Introducing Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Recommended AI tools
DeepSeek
Conversational AI
Efficient open-weight AI models for advanced reasoning and research
n8n
Productivity & Collaboration
Open-source workflow automation with native AI
Notebook LLM
Productivity & Collaboration
Turn complexity into clarity with your AI-powered research and thinking partner
JanitorAI
Conversational AI
Create, share, and roleplay with fully customizable AI characters—your stories, your rules.
AutoGPT
Productivity & Collaboration
Build, deploy, and manage autonomous AI agents – automate anything, effortlessly.
Apify
Data Analytics
Automate Anything
About the Author

Albert Schaper is a co-founder of Best-AI.org. He focuses on product strategy, AI adoption, practical tool selection, and educational content that helps users compare AI products with clearer context.
More from AlbertWas this article helpful?
Found outdated info or have suggestions? Send us a note.