PrismML released Bonsai 2, a large language model that fits on PCs and smartphones.
The team compressed Qwen3.8 27B down to just 5.9 GB of memory. This represents a nine to ten times reduction in storage compared to the original.
The startup was founded by Caltech researchers led by Professor Babak Hassibi. Their compression technique uses ternary weights instead of standard sixteen-bit values. Each weight now stores only +1, -1, or 0 to save space.
Bonsai 2 matches ninety-eight percent of the original model's benchmark scores. The previous version reached ninety-five percent and has been downloaded over eleven million times. PrismML aims to release larger models in the coming months.
Source: https://techcrunch.com/2026/09/17/prismml-hopes-its-tiny-llm-could-change-how-we-all-use-ai/


