Bonsai 27B is a significant milestone in the development of AI models, as it is the first 27B-class model to run on a phone [1]. This achievement enables local execution of AI workloads, reducing the dependence on cloud APIs and improving the overall efficiency of AI systems. The model comes in two variants: Ternary Bonsai 27B and 1-bit Bonsai 27B, with the latter being footprint-oriented and suitable for phone-class devices [2].

Technical Details

The Ternary Bonsai 27B variant uses ternary weights with FP16 group-wise scaling, giving a true 1.71 effective bits per weight [3]. This results in a model size of 5.9 GB, making it suitable for laptop-class devices. On the other hand, the 1-bit Bonsai 27B variant uses binary weights with the same group-wise scaling, giving 1.125 effective bits per weight [4]. This results in a model size of 3.9 GB, making it suitable for phone-class devices.

Benchmark Results

The benchmark results for Bonsai 27B are impressive, with the Ternary variant retaining 95% of the full-precision baseline and the 1-bit variant retaining 90% [5]. The model has been tested on a range of benchmarks, including math, coding, and vision tasks, and has demonstrated its capability in these areas.

Conclusion

Bonsai 27B is a significant achievement in the field of AI, enabling local execution of AI workloads on phones and laptops. The model's ability to run on a phone makes it an attractive option for developers and users who want to leverage the power of AI without relying on cloud APIs. With its impressive benchmark results and small model size, Bonsai 27B is poised to revolutionize the way we interact with AI systems.

Sources

  1. PrismML. (2026). Announcing Bonsai 27B: The First 27B-Class Model to Run on a Phone.
  2. byteiota. (2026). Ternary Bonsai: 1.58-Bit AI Runs on iPhone at 27 Tok/Sec.
  3. PrismML. (2026). Bonsai 27B Whitepaper.
  4. markmancapitalinsight. (2026). Bonsai 8B: The First Commercially Viable 1-bit LLM.
  5. PrismML. (2026). Bonsai 27B Benchmark Results.