Bonsai 27B: A 27B-Class model that runs on a phone(prismml.com)
694 points by xenova 57 days ago | 249 comments
tl;dr: PrismML released Bonsai 27B, a compressed version of Qwen3.6 27B using ternary (1.71 bits/weight, 5.9GB) or binary (1.125 bits/weight, 3.9GB) weights, enabling a 27B-class multimodal model to run on an iPhone 17 Pro. The variants retain 95% and 90% of full-precision benchmark performance respectively, with math and coding capabilities nearly untouched, and reach up to 163 tok/s on an RTX 5090. Weights are available under Apache 2.0 with native support for Apple MLX and NVIDIA CUDA.
HN Discussion:
  • Requests comparison with competing small models like Gemma 4 12B QAT to validate claims
  • Excited about ternary/binary quantization breakthrough enabling local model use
  • ~Questions clarity of comparisons and impact on tool calling performance
  • Shares hands-on benchmark/testing results revealing implementation issues
  • Skeptical of demo quality, noting factual errors in showcased output