PrismML
Category: Foundation Models / LLMs
A Caltech spinout developing the world's first commercially viable 1-bit large language models, enabling advanced AI to run locally on edge devices like smartphones with dramatically reduced memory, compute, and energy requirements. PrismML was founded in 2025. The company is led by Babak Hassibi. Based in Pasadena, California, United States. Team size: 11-50. Total funding raised: $16.25M. Latest round: Seed. Key investors include Khosla Ventures, Cerberus Ventures, Caltech.
- Founded
- 2025
- Headquarters
- Pasadena, California, United States
- Team size
- 11-50
- Total funding
- $16.25M
Value proposition
Enables 14x smaller, 8x faster, and 5x more energy-efficient AI models than full-precision equivalents without sacrificing reasoning performance, allowing advanced AI to run locally on consumer devices rather than requiring cloud infrastructure
Products and solutions
1-bit Bonsai 8B (8B-parameter 1-bit LLM, 1.15GB footprint), 1-bit Bonsai 4B (4B-parameter, 0.57GB), 1-bit Bonsai 1.7B (1.7B-parameter, 0.25GB), Ternary Bonsai 8B/4B/1.7B (ternary weight variants), Bonsai Image 4B (1-bit and ternary diffusion transformer for image generation), Bonsai Studio (iOS app for local inference)
Unique value
World's first commercially viable 1-bit LLMs — proprietary Caltech-developed mathematical framework compressing neural network weights to a single bit (+1/-1), achieving over 10x the intelligence density of full-precision models while maintaining competitive benchmark performance
Target customer
Smartphone and edge device manufacturers (Apple, etc.), robotics companies, IoT device makers, data center operators, enterprise AI developers seeking cost-efficient inference, and developers building on-device AI applications
Industries served
Edge AI / On-Device AI, Smartphones & Consumer Electronics, Robotics, IoT, Data Center Inference, Automotive, Wearables
Technology advantage
Proprietary 1-bit neural network compression framework based on years of Caltech mathematical research; Straight-Through Estimator implementation for training with 1-bit weights; optimized inference kernels for MLX and llama.cpp backends; trained on Google v4 TPUs; Apache 2.0 open-source release of model weights; exclusive license to Caltech-held patents
How they differentiate
Unlike standard quantization (4-bit, 8-bit) which still uses multi-bit weights, PrismML reduces each weight to a single bit (+1 or -1) through a proprietary mathematical framework developed at Caltech. This enables extreme compression ratios (14x) while preserving reasoning capability — demonstrated by running a 27B-parameter model (Qwen 3.6) from 54GB to under 4GB on an iPhone 17 Pro with all 27B parameters active simultaneously, unlike Apple's sparse architecture which only activates 1-4B at a time.
Main competitors
Google (TurboQuant KV-cache compression), Meta (Llama model family with quantization), Mistral AI (Ministral 3 edge models), Apple (AFM on-device models with sparse architecture), Llamafile (single-file local model distribution)
Key partnerships
Caltech (exclusive IP license, compute grants), Google (compute grants via TPU Research Cloud, trained on Google v4 TPUs), Apple (reportedly in discussions for on-device AI integration), HuggingFace (model distribution), GitHub (open-source release)
Major milestones
March 31, 2026: Emerged from stealth and launched world's first commercially viable 1-bit LLMs (Bonsai 8B/4B/1.7B), April 2026: Open-sourced Bonsai models under Apache 2.0 license, July 2026: Reportedly compressed Alibaba's 27B-parameter Qwen 3.6 from 54GB to under 4GB, running on iPhone 17 Pro — drawing Apple's interest for on-device AI
Market positioning
Pioneer in extreme model compression for edge deployment — positioned at the intersection of foundation model efficiency and edge AI, competing with traditional quantization approaches while enabling entirely new on-device AI capabilities previously limited to cloud infrastructure
Geographic focus
Global (US-headquartered with Caltech roots; Apple partnership interest; open-source global developer community)
Patents and IP
Caltech holds the underlying patents for the 1-bit neural network compression technology; PrismML has been granted an exclusive license from Caltech
About Babak Hassibi
Mose and Lillian S. Bohn Professor of Electrical Engineering and Computing and Mathematical Sciences at Caltech (2001-present); B.S. University of Tehran (1989); M.S. Stanford University (1993); Ph.D. Stanford University (1996); PECASE Awardee (2002); Executive Officer of Caltech EE (2008-2015)
Latest news about PrismML
More Foundation Models / LLMs companies
Official website: https://prismml.com