Skip to main content

PrismML

Category: Foundation Models / LLMs

A Caltech spinout developing the world's first commercially viable 1-bit large language models, enabling advanced AI to run locally on edge devices like smartphones with dramatically reduced memory, compute, and energy requirements. PrismML was founded in 2025. The company is led by Babak Hassibi. Based in Pasadena, California, United States. Team size: 11-50. Total funding raised: $16.25M. Latest round: Seed. Key investors include Khosla Ventures, Cerberus Ventures, Caltech.

Founded
2025
Headquarters
Pasadena, California, United States
Team size
11-50
Total funding
$16.25M

Value proposition

Enables 14x smaller, 8x faster, and 5x more energy-efficient AI models than full-precision equivalents without sacrificing reasoning performance, allowing advanced AI to run locally on consumer devices rather than requiring cloud infrastructure

Products and solutions

1-bit Bonsai 8B (8B-parameter 1-bit LLM, 1.15GB footprint), 1-bit Bonsai 4B (4B-parameter, 0.57GB), 1-bit Bonsai 1.7B (1.7B-parameter, 0.25GB), Ternary Bonsai 8B/4B/1.7B (ternary weight variants), Bonsai Image 4B (1-bit and ternary diffusion transformer for image generation), Bonsai Studio (iOS app for local inference)

Unique value

World's first commercially viable 1-bit LLMs — proprietary Caltech-developed mathematical framework compressing neural network weights to a single bit (+1/-1), achieving over 10x the intelligence density of full-precision models while maintaining competitive benchmark performance

Target customer

Smartphone and edge device manufacturers (Apple, etc.), robotics companies, IoT device makers, data center operators, enterprise AI developers seeking cost-efficient inference, and developers building on-device AI applications

Industries served

Edge AI / On-Device AI, Smartphones & Consumer Electronics, Robotics, IoT, Data Center Inference, Automotive, Wearables

Technology advantage

Proprietary 1-bit neural network compression framework based on years of Caltech mathematical research; Straight-Through Estimator implementation for training with 1-bit weights; optimized inference kernels for MLX and llama.cpp backends; trained on Google v4 TPUs; Apache 2.0 open-source release of model weights; exclusive license to Caltech-held patents

How they differentiate

Unlike standard quantization (4-bit, 8-bit) which still uses multi-bit weights, PrismML reduces each weight to a single bit (+1 or -1) through a proprietary mathematical framework developed at Caltech. This enables extreme compression ratios (14x) while preserving reasoning capability — demonstrated by running a 27B-parameter model (Qwen 3.6) from 54GB to under 4GB on an iPhone 17 Pro with all 27B parameters active simultaneously, unlike Apple's sparse architecture which only activates 1-4B at a time.

Main competitors

Google (TurboQuant KV-cache compression), Meta (Llama model family with quantization), Mistral AI (Ministral 3 edge models), Apple (AFM on-device models with sparse architecture), Llamafile (single-file local model distribution)

Key partnerships

Caltech (exclusive IP license, compute grants), Google (compute grants via TPU Research Cloud, trained on Google v4 TPUs), Apple (reportedly in discussions for on-device AI integration), HuggingFace (model distribution), GitHub (open-source release)

Major milestones

March 31, 2026: Emerged from stealth and launched world's first commercially viable 1-bit LLMs (Bonsai 8B/4B/1.7B), April 2026: Open-sourced Bonsai models under Apache 2.0 license, July 2026: Reportedly compressed Alibaba's 27B-parameter Qwen 3.6 from 54GB to under 4GB, running on iPhone 17 Pro — drawing Apple's interest for on-device AI

Market positioning

Pioneer in extreme model compression for edge deployment — positioned at the intersection of foundation model efficiency and edge AI, competing with traditional quantization approaches while enabling entirely new on-device AI capabilities previously limited to cloud infrastructure

Geographic focus

Global (US-headquartered with Caltech roots; Apple partnership interest; open-source global developer community)

Patents and IP

Caltech holds the underlying patents for the 1-bit neural network compression technology; PrismML has been granted an exclusive license from Caltech

About Babak Hassibi

Mose and Lillian S. Bohn Professor of Electrical Engineering and Computing and Mathematical Sciences at Caltech (2001-present); B.S. University of Tehran (1989); M.S. Stanford University (1993); Ph.D. Stanford University (1996); PECASE Awardee (2002); Executive Officer of Caltech EE (2008-2015)

Latest news about PrismML

More Foundation Models / LLMs companies

Official website: