Wafer
Category: AI Infrastructure
AI inference provider that uses autonomous AI agents to optimize GPU kernels and serving stacks, making non-Nvidia (AMD) chips practical for fast open-source LLM inference. Wafer was founded in 2025. The company is led by Emilio Andere. Based in San Francisco, USA. Team size: 1-10. Total funding raised: $44M. Latest round: Series A. Key investors include Marathon Management Partners, Chemistry, Fifty Years, Y Combinator, Liquid 2 Ventures, Wing Venture Capital, AMD Ventures, Outset Capital.
- Founded
- 2025
- Headquarters
- San Francisco, USA
- Team size
- 1-10
- Total funding
- $44M
Value proposition
Wafer builds autonomous AI "performance engineers" that profile and tune inference workloads across model, engine, kernel, and hardware, delivering 2-2.8x speedups over stock baselines (SGLang/vLLM) and making AMD MI355X reach ~80% of Nvidia B200 throughput at less than half the cost. Ships as Wafer Pass, a flat-rate API subscription to the fastest optimized open-source LLMs.
Products and solutions
Wafer Pass (flat-rate API subscription to optimized open-source LLMs, plans from $10/week), autonomous AI performance engineer that rewrites serving stack (kernels, batching, scheduling, memory layout), serverless and dedicated inference for open-source LLMs (Qwen3.5-397B, GLM5.1, DeepSeek V4 Pro).
Unique value
"AI that optimizes AI" — autonomous agents replace teams of specialized performance engineers, automating GPU kernel optimization and making non-Nvidia (AMD) inference chips viable, with 2-2.8x speedups over stock serving stacks.
Target customer
AI coding agent developers and teams using agentic coding harnesses (Claude Code, OpenClaw, Cline, Kilo Code, Roo Code, OpenHands, Conductor); builders of AI agents and coding tools needing a fast default model layer.
Industries served
AI infrastructure / inference, AI developer tools, agentic coding
Technology advantage
Autonomous AI agents that profile workloads and search optimal deployment across model/engine/kernel/hardware; focus on non-Nvidia chipsets (AMD MI355X) in the fragmented inference market; optimizes the deployment stack itself rather than shipping its own inference engine or serving platform.
How they differentiate
Unlike Inferact (vLLM), RadixArk (SGLang), and Baseten (serving platform), Wafer optimizes the deployment stack itself with autonomous AI agents rather than shipping its own inference engine or serving platform, and specifically targets non-Nvidia (AMD) hardware to make it practical.
Main competitors
Inferact (vLLM-based inference), RadixArk (SGLang commercial spinout), Baseten (model deployment/serving platform)
Key partnerships
AMD Ventures (investor), NVIDIA Inception Program, Y Combinator (S25), Wilson Sonsini (legal advisor)
Notable customers
Neon Health, DigitalOcean, Brilliant, Y Combinator
Major milestones
YC Summer 2025 batch, launched Wafer Pass (flat-rate open-source LLM API), $4M seed (Apr 2026), $40M Series A at $200M+ valuation (Sep 2026), turned down acquisition offers from multiple cloud providers, AMD MI355X reached ~80% of Nvidia B200 throughput at less than half cost in July 2026 testing.
Growth metrics
$8M ARR within ~4 months of launching inference cloud (as of Sep 2026); 2T+ tokens of continual inference served
Market positioning
Early-stage AI inference optimization layer positioned to make non-Nvidia chips practical; valued at $200M+ after a 50x valuation jump from seed to Series A in five months.
Geographic focus
United States (San Francisco); global inference market
About Emilio Andere
Ex-Argonne National Laboratory, UChicago Sand Lab, Elicit; Mathematics degree at University of Chicago. Co-founded Wafer (YC S25).
Latest news about Wafer
More AI Infrastructure companies
Official website: https://www.wafer.ai