Skip to main content

Wafer

Category: AI Infrastructure

AI inference provider that uses autonomous AI agents to optimize GPU kernels and serving stacks, making non-Nvidia (AMD) chips practical for fast open-source LLM inference. Wafer was founded in 2025. The company is led by Emilio Andere. Based in San Francisco, USA. Team size: 1-10. Total funding raised: $44M. Latest round: Series A. Key investors include Marathon Management Partners, Chemistry, Fifty Years, Y Combinator, Liquid 2 Ventures, Wing Venture Capital, AMD Ventures, Outset Capital.

Founded
2025
Headquarters
San Francisco, USA
Team size
1-10
Total funding
$44M

Value proposition

Wafer builds autonomous AI "performance engineers" that profile and tune inference workloads across model, engine, kernel, and hardware, delivering 2-2.8x speedups over stock baselines (SGLang/vLLM) and making AMD MI355X reach ~80% of Nvidia B200 throughput at less than half the cost. Ships as Wafer Pass, a flat-rate API subscription to the fastest optimized open-source LLMs.

Products and solutions

Wafer Pass (flat-rate API subscription to optimized open-source LLMs, plans from $10/week), autonomous AI performance engineer that rewrites serving stack (kernels, batching, scheduling, memory layout), serverless and dedicated inference for open-source LLMs (Qwen3.5-397B, GLM5.1, DeepSeek V4 Pro).

Unique value

"AI that optimizes AI" — autonomous agents replace teams of specialized performance engineers, automating GPU kernel optimization and making non-Nvidia (AMD) inference chips viable, with 2-2.8x speedups over stock serving stacks.

Target customer

AI coding agent developers and teams using agentic coding harnesses (Claude Code, OpenClaw, Cline, Kilo Code, Roo Code, OpenHands, Conductor); builders of AI agents and coding tools needing a fast default model layer.

Industries served

AI infrastructure / inference, AI developer tools, agentic coding

Technology advantage

Autonomous AI agents that profile workloads and search optimal deployment across model/engine/kernel/hardware; focus on non-Nvidia chipsets (AMD MI355X) in the fragmented inference market; optimizes the deployment stack itself rather than shipping its own inference engine or serving platform.

How they differentiate

Unlike Inferact (vLLM), RadixArk (SGLang), and Baseten (serving platform), Wafer optimizes the deployment stack itself with autonomous AI agents rather than shipping its own inference engine or serving platform, and specifically targets non-Nvidia (AMD) hardware to make it practical.

Main competitors

Inferact (vLLM-based inference), RadixArk (SGLang commercial spinout), Baseten (model deployment/serving platform)

Key partnerships

AMD Ventures (investor), NVIDIA Inception Program, Y Combinator (S25), Wilson Sonsini (legal advisor)

Notable customers

Neon Health, DigitalOcean, Brilliant, Y Combinator

Major milestones

YC Summer 2025 batch, launched Wafer Pass (flat-rate open-source LLM API), $4M seed (Apr 2026), $40M Series A at $200M+ valuation (Sep 2026), turned down acquisition offers from multiple cloud providers, AMD MI355X reached ~80% of Nvidia B200 throughput at less than half cost in July 2026 testing.

Growth metrics

$8M ARR within ~4 months of launching inference cloud (as of Sep 2026); 2T+ tokens of continual inference served

Market positioning

Early-stage AI inference optimization layer positioned to make non-Nvidia chips practical; valued at $200M+ after a 50x valuation jump from seed to Series A in five months.

Geographic focus

United States (San Francisco); global inference market

About Emilio Andere

Ex-Argonne National Laboratory, UChicago Sand Lab, Elicit; Mathematics degree at University of Chicago. Co-founded Wafer (YC S25).

Latest news about Wafer

More AI Infrastructure companies

Official website: