Skip to main content

RadixArk

Category: AI Infrastructure

A high-performance AI inference platform and programming framework designed to accelerate and optimize the deployment of large language models (LLMs) at scale. RadixArk was founded in 2025. The company is led by Ying Sheng. Based in Palo Alto, USA. Team size: 1-10. Total funding raised: $100,000,000. Latest round: Series A ($400M valuation, Jan 2026). Key investors include Accel, Spark Capital, NVentures (NVIDIA), Salience Capital, A&E Investments, HOF Capital, Walden Catalyst Ventures, AMD, LDV Partners, WTT Investment, MediaTek, Lip-Bu Tan.

AMW Analysis

RadixArk is an AI infrastructure startup based in Palo Alto that provides a high-performance inference platform and programming framework for deploying large language models at scale. Its technology focuses on reducing inference latency and operational costs through optimized KV cache management and support for structured, multi-step LLM workflows. The company is the commercial spin-out of the SGLang project and utilizes a RadixAttention mechanism for KV cache reuse, which has delivered efficiency gains for high-scale users such as xAI.

Recent news flow shows rapid scaling and product expansion. RadixArk launched in May 2026 with $100 million in seed funding from investors including Accel, Spark Capital, AMD, MediaTek, and NVentures. By January 2026, the company had reached a $400 million valuation in a round led by Accel. The startup has also introduced its Miles framework, which targets reinforcement learning for Mixture of Experts models, indicating a broadening of its platform beyond core inference optimization.

AMW analysis, generated from 2 tracked news signals.

Founded
2025
Headquarters
Palo Alto, USA
Team size
1-10
Total funding
$100,000,000

Value proposition

Drastically reduces inference latency and operational costs by optimizing KV cache management and enabling structured, multi-step LLM workflows through a specialized programming interface.

Products and solutions

SGLang (Open-Source Inference Engine), Miles (Open-Source RL Post-Training Framework), RadixArk Managed Inference Platform, Managed Training & Fine-Tuning Services

Unique value

Commercializes SGLang, a framework that allows for 'structured generation' where the model's output is controlled and accelerated by a specialized compiler and runtime, rather than just raw token generation.

Target customer

AI labs, enterprise software companies, cloud service providers, and developers building high-throughput LLM applications.

Industries served

Artificial Intelligence Infrastructure, Cloud Computing, Enterprise Software (SaaS), Software Development Tools

Technology advantage

Features 'RadixAttention,' a novel technique for automatic KV cache sharing across multiple requests (prefix caching), and a high-performance runtime that outperforms existing solutions like vLLM in complex, multi-turn interactions.

How they differentiate

RadixArk differentiates through 'RadixAttention,' a novel prefix caching technique that allows for automatic KV cache sharing across multiple requests. Unlike vLLM's block-level hashing, RadixArk's SGLang framework uses a token-level radix tree, significantly reducing latency and compute costs for multi-turn conversations and complex, structured LLM workflows.

Main competitors

vLLM / Inferact, Anyscale, NVIDIA (TensorRT-LLM), Together AI

Key partnerships

UC Berkeley SkyLab (Academic origin), LMSYS Org (Co-founding relationship), Accel (Lead Investor), Spark Capital (Co-Lead Investor), NVIDIA / NVentures (Strategic Investor), AMD (Strategic Investor), Google (Customer), Microsoft (Customer), xAI (Customer)

Notable customers

Google, Microsoft, NVIDIA, Oracle, AMD, xAI (Grok), Cursor (Anysphere), LinkedIn, Thinking Machines Lab, humans&

Major milestones

Open-source launch of SGLang at UC Berkeley SkyLab (2024), Commercial spin-out from UC Berkeley as RadixArk (2025), Secured $100M Seed funding at $400M valuation led by Accel and Spark Capital (May 2026), Integration as core inference engine for xAI's Grok models, Launched Miles, open-source RL training framework (July 2026), SGLang deployed across hundreds of thousands of GPUs serving trillions of tokens daily

Growth metrics

SGLang has seen rapid adoption within the AI developer community over the last six months, becoming a primary alternative to vLLM for high-performance inference.

Market positioning

High-performance AI inference infrastructure provider targeting enterprise AI labs and high-throughput LLM application developers.

Geographic focus

North America (Berkeley/San Francisco based), with a global developer community.

Patents and IP

No specific registered patents disclosed; intellectual property is centered on proprietary optimizations of the SGLang architecture and trade secrets in distributed inference.

About Ying Sheng

Ying Sheng is the co-founder and CEO of RadixArk. She was previously a software engineer at xAI, where she worked on inference systems for Grok, and a research scientist at Databricks. She is a co-founder of LMSYS Org and a primary contributor to the SGLang project. She holds a PhD in Computer Science from Stanford University, where her research focused on high-throughput generative inference.

Latest news about RadixArk

More AI Infrastructure companies

Official website: