RadixArk
Category: AI Infrastructure
A high-performance AI inference platform and programming framework designed to accelerate and optimize the deployment of large language models (LLMs) at scale. RadixArk was founded in 2025. The company is led by Ying Sheng. Based in Palo Alto, USA. Team size: 1-10. Total funding raised: $100,000,000. Latest round: Series A ($400M valuation, Jan 2026). Key investors include Accel, Spark Capital, NVentures (NVIDIA), Salience Capital, A&E Investments, HOF Capital, Walden Catalyst Ventures, AMD, LDV Partners, WTT Investment, MediaTek, Lip-Bu Tan.
AMW Analysis
RadixArk is an AI infrastructure startup based in Palo Alto that provides a high-performance inference platform and programming framework for deploying large language models at scale. Its technology focuses on reducing inference latency and operational costs through optimized KV cache management and support for structured, multi-step LLM workflows. The company is the commercial spin-out of the SGLang project and utilizes a RadixAttention mechanism for KV cache reuse, which has delivered efficiency gains for high-scale users such as xAI.
Recent news flow shows rapid scaling and product expansion. RadixArk launched in May 2026 with $100 million in seed funding from investors including Accel, Spark Capital, AMD, MediaTek, and NVentures. By January 2026, the company had reached a $400 million valuation in a round led by Accel. The startup has also introduced its Miles framework, which targets reinforcement learning for Mixture of Experts models, indicating a broadening of its platform beyond core inference optimization.
AMW analysis, generated from 2 tracked news signals.
- Founded
- 2025
- Headquarters
- Palo Alto, USA
- Team size
- 1-10
- Total funding
- $100,000,000
Value proposition
Drastically reduces inference latency and operational costs by optimizing KV cache management and enabling structured, multi-step LLM workflows through a specialized programming interface.
Products and solutions
SGLang (Open-Source Inference Engine), Miles (Open-Source RL Post-Training Framework), RadixArk Managed Inference Platform, Managed Training & Fine-Tuning Services
Unique value
Commercializes SGLang, a framework that allows for 'structured generation' where the model's output is controlled and accelerated by a specialized compiler and runtime, rather than just raw token generation.
Target customer
AI labs, enterprise software companies, cloud service providers, and developers building high-throughput LLM applications.
Industries served
Artificial Intelligence Infrastructure, Cloud Computing, Enterprise Software (SaaS), Software Development Tools
Technology advantage
Features 'RadixAttention,' a novel technique for automatic KV cache sharing across multiple requests (prefix caching), and a high-performance runtime that outperforms existing solutions like vLLM in complex, multi-turn interactions.
How they differentiate
RadixArk differentiates through 'RadixAttention,' a novel prefix caching technique that allows for automatic KV cache sharing across multiple requests. Unlike vLLM's block-level hashing, RadixArk's SGLang framework uses a token-level radix tree, significantly reducing latency and compute costs for multi-turn conversations and complex, structured LLM workflows.
Main competitors
vLLM / Inferact, Anyscale, NVIDIA (TensorRT-LLM), Together AI
Key partnerships
UC Berkeley SkyLab (Academic origin), LMSYS Org (Co-founding relationship), Accel (Lead Investor), Spark Capital (Co-Lead Investor), NVIDIA / NVentures (Strategic Investor), AMD (Strategic Investor), Google (Customer), Microsoft (Customer), xAI (Customer)
Notable customers
Google, Microsoft, NVIDIA, Oracle, AMD, xAI (Grok), Cursor (Anysphere), LinkedIn, Thinking Machines Lab, humans&
Major milestones
Open-source launch of SGLang at UC Berkeley SkyLab (2024), Commercial spin-out from UC Berkeley as RadixArk (2025), Secured $100M Seed funding at $400M valuation led by Accel and Spark Capital (May 2026), Integration as core inference engine for xAI's Grok models, Launched Miles, open-source RL training framework (July 2026), SGLang deployed across hundreds of thousands of GPUs serving trillions of tokens daily
Growth metrics
SGLang has seen rapid adoption within the AI developer community over the last six months, becoming a primary alternative to vLLM for high-performance inference.
Market positioning
High-performance AI inference infrastructure provider targeting enterprise AI labs and high-throughput LLM application developers.
Geographic focus
North America (Berkeley/San Francisco based), with a global developer community.
Patents and IP
No specific registered patents disclosed; intellectual property is centered on proprietary optimizations of the SGLang architecture and trade secrets in distributed inference.
About Ying Sheng
Ying Sheng is the co-founder and CEO of RadixArk. She was previously a software engineer at xAI, where she worked on inference systems for Grok, and a research scientist at Databricks. She is a co-founder of LMSYS Org and a primary contributor to the SGLang project. She holds a PhD in Computer Science from Stanford University, where her research focused on high-throughput generative inference.
Latest news about RadixArk
More AI Infrastructure companies
Official website: https://www.radixark.ai/