Skip to main content
Back to News
Goodfire prices interpretability probes as a cheaper alternative to LLM-judge monitoring for AI agents
Technology
2 min read
US

Goodfire prices interpretability probes as a cheaper alternative to LLM-judge monitoring for AI agents

The AMW Read

Interpretability-based runtime monitoring is a new architecture next to gateways and sidecars, and it comes with real cost figures. Adoption depends on access to open weights, which keeps the impact at the segment level.
NoveltySignificance
AI Agents · Player MapSafety / Alignment

Goodfire prices interpretability probes as a cheaper alternative to LLM-judge monitoring for AI agents

On October 8, Goodfire launched "inside-out monitors," activation probes that read a model's internal neural activations during the forward pass and run a classifier on them. The company puts the cost at about $185 per 1 million exchanges. It compares that with about $5,420 for standard AI monitors and up to $200,000 for top-tier options. It reports a 93% catch rate on malicious hacking sessions, a 5.5% false-positive rate and under 2% added latency with four probes running at once. Baseten customers can choose which risks to watch, including offensive hacking, chemical or biological misuse and reward hacking. They can then set the response: logging, human review or refusal.

The headline says 50x, but the numbers in the article work out to about 29x against standard monitors. The 50x figure only holds against an unstated LLM-judge baseline, so buyers should ask which comparison applies to them. Our earlier report on this launch cited 94% detection at about 3.4 cents per session, so the published metrics are still moving. The bigger point is the architecture. Gateway and sidecar products like the one behind Rein Security's $25M Series A watch inputs and outputs from outside the model. Goodfire reuses computation the model has already done. That makes monitoring every call cheap enough to be practical. Google DeepMind's use of similar misuse-detection probes in Gemini earlier this year suggests the method works in production.

For builders, the catch is access. Activation probes need the model's internal activations, so they work on open-weight models served on your own or a partner's infrastructure, like Kimi K3 via Baseten. They don't work on closed APIs. That puts Goodfire, which has raised about $207M including a $150M Series B led by B Capital, in a better spot with open-weight deployments than with teams building on proprietary model endpoints.

#Goodfire #Interpretability #AgentSecurity #AISafety #Baseten #OpenWeights

#Goodfire#interpretability#activation probes#agent monitoring#Baseten#reward hacking

How This Connects

Based on AI Agents · Player Map

  1. 1h agoGoodfire prices interpretability probes as a cheaper alternative to LLM-judge monitoring for AI agents · THIS ARTICLE
  2. 1w agoManus launches Manus 2.0 and Cue personal-agent appManus
  3. 1w agoNvidia Expands OpenShell and Introduces Sentry for AI Agent SecurityNvidia
  4. 1w agoSalesforce Agentforce Faces SalesBleed Data-Theft and Phishing Flaw ClaimsSalesforce
  5. 3w agoHugging Face details OpenAI agent intrusion: 17,600 actions over 4.5 daysHugging Face
  6. 0mo agoHiddenLayer closes $100M Series B to harden AI agent and model security across the enterprise stack.HiddenLayer

Related News

More news from Goodfire

Stay updated with the latest news and announcements from Goodfire.

View all Goodfire news

Discover AI Startups

Explore 5,000+ AI companies with VC-grade analysis, funding data, and investment insights.

Explore Dashboard