
Goodfire prices interpretability probes as a cheaper alternative to LLM-judge monitoring for AI agents
The AMW Read
Interpretability-based runtime monitoring is a new architecture next to gateways and sidecars, and it comes with real cost figures. Adoption depends on access to open weights, which keeps the impact at the segment level.
Goodfire prices interpretability probes as a cheaper alternative to LLM-judge monitoring for AI agents
On October 8, Goodfire launched "inside-out monitors," activation probes that read a model's internal neural activations during the forward pass and run a classifier on them. The company puts the cost at about $185 per 1 million exchanges. It compares that with about $5,420 for standard AI monitors and up to $200,000 for top-tier options. It reports a 93% catch rate on malicious hacking sessions, a 5.5% false-positive rate and under 2% added latency with four probes running at once. Baseten customers can choose which risks to watch, including offensive hacking, chemical or biological misuse and reward hacking. They can then set the response: logging, human review or refusal.
The headline says 50x, but the numbers in the article work out to about 29x against standard monitors. The 50x figure only holds against an unstated LLM-judge baseline, so buyers should ask which comparison applies to them. Our earlier report on this launch cited 94% detection at about 3.4 cents per session, so the published metrics are still moving. The bigger point is the architecture. Gateway and sidecar products like the one behind Rein Security's $25M Series A watch inputs and outputs from outside the model. Goodfire reuses computation the model has already done. That makes monitoring every call cheap enough to be practical. Google DeepMind's use of similar misuse-detection probes in Gemini earlier this year suggests the method works in production.
For builders, the catch is access. Activation probes need the model's internal activations, so they work on open-weight models served on your own or a partner's infrastructure, like Kimi K3 via Baseten. They don't work on closed APIs. That puts Goodfire, which has raised about $207M including a $150M Series B led by B Capital, in a better spot with open-weight deployments than with teams building on proprietary model endpoints.

