
Goodfire launches activation-based AI agent monitors through Baseten
The AMW Read
Activation-based monitoring extends Goodfire's research tooling into agent deployment safeguards, meaningfully updating its product role with potential segment-level impact subject to broader validation.
Goodfire launches activation-based AI agent monitors through Baseten
Goodfire announced on October 8 that its monitoring system is available to Baseten customers. The San Francisco company uses small detectors, called probes, to inspect a model's internal activations at each reasoning step, calling a separate model for closer review when a probe flags suspicious activity. Goodfire says testing on Moonshot AI's Kimi K3 detected 94% of malicious hacking sessions across 1,500 runs, with approximately $51 in total monitoring costs—about 3.4 cents per session. Customers can configure monitors for hacking, chemical or biological weapons misuse, and reward hacking.
The launch moves Goodfire's mechanistic interpretability work toward operational safeguards for autonomous agents. Its earlier Ember API let researchers inspect and steer internal model features; the new product applies that approach to monitoring deployed systems. In the agent market, this addresses a practical constraint: systems that take actions need supervision during execution, while continuous review by another large model adds inference expense. Goodfire's approach concentrates that additional review on flagged activity. The reported detection rate and cost suggest a potentially economical monitoring layer, but remain company-reported results from a specific model and test setting.
For builders, the concrete next step is to evaluate activation monitoring against their own agent workflows before treating it as a safety control. Procurement should examine false positives, missed attacks, intervention latency, and total cost alongside the headline detection rate. Baseten availability provides a distribution channel; the investment question is whether Goodfire can turn its interpretability technology into reliable deployment tooling across customer workloads.
