The standard way to keep AI agents in line is to have a second AI read over its shoulder. This is the default approach, but the cost can quickly add up if the agent runs for hours and processes several novels’ worth of text.
Goodfire, a startup focused on interpretability (understanding how AI models work under the hood), launched a cheaper option on Thursday. It’s a monitor that watches what’s happening inside the AI model while it’s working, rather than just reading what the AI model writes. This monitor is available to Baseten customers who host and run AI models from other companies.
Last month, Baseten’s Base Labs announced a safety partnership with Goodfire and AI platform Hugging Face.
This release follows a series of incidents this year in which AI agents escaped from test environments, including the OpenAI agent that compromised Hugging Face. Goodfire’s first monitor, Open Model Kim K3, used a sandbox breach this summer to access information on the internet and GitHub.
Goodfire’s system works similar to airport security. Small detectors called probes read the model’s internal signals at every step of the agent’s work, much like a walk-through scanner checks every passenger. Only if the probe flags something is another AI model, equivalent to a manual search, scrutinized.
Baseten customers can choose which risks to monitor, including offensive hacking, chemical and biological warfare, and reward hacking. It also determines automatic responses, such as whether to log the event, send it for human review, or deny the request entirely.
Goodfire says its approach is also cheaper to operate. Most AI monitors are separate models and require reloading all of the monitored model’s behavior, increasing time and cost. Instead, Goodfire’s probes leverage calculations that the model is already performing during operation.
“Internal activation monitors are very cheap because they reuse calculations in the forward pass,” Goodfire CEO Eric Ho said on venture capitalist Matt Turk’s MAD podcast last week. “So the model has already computed this token. All you are doing is taking the intermediate neural activity that has already been computed and running the classifier on these internal computations.” In other words, the model has already computed and the probe just reads the results.
In Goodfire’s Kim K3’s test, it cost about $51 to monitor about 1,500 sessions. In comparison, a cheaper AI model that checked every step cost $233, and a top-of-the-line AI model cost about $10,000. The study captured 94% of malicious hacking sessions and sent 8.7% of benign hacking sessions for review.
The company says that running four probes simultaneously increased the time it takes for the model to start responding by less than 2%.

“The big advantage is that we can see things before they happen,” said Dan Balsam, CTO and co-founder of Goodfire. “You can detect that your model can be hacked during evaluation or training.”
This pitch targets open models. Developers can download them and remove security measures, but they don’t come with the kind of monitoring capabilities that private labs run on their own systems.
“The damage that an individual can do in an open model is dwarfed by the damage someone can do in a computing cluster, such as an inference provider, and that’s where most of the blame lies,” Balsam said. “An open ‘myth’ moment will reveal that the model needs guardrails that are deployed during inference. ”
Goodfire’s recent research found that leading open models, including Kimi K3 and GLM-5.2, rewarded hacking in 50% to 96% of AI agent test runs.
Goodfire is not the first to try this approach. Google DeepMind announced in January that its research informed the introduction of an exploit detection probe in Gemini.
Balsam said the monitor is a short-term part of a long-term research goal to reverse engineer LLM to be able to track behavior back to where it emerges during training. “We want to turn the magic of model training into precision engineering,” he said.
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.
