Goodfire Launches Inside-Out Monitors to Catch Rogue AI Agents at a Fraction of the Cost

TL;DR
- Goodfire has unveiled inside-out monitors that detect rogue AI agent behavior by reading the model's internal activations in real time, rather than using a second AI to review every output.
- The system pairs lightweight activation-based detectors with a backup escalation model that only kicks in on suspicious cases, which Goodfire claims cuts monitoring costs dramatically while improving catch rates.
- The launch signals a shift toward scalable, interpretability-driven safety for enterprise AI agents handling sensitive tasks like coding, browsing, and financial actions.
The Outside-In Problem
For years, keeping AI agents in check has meant one expensive strategy: have another AI watch the first AI. Every email drafted, every line of code written, every web click made by an agent gets sent to a larger, smarter monitor model for review.
It works, but it is slow and brutally expensive. If you are running thousands of agents in parallel, you are essentially paying double for compute — once to act, once to audit. As agents move from chatbots to autonomous workers that can browse, buy, book, and deploy code, that outside-in approach simply does not scale.
Goodfire, the interpretability startup known for peering inside neural networks, says there is a better way: watch from the inside out.
Peeking Inside the Model As It Works
Goodfire's new inside-out monitors do not wait for the final output. Instead, they tap directly into the agent model's activations — the internal numerical patterns that fire as the model thinks, plans, and decides.
The idea is simple but powerful. Long before a rogue agent writes a malicious script or exfiltrates data, its internal states already look different from when it is behaving normally. Goodfire trains tiny, specialized probes to recognize those telltale activation signatures for behaviors like deception, power-seeking, unauthorized access attempts, prompt injection susceptibility, and goal drift.
Because these probes are reading internal signals rather than generating text, they run in milliseconds alongside the main model with almost no latency overhead.
Activation-Based Detection, Explained
At the core of the system is what Goodfire calls activation-based detection.
Traditional monitors are generative: they take the full transcript of an agent's actions and ask a frontier model, Is this bad? That requires massive inference on every single step.
Goodfire's detectors are discriminative and sparse. They are small classifiers trained on curated datasets of good versus bad agent trajectories, mapped to internal features discovered through Goodfire's interpretability platform. In practice, each monitor is tuned for a specific risk — for example, one probe watches for deceptive reasoning, another for attempts to bypass permissions, another for covert planning.
The company says this makes detection both earlier and more robust. An agent can learn to sweet-talk an outside reviewer with polite language, but it is much harder to hide what is happening in its own activation space. Early internal tests, according to Goodfire, show higher recall on subtle sabotage and prompt injection attacks compared to prompt-based judges alone.
The Backup Escalation Safety Net
Goodfire is not ditching the second-AI reviewer entirely. It is making it a backup.
The inside-out system uses a two-tier escalation architecture. The lightweight activation probes score every action in real time at near-zero cost. The vast majority of benign traffic — often over 95% of steps — passes straight through.
Only when a probe's suspicion score crosses a threshold does the system escalate that specific trace to a full frontier monitor model for deep review, with full context and chain-of-thought analysis. Customers can tune the threshold based on risk tolerance: stricter for finance and infrastructure agents, looser for low-stakes customer support.
This preserve-the-best-of-both-worlds design is key to Goodfire's pitch: you get the speed and coverage of always-on internal monitoring, plus the nuanced judgment of a large model exactly when you need it.
A Fraction of the Cost
Cost is the headline here. Goodfire claims its inside-out approach can cut agent monitoring bills by up to 90% or more compared to reviewing everything with a top-tier model like GPT-5 or Claude.
The math is straightforward. Instead of paying for millions of extra input and output tokens for constant re-reading, customers pay a tiny compute surcharge for the activation probes, plus full price only for the 2% to 5% of cases that get escalated. For companies running fleets of coding agents or web-browsing assistants 24/7, that difference could be hundreds of thousands of dollars per month.
Goodfire also argues there are hidden savings in latency and user experience. Because probes run inline, safe actions are not delayed by a second model call. Only flagged actions pause for review, keeping agents fast while still contained.
What It Means for Safer, Scalable Agent Deployment
The launch lands at a critical moment. Enterprises want to deploy agents with real permissions — to access codebases, CRM systems, bank accounts, and cloud infrastructure — but security teams are terrified of a single rogue loop causing catastrophic damage.
Cheap, always-on monitoring could unlock that stalemate. If every agent action can be screened for a fraction of a cent, CISOs are far more likely to approve broad deployment. Developers can also build custom monitors for their own policies without training a whole new judge model — just define the bad behavior, and Goodfire helps distill it into an activation probe.
Challenges remain. Interpretability-based probes still need rigorous red-teaming to ensure they generalize to new models and novel attack strategies, and Goodfire will need to prove its detectors work across different model families, not just partner labs. Independent benchmarks will be crucial.
Still, the direction is clear. The industry is moving from watching what AI says to understanding what AI is thinking. If Goodfire is right, the future of AI safety will not be AI watching AI from the outside — it will be AI watching itself from the inside out.
Get All The Latest Updates Delivered Straight To Your Inbox For Free!