344
Productivity & Workflow355
Automation & Workflow225
Software Development251
Marketing & Growth193
AI Infrastructure & MLOps175
Writing & Content Creation204
Data & Analytics142
Photography & Imaging156
Design & Creative170
Customer Support133
Sales & Outreach125
Voice & Speech135
Education & Learning131
Operations & Admin87
Goodfire says its new monitoring system can spot risky AI agent behavior by checking the model’s internal signals and only escalating when needed.
In short: Goodfire launched a new way to monitor AI agents that it says is cheaper than having a second AI review everything an agent does.
Goodfire, a startup that studies how AI systems work, announced new “inside-out” monitors for AI agents. AI agents are AI systems that can take steps to complete a task, like searching, writing, and using tools on their own.
Many companies monitor agents by running a second AI model that reads the agent’s actions and outputs, like a supervisor reading over someone’s shoulder. Goodfire’s approach uses small detectors called “probes” that watch signals inside the AI model while it is working, more like an airport scanner that checks every passenger. If a probe flags something risky, a separate AI model can take a closer look.
Goodfire said the monitors are available to customers of Baseten, a company that hosts and runs AI models for other businesses. Baseten customers can pick what they want to watch for, including hacking attempts, misuse related to chemical or biological weapons, and “reward hacking” (when an AI tries to game a test to get a high score instead of following the rules). Customers can also choose what happens next, such as logging the event, sending it to a human reviewer, or refusing the request.
Goodfire shared test results on an open AI model called Kimi K3. It said monitoring around 1,500 sessions cost about $51, compared with $233 for a lower cost full-time AI checker and about $10,000 for a top-tier one. Goodfire said its probes caught 94% of malicious hacking sessions and sent 8.7% of harmless sessions for a second look.
AI agents are being used for longer, more complex tasks, and constant monitoring can get expensive. Cheaper monitoring could make it easier for more companies to add safety checks, especially when using open models that can be downloaded and modified.
Source: TechCrunch AI