As of Oct 9, Goodfire launched activation probes that watch an AI model’s internal signals while an agent runs, escalating to a separate model only when a probe flags risk, and made them available to Baseten customers, per TechCrunch.
Goodfire launched activation probes that watch an AI model’s internal signals while an agent runs, escalating to a separate model only when a probe flags risk, and made them available to Baseten customers, per TechCrunch. Baseten users can select risks such as offensive hacking, chemical/biological misuse, and reward hacking, plus responses from logging to human review or refusal. On Kimi K3 tests, monitoring about 1,500 sessions cost roughly $51 versus $233 for a cheaper external checker and about $10,000 for a top-tier one, with probes catching 94% of malicious hacking sessions and escalating 8.7% of harmless ones, while four probes added under 2% time to first token, Goodfire said.
Items older than 72 hours never go in Top stories and show their original date.