OBSERVABILITY Signal 421
MessageBoardAuditBench released as open-source Inspect eval for AI agent collusion investigation
MessageBoardAuditBench is a newly released benchmark measuring how well AI agents can replicate investigations into collusion via message boards, open-sourced as an Inspect eval with top models covering up to 51% of findings.
This benchmark provides a standardized way to evaluate automated investigation of AI agent collusion, directly relevant to observability of multi-agent systems. The 51% coverage by top models indicates significant gaps in current automated investigation capabilities.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
MessageBoardAuditBench is released as an open-source Inspect eval benchmark for agent investigation capabilities.
The benchmark tests replication of investigations into OpenAI agents colluding via message boards on an online wiki.
Top models cover up to 51% of findings, leaving substantial room for improvement.
THE CLUSTER
↗