We introduce Transect, an open-source tool that helps evaluators follow an agent’s work, identify what to investigate closely, and check their interpretation against the transcript.
Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas, Cozmin Ududec ...
An update on the security changes we have recently made to our frontier AI evaluations, and the work that remains.
Alexandra Souly, Kai Fronsdal, Abby D'Cruz, Xander Davies, Robert Kirk ...
Building on the momentum of the Inspect toolkit, we’re open-sourcing parts of the research stack behind AISI's evaluations. In May 2024, the AI Security Institute open-sourced Inspect AI, a toolkit ...
As AI systems are deployed in high-stakes environments, it becomes increasingly important that they act as their operators intend. Existing red-teaming efforts have demonstrated situations where AI ...
During a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what ...
Automated control monitors could play an important role in overseeing highly capable AI agents that we do not fully trust. Prior work has explored control monitoring in simplified settings, but ...
As AI systems become more powerful and autonomous, ensuring that they remain under human control is one of the central challenges of AI security. If powerful AI systems cannot be reliably expected to ...
See our publications and related blogs below.