AISI Blog https://www.aisi.gov.uk/blog Blog posts from the AI Security Institute http://www.rssboard.org/rss-specification python-feedgen en Sun, 23 Aug 2026 01:29:56 +0000 Incident Report: unsanctioned agent behaviour during cyber testing https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing During a routine cyber evaluation, AISI identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations. We are disclosing what we found, what it means, and the actions now underway. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing Tue, 04 Aug 2026 00:00:00 +0000 How our Control Red Team is stress-testing frontier monitors https://www.aisi.gov.uk/blog/how-our-new-control-red-team-is-stress-testing-frontier-monitors Early learnings from red-teaming the internal monitors of frontier AI companies, and our perspectives on the open problems that remain. https://www.aisi.gov.uk/blog/how-our-new-control-red-team-is-stress-testing-frontier-monitors Thu, 23 Jul 2026 00:00:00 +0000 UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities Our joint evaluation with CAISI finds Kimi K3 trails leading US frontier closed weight models on cyber capability. https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities Thu, 23 Jul 2026 00:00:00 +0000 International evaluation best practice and open questions in AI measurement https://www.aisi.gov.uk/blog/international-evaluation-best-practice-and-open-questions-in-ai-measurement The International Network for Advanced AI Measurement, Evaluation and Science convened in Seoul to continue outlining international best practice. https://www.aisi.gov.uk/blog/international-evaluation-best-practice-and-open-questions-in-ai-measurement Thu, 23 Jul 2026 00:00:00 +0000 Cheating behaviour in frontier model evaluations https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations We find cheating behaviour in all of our cyber capability evaluations, and outline the implications as models grow more capable. https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations Tue, 21 Jul 2026 00:00:00 +0000 How Far Behind the Frontier are Leading Open Weight Models on Cyber? https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber We evaluated the cyber capabilities of leading open and closed weight AI models, and found that recent open models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released 4 to 7 months before them – a narrower gap than the 6 to 10 months we measured through most of 2025. https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber Fri, 17 Jul 2026 00:00:00 +0000 Finding Cloud Misconfigurations with Frontier AI: A Case Study https://www.aisi.gov.uk/blog/finding-cloud-misconfigurations-with-frontier-ai-a-case-study A cybersecurity exercise from AISI’s engineering team, using frontier models to test our research platform for misconfigurations. https://www.aisi.gov.uk/blog/finding-cloud-misconfigurations-with-frontier-ai-a-case-study Tue, 07 Jul 2026 00:00:00 +0000 More compute, more capability: Why AI agent evaluations need to account for test-time compute https://www.aisi.gov.uk/blog/more-compute-more-capability-why-ai-agent-evals-need-to-account-for-test-time-compute Standard evaluations cap how much compute AI agents can use. We show that raising those caps changes measured capability, the difficulty of tasks agents can solve, and how fast the frontier appears to move. https://www.aisi.gov.uk/blog/more-compute-more-capability-why-ai-agent-evals-need-to-account-for-test-time-compute Thu, 02 Jul 2026 00:00:00 +0000 UK-Germany Joint Statement on advanced AI safety and security https://www.aisi.gov.uk/blog/uk-germany-joint-statement-on-advanced-ai-safety-and-security A joint statement by the UK and Germany on collaborating to ensure advanced AI is developed safely and its risks are rigorously understood and managed. https://www.aisi.gov.uk/blog/uk-germany-joint-statement-on-advanced-ai-safety-and-security Tue, 30 Jun 2026 00:00:00 +0000 Releasing AISI’s Engineering Playbook https://www.aisi.gov.uk/blog/releasing-aisis-engineering-playbook Building on the momentum of the Inspect toolkit, we’re open-sourcing parts of the research stack behind AISI's evaluations. https://www.aisi.gov.uk/blog/releasing-aisis-engineering-playbook Thu, 18 Jun 2026 00:00:00 +0000 RealityTest: Do AI systems disclose their identity when asked? https://www.aisi.gov.uk/blog/realitytest-do-ai-systems-disclose-their-identity-when-asked A new benchmark grounded in how real users actually probe AI identity during interactions – covering five languages, across text and speech. https://www.aisi.gov.uk/blog/realitytest-do-ai-systems-disclose-their-identity-when-asked Mon, 08 Jun 2026 00:00:00 +0000 Deepening our partnership with the Australian AI Safety Institute https://www.aisi.gov.uk/blog/deepening-our-partnership-with-the-australian-ai-safety-institute An agreement between Institutes to collaborate on best practices in AI evaluation, and share research findings. https://www.aisi.gov.uk/blog/deepening-our-partnership-with-the-australian-ai-safety-institute Mon, 25 May 2026 00:00:00 +0000 Will it become harder to oversee AI systems? https://www.aisi.gov.uk/blog/will-it-become-harder-to-oversee-ai-systems Our new report examines today’s AI oversight landscape, how robust it is to capability advances, and the pathways that could lead to its degradation. https://www.aisi.gov.uk/blog/will-it-become-harder-to-oversee-ai-systems Thu, 21 May 2026 00:00:00 +0000 How fast is autonomous AI cyber capability advancing? https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing The length of tasks frontier models can autonomously complete in our narrow cyber suite has been doubling every few months. This doubling rate has become faster over time, and recent models exceeded our previous trends. https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing Wed, 13 May 2026 00:00:00 +0000 Partnering with Microsoft to strengthen frontier AI safety https://www.aisi.gov.uk/blog/partnering-with-microsoft-to-strengthen-frontier-ai-safety A new collaboration on high-risk capability evaluation, safeguard testing, and societal resilience research. https://www.aisi.gov.uk/blog/partnering-with-microsoft-to-strengthen-frontier-ai-safety Tue, 05 May 2026 00:00:00 +0000 Our evaluation of OpenAI's GPT-5.5 cyber capabilities https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5-5-cyber-capabilities AISI conducted cyber evaluations on OpenAI's GPT-5.5. GPT-5.5 is one of the strongest models we have tested on our cyber tasks and is the second model to solve one of our multi-step cyber-attack simulations end-to-end. https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5-5-cyber-capabilities Thu, 30 Apr 2026 00:00:00 +0000 Ask Don't Tell: Reducing Sycophancy in Large Language Models https://www.aisi.gov.uk/blog/ask-dont-tell-reducing-sycophancy-in-large-language-models-2 Our new paper finds that the phrasing of user inputs affects levels of model sycophancy. https://www.aisi.gov.uk/blog/ask-dont-tell-reducing-sycophancy-in-large-language-models-2 Tue, 28 Apr 2026 00:00:00 +0000 Evaluating whether AI models would sabotage AI safety research https://www.aisi.gov.uk/blog/evaluating-whether-ai-models-would-sabotage-ai-safety-research An update on our alignment testing methodology for recent frontier models. https://www.aisi.gov.uk/blog/evaluating-whether-ai-models-would-sabotage-ai-safety-research Mon, 27 Apr 2026 00:00:00 +0000 How do environmental factors impact AI behaviour? https://www.aisi.gov.uk/blog/how-do-environmental-factors-impact-ai-behaviour Developing methods to better understand when and why AI models sometimes act against user intentions. https://www.aisi.gov.uk/blog/how-do-environmental-factors-impact-ai-behaviour Fri, 24 Apr 2026 00:00:00 +0000 What can sandboxed AI agents learn about their evaluation environments? https://www.aisi.gov.uk/blog/what-can-sandboxed-ai-agents-learn-about-their-evaluation-environments We deployed open-source AI agent OpenClaw inside a sandbox on our research platform. Despite our initial countermeasures, it successfully identified our organisation by name, inferred the identity of a human operator and reconstructed a timeline of some of our research activities. https://www.aisi.gov.uk/blog/what-can-sandboxed-ai-agents-learn-about-their-evaluation-environments Mon, 20 Apr 2026 00:00:00 +0000 Our evaluation of Claude Mythos Preview’s cyber capabilities https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities We conducted cyber evaluations of Anthropic’s Claude Mythos Preview and found continued improvement in capture-the-flag (CTF) challenges and significant improvement on multi-step cyber-attack simulations. https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities Mon, 13 Apr 2026 00:00:00 +0000 Harnessing frontier AI for cyber defence https://www.aisi.gov.uk/blog/harnessing-frontier-ai-for-cyber-defence Sharing work with the National Cyber Security Centre (NCSC) on how cyber defenders can use advanced AI capabilities to stay ahead of attackers. https://www.aisi.gov.uk/blog/harnessing-frontier-ai-for-cyber-defence Tue, 31 Mar 2026 00:00:00 +0000 How are AI Agents used? Evidence from 177,000 AI agent tools https://www.aisi.gov.uk/blog/how-are-ai-agents-used-evidence-from-177000-ai-agent-tools A monitoring method and large‑scale analysis to understand the tasks AI agents are performing today. https://www.aisi.gov.uk/blog/how-are-ai-agents-used-evidence-from-177000-ai-agent-tools Thu, 26 Mar 2026 00:00:00 +0000 Can AI agents escape their sandboxes? A benchmark for safely measuring container breakout capabilities https://www.aisi.gov.uk/blog/can-ai-agents-escape-their-sandboxes-a-benchmark-for-safely-measuring-container-breakout-capabilities We introduce SandboxEscapeBench, the first benchmark to systematically evaluate whether AI agents can break out of their sandboxes, and share some early results. https://www.aisi.gov.uk/blog/can-ai-agents-escape-their-sandboxes-a-benchmark-for-safely-measuring-container-breakout-capabilities Mon, 23 Mar 2026 00:00:00 +0000 How do frontier AI agents perform in multi-step cyber-attack scenarios? https://www.aisi.gov.uk/blog/how-do-frontier-ai-agents-perform-in-multi-step-cyber-attack-scenarios We tested seven large language models (LLMs) on two custom-built cyber ranges, measuring their ability to execute extended attack sequences in complex environments. https://www.aisi.gov.uk/blog/how-do-frontier-ai-agents-perform-in-multi-step-cyber-attack-scenarios Mon, 16 Mar 2026 00:00:00 +0000 Evidence for inference scaling in AI cyber tasks: Increased evaluation budgets reveal higher success rates https://www.aisi.gov.uk/blog/evidence-for-inference-scaling-in-ai-cyber-tasks-increased-evaluation-budgets-reveal-higher-success-rates Alongside Irregular, we found evidence demonstrating that evaluators need to use large token budgets to understand the cyber capabilities of recent Large Language Models (LLMs). https://www.aisi.gov.uk/blog/evidence-for-inference-scaling-in-ai-cyber-tasks-increased-evaluation-budgets-reveal-higher-success-rates Thu, 05 Mar 2026 00:00:00 +0000 An evaluation framework for AI misuse in fraud and cybercrime https://www.aisi.gov.uk/blog/an-evaluation-framework-for-ai-misuse-in-fraud-and-cybercrime We developed a scalable approach to measuring how text-based AI models can assist in three complex fraud and cybercrime scenarios. https://www.aisi.gov.uk/blog/an-evaluation-framework-for-ai-misuse-in-fraud-and-cybercrime Thu, 26 Feb 2026 00:00:00 +0000 A pipeline for transcript analysis using Inspect Scout https://www.aisi.gov.uk/blog/a-pipeline-for-transcript-analysis-using-inspect-scout We outline a step-by-step pipeline for using our open-source transcript analysis tool, Inspect Scout. https://www.aisi.gov.uk/blog/a-pipeline-for-transcript-analysis-using-inspect-scout Wed, 25 Feb 2026 00:00:00 +0000 Funding 60 projects to advance AI alignment research https://www.aisi.gov.uk/blog/funding-60-projects-to-advance-ai-alignment-research The Alignment Project welcomes its first cohort of grantees, and new partners join the coalition, bringing total funding to £27m. https://www.aisi.gov.uk/blog/funding-60-projects-to-advance-ai-alignment-research Thu, 19 Feb 2026 00:00:00 +0000 Advancing AI voice security with ElevenLabs https://www.aisi.gov.uk/blog/advancing-voice-ai-security-with-elevenlabs New partnership exploring the security and societal implications of voice AI systems https://www.aisi.gov.uk/blog/advancing-voice-ai-security-with-elevenlabs Wed, 18 Feb 2026 00:00:00 +0000 Boundary Point Jailbreaking: A new way to break the strongest AI defences https://www.aisi.gov.uk/blog/boundary-point-jailbreaking-a-new-way-to-break-the-strongest-ai-defences Introducing an automated attack technique that generates universal jailbreaks against the best defended systems https://www.aisi.gov.uk/blog/boundary-point-jailbreaking-a-new-way-to-break-the-strongest-ai-defences Tue, 17 Feb 2026 00:00:00 +0000 International consensus and open questions in AI evaluations https://www.aisi.gov.uk/blog/international-ai-network-consensus-and-open-questions The International Network for Advanced AI Measurement, Evaluation and Science reflects on recent meeting and looks ahead to the India AI Impact Summit https://www.aisi.gov.uk/blog/international-ai-network-consensus-and-open-questions Thu, 12 Feb 2026 00:00:00 +0000 AI and the future of work: Measuring AI-driven productivity gains for workplace tasks https://www.aisi.gov.uk/blog/ai-and-the-future-of-work-measuring-ai-driven-productivity-gains-for-workplace-tasks Alongside the government’s new Future of Work Unit, we conducted a pilot study to explore how much AI models increase worker productivity for common tasks. https://www.aisi.gov.uk/blog/ai-and-the-future-of-work-measuring-ai-driven-productivity-gains-for-workplace-tasks Mon, 02 Feb 2026 00:00:00 +0000 Our 2025 year in review https://www.aisi.gov.uk/blog/our-2025-year-in-review Adam Beaumont, Director of the UK AI Security Institute, reflects on the year's biggest achievements. https://www.aisi.gov.uk/blog/our-2025-year-in-review Mon, 22 Dec 2025 00:00:00 +0000 5 key findings from our first Frontier AI Trends Report https://www.aisi.gov.uk/blog/5-key-findings-from-our-first-frontier-ai-trends-report Our inaugural Frontier AI Trends Report draws on 2 years' worth of evaluations to provide accessible insights into the trajectory of AI development. https://www.aisi.gov.uk/blog/5-key-findings-from-our-first-frontier-ai-trends-report Thu, 18 Dec 2025 00:00:00 +0000 Our approach to tackling AI-generated child sexual abuse material https://www.aisi.gov.uk/blog/our-approach-to-tackling-ai-generated-child-sexual-abuse-material How we’re partnering with government and experts to prevent the creation and spread of AI‑generated CSAM https://www.aisi.gov.uk/blog/our-approach-to-tackling-ai-generated-child-sexual-abuse-material Wed, 17 Dec 2025 00:00:00 +0000 Stress-testing asynchronous monitoring of AI coding agents https://www.aisi.gov.uk/blog/stress-testing-asynchronous-monitoring-of-ai-coding-agents Our new paper shares findings from an adversarial evaluation of monitoring systems for detecting sabotage by AI coding agents. https://www.aisi.gov.uk/blog/stress-testing-asynchronous-monitoring-of-ai-coding-agents Tue, 16 Dec 2025 00:00:00 +0000 Deepening our partnership with Google DeepMind https://www.aisi.gov.uk/blog/deepening-our-partnership-with-google-deepmind Expanding our collaboration with a new research MOU https://www.aisi.gov.uk/blog/deepening-our-partnership-with-google-deepmind Thu, 11 Dec 2025 00:00:00 +0000 Auditing games for sandbagging detection https://www.aisi.gov.uk/blog/auditing-games-for-sandbagging-detection Our new paper shares the results of an auditing game to evaluate ten methods for sandbagging detection in AI models. https://www.aisi.gov.uk/blog/auditing-games-for-sandbagging-detection Tue, 09 Dec 2025 00:00:00 +0000 How do AI models persuade? Exploring the levers of AI-enabled persuasion through large-scale experiments https://www.aisi.gov.uk/blog/how-do-ai-models-persuade-exploring-the-levers-of-ai-enabled-persuasion-through-large-scale-experiments A deep dive into AISI’s study of the persuasive capabilities of conversational AI, published today in Science. https://www.aisi.gov.uk/blog/how-do-ai-models-persuade-exploring-the-levers-of-ai-enabled-persuasion-through-large-scale-experiments Thu, 04 Dec 2025 00:00:00 +0000 UKAISI at NeurIPS 2025 https://www.aisi.gov.uk/blog/ukaisi-at-neurips-2025 An overview of the research we’ll be presenting at this year’s NeurIPS conference. https://www.aisi.gov.uk/blog/ukaisi-at-neurips-2025 Wed, 26 Nov 2025 00:00:00 +0000 Investigating models for misalignment https://www.aisi.gov.uk/blog/investigating-models-for-misalignment Insights from our alignment evaluations of Claude Opus 4.1, Sonnet 4.5, and a pre‑release snapshot of Opus 4.5. https://www.aisi.gov.uk/blog/investigating-models-for-misalignment Wed, 26 Nov 2025 00:00:00 +0000 Mapping the limitations of current AI systems https://www.aisi.gov.uk/blog/mapping-the-limitations-of-current-ai-systems Takeaways from expert interviews on barriers to AI capable of automating most cognitive labour. https://www.aisi.gov.uk/blog/mapping-the-limitations-of-current-ai-systems Thu, 23 Oct 2025 00:00:00 +0000 Introducing ControlArena: A library for running AI control experiments https://www.aisi.gov.uk/blog/introducing-controlarena-a-library-for-running-ai-control-experiments Our dedicated library to make AI control experiments easy, consistent, and repeatable. https://www.aisi.gov.uk/blog/introducing-controlarena-a-library-for-running-ai-control-experiments Wed, 22 Oct 2025 00:00:00 +0000 Transcript analysis for AI agent evaluations https://www.aisi.gov.uk/blog/transcript-analysis-for-ai-agent-evaluations Why we use transcript analysis for our agent evaluations, and results from an early case study. https://www.aisi.gov.uk/blog/transcript-analysis-for-ai-agent-evaluations Fri, 10 Oct 2025 00:00:00 +0000 Examining backdoor data poisoning at scale https://www.aisi.gov.uk/blog/examining-backdoor-data-poisoning-at-scale Our work with Anthropic and the Alan Turing Institute suggests that data poisoning attacks may be easier than previously believed. https://www.aisi.gov.uk/blog/examining-backdoor-data-poisoning-at-scale Thu, 09 Oct 2025 00:00:00 +0000 Do chatbots inform or misinform voters? https://www.aisi.gov.uk/blog/do-chatbots-inform-or-misinform-voters What we learned from a large-scale empirical study of AI use for political information-seeking. https://www.aisi.gov.uk/blog/do-chatbots-inform-or-misinform-voters Tue, 30 Sep 2025 00:00:00 +0000 How we’re working with frontier AI developers to improve model security https://www.aisi.gov.uk/blog/how-were-working-with-frontier-ai-developers-to-improve-model-security Insights into our ongoing voluntary collaborations with Anthropic and OpenAI. https://www.aisi.gov.uk/blog/how-were-working-with-frontier-ai-developers-to-improve-model-security Sat, 13 Sep 2025 00:00:00 +0000 From bugs to bypasses: adapting vulnerability disclosure for AI safeguards https://www.aisi.gov.uk/blog/from-bugs-to-bypasses-adapting-vulnerability-disclosure-for-ai-safeguards Exploring how far cyber security approaches can help mitigate risks in generative AI systems, in collaboration with the National Cyber Security Centre (NCSC). https://www.aisi.gov.uk/blog/from-bugs-to-bypasses-adapting-vulnerability-disclosure-for-ai-safeguards Tue, 02 Sep 2025 00:00:00 +0000 Managing risks from increasingly capable open-weight AI systems https://www.aisi.gov.uk/blog/managing-risks-from-increasingly-capable-open-weight-ai-systems Current methods and open problems in open-weight model risk management. https://www.aisi.gov.uk/blog/managing-risks-from-increasingly-capable-open-weight-ai-systems Fri, 29 Aug 2025 00:00:00 +0000 The Inspect Sandboxing Toolkit: Scalable and secure AI agent evaluations https://www.aisi.gov.uk/blog/the-inspect-sandboxing-toolkit-scalable-and-secure-ai-agent-evaluations A comprehensive toolkit for safely evaluating AI agents. https://www.aisi.gov.uk/blog/the-inspect-sandboxing-toolkit-scalable-and-secure-ai-agent-evaluations Thu, 07 Aug 2025 00:00:00 +0000 Announcing the Alignment Project: A global fund of over £15 million for AI alignment research https://www.aisi.gov.uk/blog/announcing-the-alignment-project Announcing the Alignment Project: A global fund of over £15 million for AI alignment research https://www.aisi.gov.uk/blog/announcing-the-alignment-project Wed, 30 Jul 2025 00:00:00 +0000 Navigating the uncharted: Building societal resilience to frontier AI https://www.aisi.gov.uk/blog/navigating-the-uncharted-building-societal-resilience-to-frontier-ai We outline our approach to study and address AI risks in real-world applications https://www.aisi.gov.uk/blog/navigating-the-uncharted-building-societal-resilience-to-frontier-ai Thu, 24 Jul 2025 00:00:00 +0000 International joint testing exercise: Agentic testing https://www.aisi.gov.uk/blog/international-joint-testing-exercise-agentic-testing Advancing methodologies for agentic evaluations across domains, including leakage of sensitive Information, fraud and cybersecurity threats. https://www.aisi.gov.uk/blog/international-joint-testing-exercise-agentic-testing Thu, 17 Jul 2025 00:00:00 +0000 A structured protocol for elicitation experiments https://www.aisi.gov.uk/blog/our-approach-to-ai-capability-elicitation Calibrating AI risk assessment through rigorous elicitation practices. https://www.aisi.gov.uk/blog/our-approach-to-ai-capability-elicitation Wed, 16 Jul 2025 00:00:00 +0000 Why we're working on white box control https://www.aisi.gov.uk/blog/why-were-working-on-white-box-control An introduction to white box control, and an update on our research so far. https://www.aisi.gov.uk/blog/why-were-working-on-white-box-control Thu, 10 Jul 2025 00:00:00 +0000 LLM judges on trial: A new statistical framework to assess autograders https://www.aisi.gov.uk/blog/llm-judges-on-trial-a-new-statistical-framework-to-assess-autograders Our new framework can assess the reliability of LLM evaluators, while simultaneously answering a primary research question. https://www.aisi.gov.uk/blog/llm-judges-on-trial-a-new-statistical-framework-to-assess-autograders Wed, 09 Jul 2025 00:00:00 +0000 How will AI enable the crimes of the future? https://www.aisi.gov.uk/blog/how-will-ai-enable-the-crimes-of-the-future How we're working to track and mitigate against criminal misuse of AI. https://www.aisi.gov.uk/blog/how-will-ai-enable-the-crimes-of-the-future Thu, 03 Jul 2025 00:00:00 +0000 Inspect Cyber: A New Standard for Agentic Cyber Evaluations https://www.aisi.gov.uk/blog/inspect-cyber Inspect Cyber: A New Standard for Agentic Cyber Evaluations https://www.aisi.gov.uk/blog/inspect-cyber Thu, 26 Jun 2025 00:00:00 +0000 New updates to the AISI Challenge Fund https://www.aisi.gov.uk/blog/new-updates-to-the-aisi-challenge-fund New updates to the AISI Challenge Fund https://www.aisi.gov.uk/blog/new-updates-to-the-aisi-challenge-fund Thu, 05 Jun 2025 00:00:00 +0000 Making safeguard evaluations actionable https://www.aisi.gov.uk/blog/making-safeguard-evaluations-actionable An Example Safety Case for Safeguards Against Misuse https://www.aisi.gov.uk/blog/making-safeguard-evaluations-actionable Thu, 29 May 2025 00:00:00 +0000 HiBayES: Improving LLM evaluation with hierarchical Bayesian modelling https://www.aisi.gov.uk/blog/hibayes-improving-llm-evaluation-with-hierarchical-bayesian-modelling HiBayES: a flexible, robust statistical modelling framework that accounts for the nuances and hierarchical structure of advanced evaluations. https://www.aisi.gov.uk/blog/hibayes-improving-llm-evaluation-with-hierarchical-bayesian-modelling Mon, 12 May 2025 00:00:00 +0000 Research Agenda https://www.aisi.gov.uk/blog/research-agenda We outline our research priorities, our approach to developing technical solutions to the most pressing AI concerns, and the key risks that must be addressed as AI capabilities advance. https://www.aisi.gov.uk/blog/research-agenda Tue, 06 May 2025 00:00:00 +0000 RepliBench: measuring autonomous replication capabilities in AI systems https://www.aisi.gov.uk/blog/replibench-measuring-autonomous-replication-capabilities-in-ai-systems A comprehensive benchmark to detect emerging replication abilities in AI systems and provide a quantifiable understanding of potential risks https://www.aisi.gov.uk/blog/replibench-measuring-autonomous-replication-capabilities-in-ai-systems Tue, 22 Apr 2025 00:00:00 +0000 How to evaluate control measures for AI agents? https://www.aisi.gov.uk/blog/how-to-evaluate-control-measures-for-ai-agents Our new paper outlines how AI control methods can mitigate misalignment risks as capabilities of AI systems increase https://www.aisi.gov.uk/blog/how-to-evaluate-control-measures-for-ai-agents Fri, 11 Apr 2025 00:00:00 +0000 Strengthening AI resilience https://www.aisi.gov.uk/blog/strengthening-ai-resilience 20 Systemic Safety Grant Awardees Announced https://www.aisi.gov.uk/blog/strengthening-ai-resilience Thu, 03 Apr 2025 00:00:00 +0000 How we’re addressing the gap between AI capabilities and mitigations https://www.aisi.gov.uk/blog/aisis-research-direction-for-technical-solutions We outline our approach to technical solutions for misuse and loss of control. https://www.aisi.gov.uk/blog/aisis-research-direction-for-technical-solutions Tue, 11 Mar 2025 00:00:00 +0000 How can safety cases be used to help with frontier AI safety? https://www.aisi.gov.uk/blog/how-can-safety-cases-be-used-to-help-with-frontier-ai-safety Our new papers show how safety cases can help AI developers turn plans in their safety frameworks into action https://www.aisi.gov.uk/blog/how-can-safety-cases-be-used-to-help-with-frontier-ai-safety Mon, 10 Feb 2025 00:00:00 +0000 Principles for safeguard evaluation https://www.aisi.gov.uk/blog/principles-for-safeguard-evaluation Our new paper proposes core principles for evaluating misuse safeguards https://www.aisi.gov.uk/blog/principles-for-safeguard-evaluation Tue, 04 Feb 2025 00:00:00 +0000 Pre-Deployment evaluation of OpenAI’s o1 model https://www.aisi.gov.uk/blog/pre-deployment-evaluation-of-openais-o1-model The UK Artificial Intelligence Safety Institute and the U.S. Artificial Intelligence Safety Institute conducted a joint pre-deployment evaluation of OpenAI's o1 model https://www.aisi.gov.uk/blog/pre-deployment-evaluation-of-openais-o1-model Wed, 18 Dec 2024 00:00:00 +0000 Long-Form Tasks https://www.aisi.gov.uk/blog/long-form-tasks A Methodology for Evaluating Scientific Assistants https://www.aisi.gov.uk/blog/long-form-tasks Tue, 03 Dec 2024 00:00:00 +0000 Pre-deployment evaluation of Anthropic’s upgraded Claude 3.5 Sonnet https://www.aisi.gov.uk/blog/pre-deployment-evaluation-of-anthropics-upgraded-claude-3-5-sonnet The UK Artificial Intelligence Safety Institute and U.S. Artificial Intelligence Safety Institute conducted a joint pre-deployment evaluation of Anthropic’s latest model https://www.aisi.gov.uk/blog/pre-deployment-evaluation-of-anthropics-upgraded-claude-3-5-sonnet Tue, 19 Nov 2024 00:00:00 +0000 Safety case template for ‘inability’ arguments https://www.aisi.gov.uk/blog/safety-case-template-for-inability-arguments How to write part of a safety case showing a system does not have offensive cyber capabilities https://www.aisi.gov.uk/blog/safety-case-template-for-inability-arguments Thu, 14 Nov 2024 00:00:00 +0000 Announcing Inspect Evals https://www.aisi.gov.uk/blog/inspect-evals We’re open-sourcing dozens of LLM evaluations to advance safety research in the field https://www.aisi.gov.uk/blog/inspect-evals Wed, 13 Nov 2024 00:00:00 +0000 Our First Year https://www.aisi.gov.uk/blog/our-first-year The AI Safety Institute reflects on its first year https://www.aisi.gov.uk/blog/our-first-year Wed, 13 Nov 2024 00:00:00 +0000 Bounty programme for novel evaluations and agent scaffolding https://www.aisi.gov.uk/blog/evals-bounty We are launching a bounty for novel evaluations and agent scaffolds to help assess dangerous capabilities in frontier AI systems. https://www.aisi.gov.uk/blog/evals-bounty Tue, 05 Nov 2024 00:00:00 +0000 Early lessons from evaluating frontier AI systems https://www.aisi.gov.uk/blog/early-lessons-from-evaluating-frontier-ai-systems We look into the evolving role of third-party evaluators in assessing AI safety, and explore how to design robust, impactful testing frameworks. https://www.aisi.gov.uk/blog/early-lessons-from-evaluating-frontier-ai-systems Thu, 24 Oct 2024 00:00:00 +0000 Advancing the field of systemic AI safety: grants open https://www.aisi.gov.uk/blog/advancing-the-field-of-systemic-ai-safety-grants-open Calling researchers from academia, industry, and civil society to apply for up to £200,000 of funding. https://www.aisi.gov.uk/blog/advancing-the-field-of-systemic-ai-safety-grants-open Tue, 15 Oct 2024 00:00:00 +0000 Why I joined AISI by Geoffrey Irving https://www.aisi.gov.uk/blog/why-i-joined-aisi---geoffrey-irving Our Chief Scientist, Geoffrey Irving, on why he joined the UK AI Safety Institute and why he thinks other technical folk should too https://www.aisi.gov.uk/blog/why-i-joined-aisi---geoffrey-irving Thu, 03 Oct 2024 00:00:00 +0000 Should AI systems behave like people? https://www.aisi.gov.uk/blog/should-ai-systems-behave-like-people We studied whether people want AI to be more human-like. https://www.aisi.gov.uk/blog/should-ai-systems-behave-like-people Wed, 25 Sep 2024 00:00:00 +0000 Early Insights from Developing Question-Answer Evaluations for Frontier AI https://www.aisi.gov.uk/blog/early-insights-from-developing-question-answer-evaluations-for-frontier-ai A common technique for quickly assessing AI capabilities is prompting models to answer hundreds of questions, then automatically scoring the answers. We share insights from months of using this method. https://www.aisi.gov.uk/blog/early-insights-from-developing-question-answer-evaluations-for-frontier-ai Mon, 23 Sep 2024 00:00:00 +0000 Conference on frontier AI safety frameworks https://www.aisi.gov.uk/blog/conference-on-frontier-ai-safety-frameworks AISI is bringing together AI companies and researchers for an invite-only conference to accelerate the design and implementation of frontier AI safety frameworks. This post shares the call for submissions that we sent to conference attendees. https://www.aisi.gov.uk/blog/conference-on-frontier-ai-safety-frameworks Thu, 19 Sep 2024 00:00:00 +0000 Cross-post: "Interviewing AI researchers on automation of AI R&D" by Epoch AI https://www.aisi.gov.uk/blog/interviewing-researchers-on-automation AISI funded Epoch AI to explore AI researchers’ differing predictions on the automation of AI research and development and their suggestions for how to evaluate relevant capabilities. https://www.aisi.gov.uk/blog/interviewing-researchers-on-automation Tue, 27 Aug 2024 00:00:00 +0000 Safety cases at AISI https://www.aisi.gov.uk/blog/safety-cases-at-aisi As a complement to our empirical evaluations of frontier AI models, AISI is planning a series of collaborations and research projects sketching safety cases for more advanced models than exist today, focusing on risks from loss of control and autonomy. By a safety case, we mean a structured argument that an AI system is safe within a particular training or deployment context. https://www.aisi.gov.uk/blog/safety-cases-at-aisi Fri, 23 Aug 2024 00:00:00 +0000 Advanced AI evaluations at AISI: May update https://www.aisi.gov.uk/blog/advanced-ai-evaluations-may-update We tested leading AI models for cyber, chemical, biological, and agent capabilities and safeguards effectiveness. Our first technical blog post shares a snapshot of our methods and results. https://www.aisi.gov.uk/blog/advanced-ai-evaluations-may-update Mon, 20 May 2024 00:00:00 +0000 Fourth progress report https://www.aisi.gov.uk/blog/fourth-progress-report Since February, we released our first technical blog post, published the International Scientific Report on the Safety of Advanced AI, open-sourced our testing platform Inspect, announced our San Francisco office, announced a partnership with the Canadian AI Safety Institute, grew our technical team to >30 researchers and appointed Jade Leung as our Chief Technology Officer. https://www.aisi.gov.uk/blog/fourth-progress-report Mon, 20 May 2024 00:00:00 +0000 International Scientific Report on the Safety of Advanced AI: Interim Report https://www.aisi.gov.uk/blog/international-scientific-report-on-the-safety-of-advanced-ai-interim-report This is an up-to-date, evidence-based report on the science of advanced AI safety. It highlights findings about AI progress, risks, and areas of disagreement in the field. The report is chaired by Yoshua Bengio and coordinated by AISI. https://www.aisi.gov.uk/blog/international-scientific-report-on-the-safety-of-advanced-ai-interim-report Fri, 17 May 2024 00:00:00 +0000 Open sourcing our testing framework Inspect https://www.aisi.gov.uk/blog/open-sourcing-our-testing-framework-inspect We open-sourced our framework for large language model evaluation, which provides facilities for prompt engineering, tool usage, multi-turn dialogue, and model-graded evaluations. https://www.aisi.gov.uk/blog/open-sourcing-our-testing-framework-inspect Sun, 21 Apr 2024 00:00:00 +0000 Announcing the UK and US AISI partnership https://www.aisi.gov.uk/blog/announcing-the-uk-and-us-aisi-partnership The UK and US AI Safety Institutes signed a landmark agreement to jointly test advanced AI models, share research insights, share model access and enable expert talent transfers. https://www.aisi.gov.uk/blog/announcing-the-uk-and-us-aisi-partnership Tue, 02 Apr 2024 00:00:00 +0000 Announcing the UK and France AI Research Institutes’ collaboration https://www.aisi.gov.uk/blog/announcing-the-uk-and-france-ai-research-institutes-collaboration The UK AI Safety Institute and France’s Inria (The National Institute for Research in Digital Science and Technology) are partnering to advance AI safety research. https://www.aisi.gov.uk/blog/announcing-the-uk-and-france-ai-research-institutes-collaboration Thu, 29 Feb 2024 00:00:00 +0000 Our approach to evaluations https://www.aisi.gov.uk/blog/our-approach-to-evaluations This post offers an overview of why we are doing this work, what we are testing for, how we select models, our recent demonstrations and some plans for our future work. https://www.aisi.gov.uk/blog/our-approach-to-evaluations Fri, 09 Feb 2024 00:00:00 +0000 Third progress report https://www.aisi.gov.uk/blog/third-progress-report Since October, we have recruited leaders from DeepMind and Oxford, onboarded 23 new researchers, published the principles behind the International Scientific Report on Advanced AI Safety, and began pre-deployment testing of advanced AI systems. https://www.aisi.gov.uk/blog/third-progress-report Mon, 05 Feb 2024 00:00:00 +0000 First AI Safety Summit https://www.aisi.gov.uk/blog/ai-safety-summit-2023 At the first AI Safety Summit at Bletchley Park, world leaders and top companies agreed on the significance of advanced AI risks and the importance of testing. https://www.aisi.gov.uk/blog/ai-safety-summit-2023 Thu, 02 Nov 2023 00:00:00 +0000 Second progress report https://www.aisi.gov.uk/blog/second-progress-report Since September, we have recruited leaders from OpenAI and Humane Intelligence, tripled the capacity of our research team, announced 6 new research partnerships, and helped establish the UK’s fastest supercomputer. https://www.aisi.gov.uk/blog/second-progress-report Mon, 30 Oct 2023 00:00:00 +0000 First Progress Report https://www.aisi.gov.uk/blog/first-progress-report In our first 11 weeks, we have recruited an advisory board of national security and ML leaders, including Yoshua Bengio, recruited top professors from Cambridge and Oxford and announced 4 research partnerships. https://www.aisi.gov.uk/blog/first-progress-report Thu, 07 Sep 2023 00:00:00 +0000