DS1 spectrogram: Evaluating Frontier Models for Dangerous Capabilities

Evaluating Frontier Models for Dangerous Capabilities

2403.13793

Authors

Seliem El-Sayed,Sasha Brown,Rohin Shah,Allan Dafoe,Toby Shevlane

Abstract

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 models.

Our evaluations cover four areas: (1) persuasion and deception; (2) cyber-security; (3) self-proliferation; and (4) self-reasoning. We do not find evidence of strong dangerous capabilities in the models we evaluated, but we flag early warning signs.

Our goal is to help advance a rigorous science of dangerous capability evaluation, in preparation for future models.

Resources

Stay in the loop

Every AI paper that matters, free in your inbox daily.

Details

  • takara.ai
  • Custom AI and machine learning from the Frontier Research Team.
  • © 2026 takara.ai Ltd
  • Content is sourced from third-party publications.