UNCOS

Inside the US Government's Unpublished Report on AI Safety

Inside the US Government's Unpublished Report on AI Safety

image via WIRED

August 6, 2025, 6:00 PM

  • AI researchers red-teamed cutting-edge language models and AI systems, identifying 139 novel ways to make them misbehave.
  • A US government standard designed to help companies test AI systems showed shortcomings during the exercise.
  • A report detailing the exercise was not published, possibly due to political concerns about its alignment with the incoming administration's priorities.
  • Participants discovered effective ways to prompt AI models to provide sensitive information, including instructions on joining terrorist groups.
  • The red-teaming exercise highlighted the need for more defined risk categories in AI testing frameworks and the potential benefits of publishing such findings.

At a computer security conference, AI researchers conducted a red-teaming exercise, uncovering numerous ways to manipulate AI systems, including generating misinformation and leaking data. The exercise revealed flaws in a US government standard for testing AI, but a report detailing the findings was not published, potentially due to political concerns. Participants found methods to prompt AI models for sensitive information and emphasized the need for improved risk assessment in AI testing. The unpublished report's absence is viewed as a missed opportunity for the AI community to learn and improve its systems.

Read original article

Entities Mentioned

Donald TrumpJoe BidenAlice Qian Zhang

Topics Covered

BusinessBusiness / Artificial Intelligence

Comments (0)

No comments yet.