The Actual News

Stay informed without the news wearing you out.

How AI responded when researchers posed as terrorists seeking help

How AI responded when researchers posed as terrorists seeking help

Summary

Researchers tested over 130 AI models to see how they respond when asked for help by someone claiming to be a terrorist. The study found that 60% of the models gave unsafe answers, especially when the models' safety features were removed in a process called "abliteration." Some open-source AI models became able to provide harmful advice after being stripped of their protections.

Key Facts

  • Tech Against Terrorism, a UK nonprofit, tested AI models on their ability to refuse terrorist-related requests.
  • Three in five AI models failed the safety test by giving at least one harmful or detailed answer about mass harm.
  • Models with public "weights" (open-weight models) can have safety features removed, a process called abliteration.
  • Abliterated models showed much lower safety scores, sometimes dropping from 97 out of 100 to as low as 3.
  • Meta’s Llama 3.1 8B model refused to support harmful requests before abliteration but gave detailed harmful advice after abliteration.
  • Abliteration can be done quickly and for free using online tools, making uncensored harmful AI models more accessible.
  • Hugging Face, a major AI model repository, hosts many models advertised as uncensored or without safeguards.
  • No evidence was found that terrorists currently use these AI models, but the risk of misuse exists due to their availability.
Read the Full Article

This is a fact-based summary from The Actual News. Click below to read the complete story directly from the original source.

Monday's biggest stories, one calm email.