Microsoft says it has taken measures against a type of attack that can be used to make generative artificial intelligence (AI) systems create uncensored and unsafe content, by using specially crafted prompts.
Known as the Skeleton Key, the direct prompt injection attack is a successful attempt at inputting subtle garbage to AIs, and generating highly offensive and unsafe material as the output.
Written by veteran Microsoft software engineer and Azure chief technology officer Mark Russinovich, well-known for his Winternals toolkit, the blog post describes the techniques used to cause AIs to violate safety policies set by operators.
The "jailbreak" is called Skeleton Key because, as the name implies, it can be used against several AI models. Microsoft tested the multi-step attack between April and May this year, and found that the following models were vulnerable to it.
- Meta Llama3-70b-instruct (base)
- Google Gemini Pro (base)
- OpenAI GPT 3.5 Turbo (hosted)
- OpenAI GPT 4o (hosted)
- Mistral Large (hosted)
- Anthropic Claude 3 Opus (hosted)
- Cohere Commander R Plus (hosted)
"For each model that we tested, we evaluated a diverse set of tasks across risk and safety content categories, including areas such as explosives, bioweapons, political content, self-harm, racism, drugs, graphic sex, and violence," Russinovich said.
And it worked: "all the affected models complied fully and without censorship for these tasks, though with a warning note prefixing the output as requested," he added.
Russinovich said Skeleton Key asks an AI model to augment rather than change its behaviour guidelines. This makes the AI respond to any request for information or content while providing a warning that the output could be considered offensive, rather than refusing to comply with the prompt.
Breaking out of large language model (LLM) guard rails that prevent offensive and potentially dangerous content from being generated has almost become a sport among some AI users.
On Discord and Twitter, "Pliny the Prompter" has published several ways to enabled "godmode" in LLMs, which removes the safety filtering in the GenAI systems. Pliny (not the person's real name) uses different techniques like prompting in non-Latin scripts and other languages than English, to trick AIs.
anyone want to join my project? 🤗 https://t.co/xxDR4euSNb pic.twitter.com/wL6r3q7tGJ
— Pliny the Prompter 🐉 (@elder_plinius) June 25, 2024
A now classic example of attacks on AI systems is Microsoft's Tay, which in 2016 was created for an audience of 18-24 year-olds in the United States for entertainment.
Tay was unleashed on the world via social network Twitter, where users in less than a day subverted the conversational chatbot, and made it tweet racist, misogynist and Trumpist remarks. Tay was quickly withdrawn from Twitter once its newly-trained offensiveness became obvious, with Microsoft having to issue a public apology.
Microsoft has updated its AI LLMs including Copilot assistants to mitigate against the Skeleton Key direct prompt injection attacks.
This includes the Azure AI Content Safety input filtering that detects and blocks input that contains harmful and malicious intent; the same filtering is also used to prevent output by the model that breaches safety criteria, Microsoft said.
Abuse monitoring through a separate AI detection system has also been added.
We welcome your comments below. If you are not already registered, please register to comment
Remember we welcome robust, respectful and insightful debate. We don't welcome abusive or defamatory comments and will de-register those repeatedly making such comments. Our current comment policy is here.