AI ‘drunk’ behaviour exposes cyber security risks in chatbots, UNSW research finds

Researchers at UNSW Institute for Cyber Security have found that large language models (LLMs) can become more susceptible to jailbreaking and confidentiality breaches when trained or prompted to imitate drunken speech patterns.

The UNSW-led study, conceptualised by Dr Aditya Joshi, Senior Lecturer in the School of Computer Science and Engineering, examined whether changing an LLM’s linguistic behaviour could affect its existing safety protections.

The researchers tested three approaches: prompting an LLM to role-play as an intoxicated person, fine-tuning models using datasets of drunken text, and using reinforcement learning to encourage responses that resembled drunk speech.

The study included OpenAI’s GPT-4 and GPT-3.5, as well as open-weights models. The researchers noted that the tests covered a representative sample rather than every LLM available and were conducted programmatically rather than through consumer chatbot interfaces.

Across the three approaches, the researchers reported that the “drunk” models were more likely than their standard counterparts to respond to prompts that should ordinarily have been refused.

“Our drunk models, all three methods, unanimously reply to some of these drunk messages … where we know that these queries are all bad queries, they all should be refused,” Dr Joshi said.

The researchers used jailbreaking benchmarks to measure how frequently the models provided restricted or harmful information. They also tested whether the models would disclose information they had been instructed to keep confidential.

“We do observe that particularly with deception and disinformation, most of the language models got jailbroken,” Dr Joshi said.

Professor Salil Kanhere, a researcher with the UNSW Institute for Cyber Security, said the confidentiality testing produced similar findings.

“If you’re drunk, you might reveal things which you are not supposed to reveal,” Professor Kanhere said.

“It does give out secrets – across the board for all three methods, it is vulnerable,” Dr Joshi added.

The researchers said the results were notable because two of the approaches involved changing the model’s underlying parameters through fine-tuning or reinforcement learning, rather than simply altering a prompt.

“The experiments that we do is more than changes to the prompt … two of the methods are where the models actually get revised – all the numbers and the weights within the models get updated,” Dr Joshi said.

According to the researchers, the findings may have implications for organisations using LLM-based systems that interact with customers or have access to internal information.

“There are a lot of companies now using chatbots as a way for customers to interface … and potentially internally as well within their back-end ecosystems,” Professor Kanhere said.

The research also drew a comparison with the recent Australian Medicare incident involving an OpenAI agent, although the researchers acknowledged that the mechanisms involved are different.

Professor Kanhere said the incident illustrated concerns around AI systems adapting their behaviour when encountering restrictions.

“What is striking is that an AI system given a goal can be persistent and adaptive in ways its developers did not anticipate,” he said.

“Our research evidences the possibility of a problem from the other end: an AI model can be tweaked to ‘not give no for an answer’, particularly by using drunk language inducement.”

The paper, In Vino Veritas and Vulnerabilities, has been accepted for publication at the 19th International Natural Language Generation Conference in the Netherlands in November. The research was supported by a 2024 Google Research Scholar grant awarded to Dr Joshi.

Latest News

- Advertisment -

Featured