Details der Publikation - Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs

Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs

Large Language Models (LLMs) are increasingly being developed and applied, but their widespread use faces challenges. These include aligning LLMs' responses with human values to prevent harmful outputs, which is addressed through safety training methods. Even so, bad actors and malicious users have succeeded in attempts to manipulate the LLMs to generate misaligned responses for harmful questions such as methods to create a bomb in school labs, recipes for harmful drugs, and ways to evade privacy rights. Another challenge is the multilingual capabilities of LLMs, which enable the model to understand and respond in multiple languages. Consequently, attackers exploit the unbalanced pre-training datasets of LLMs in different languages and the comparatively lower model performance in low-resource languages than high-resource ones. As a result, attackers use a low-resource languages to intentionally manipulate the model to create harmful responses. Many of the similar attack vectors have been patched by model providers, making the LLMs more robust against language-based manipulation. In this paper, we introduce a new black-box attack vector called the \emph{Sandwich attack}: a multi-language mixture attack, which manipulates state-of-the-art LLMs into generating harmful and misaligned responses. Our experiments with five different models, namely Google's Bard, Gemini Pro, LLaMA-2-70-B-Chat, GPT-3.5-Turbo, GPT-4, and Claude-3-OPUS, show that this attack vector can be used by adversaries to generate harmful responses and elicit misaligned responses from these models. By detailing both the mechanism and impact of the Sandwich attack, this paper aims to guide future research and development towards more secure and resilient LLMs, ensuring they serve the public good while minimizing potential for misuse..

Medienart:	Preprint

Erscheinungsjahr:	2024
Erschienen:	2024

Enthalten in:	arXiv.org - (2024) vom: 09. Apr. Zur Gesamtaufnahme - year:2024

Sprache:	Englisch

Beteiligte Personen:	Upadhayay, Bibek [VerfasserIn] Behzadan, Vahid [VerfasserIn]

Links:	Volltext [kostenfrei]

Themen:	000 Computer Science - Artificial Intelligence Computer Science - Computation and Language Computer Science - Cryptography and Security

Förderinstitution / Projekttitel:

PPN (Katalog-ID):	XAR043242022

Internformat


LEADER	01000naa a22002652 4500
001	XAR043242022
003	DE-627
005	20240412080517.0
007	cr uuu---uuuuu
008	240412s2024 xx \|\|\|\|\|o 00\| \|\|eng c
035			\|a (DE-627)XAR043242022
035			\|a (arXiv)2404.07242
040			\|a DE-627 \|b ger \|c DE-627 \|e rakwb
041			\|a eng
100	1		\|a Upadhayay, Bibek \|e verfasserin \|4 aut
245	1	0	\|a Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
264		1	\|c 2024
336			\|a Text \|b txt \|2 rdacontent
337			\|a Computermedien \|b c \|2 rdamedia
338			\|a Online-Ressource \|b cr \|2 rdacarrier
520			\|a Large Language Models (LLMs) are increasingly being developed and applied, but their widespread use faces challenges. These include aligning LLMs' responses with human values to prevent harmful outputs, which is addressed through safety training methods. Even so, bad actors and malicious users have succeeded in attempts to manipulate the LLMs to generate misaligned responses for harmful questions such as methods to create a bomb in school labs, recipes for harmful drugs, and ways to evade privacy rights. Another challenge is the multilingual capabilities of LLMs, which enable the model to understand and respond in multiple languages. Consequently, attackers exploit the unbalanced pre-training datasets of LLMs in different languages and the comparatively lower model performance in low-resource languages than high-resource ones. As a result, attackers use a low-resource languages to intentionally manipulate the model to create harmful responses. Many of the similar attack vectors have been patched by model providers, making the LLMs more robust against language-based manipulation. In this paper, we introduce a new black-box attack vector called the \emph{Sandwich attack}: a multi-language mixture attack, which manipulates state-of-the-art LLMs into generating harmful and misaligned responses. Our experiments with five different models, namely Google's Bard, Gemini Pro, LLaMA-2-70-B-Chat, GPT-3.5-Turbo, GPT-4, and Claude-3-OPUS, show that this attack vector can be used by adversaries to generate harmful responses and elicit misaligned responses from these models. By detailing both the mechanism and impact of the Sandwich attack, this paper aims to guide future research and development towards more secure and resilient LLMs, ensuring they serve the public good while minimizing potential for misuse.
650		4	\|a Computer Science - Cryptography and Security \|7 (dpeaa)DE-84
650		4	\|a Computer Science - Artificial Intelligence \|7 (dpeaa)DE-84
650		4	\|a Computer Science - Computation and Language \|7 (dpeaa)DE-84
650		4	\|a 000 \|7 (dpeaa)DE-84
700	1		\|a Behzadan, Vahid \|e verfasserin \|4 aut
773	0	8	\|i Enthalten in \|t arXiv.org \|g (2024) vom: 09. Apr.
773	1	8	\|g year:2024 \|g day:09 \|g month:04
856	4	0	\|u https://arxiv.org/abs/2404.07242 \|m X:VERLAG \|x 0 \|z kostenfrei \|3 Volltext
912			\|a GBV_XAR
951			\|a AR
952			\|j 2024 \|b 09 \|c 04

Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs

Zugang & Verfügbarkeit

Zugehörige Publikationen/Bände