DOE OSTI · 2585484
Mini Report: Jailbreaking Attacks and Defenses
Abstract
Overall, jailbreaking defenses are unreliable, and no defenses proposed thus far can completely stop such attacks in any verifiable way. Even manual attacks can trivially bypass some claimed ‘defenses’, and over the course of 2023 automated attacks on prior models have been adapted to work on LLMs while other attacks draw on ideas such as fuzz testing. At best, some defenses can make jailbreaks more difficult, but with the developing landscape of attacks existing papers have not robustly evaluated how effective they are against all these methods. However, our opinion is that given the way these large generative models are trained and ‘aligned’ to stated goals of safety via fine tuning, it will be exceedingly difficult if not impossible to eliminate the possibility of jailbreaking attacks. A major barrier is that the feature space of LLMs is not sufficiently understood in a way where guarantees can be made about the outputs. Barring major changes, the expectation around jailbreaking defenses should be that they can mitigate misuse, but not verifiably prevent it. However, one defense we believe merits further investigation depends on the fact as automated attacks produce text, they need an automated way to identify a successful jailbreak - a judgment model. We have seen some attempts to repurpose models like these to defend against jailbreaks, but the evaluations are small scale and not robust.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Casalnuovo, Casey, Heidbrink, Scott Jared, Kavaler, David Minh, Lee, Jina. 2024-02-01. Mini Report: Jailbreaking Attacks and Defenses. https://doi.org/10.2172/2585484
Cite the original work for its findings. Save a collection to share your selection of sources.