A Framework for Evaluating the Implementation Cost of Attacks on Large Language Models
Large Language Models (LLMs) have been increasingly proposed as a method to enhance productivity in tasks that involve language and code. However, these models are large, complex, and their capabilities are not easily understood and controlled, meaning that their adoption opens many possibilities for new cyberattacks and misuse. Numerous attacks on LLMs have been reported and summarized in literature reviews, but we found existing reviews lacking in understanding the implementation cost of the attacks - i.e., how much effort would an attacker need in terms of coding, expertise, and resources to adopt attacks presented in the literature. Therefore, we divide existing attacks on LLMs into a taxonomy, and define a cost evaluation framework to determine the cost of the attack. An attack’s cost can be 1) estimated from reading the publication about the attack or 2) determined by implementing the attack from that publication. We provide an example evaluation of a couple jailbreaking frameworks based on experiments, and then apply the more lightweight cost estimate to a representative selection of attacks across the taxonomy we define. We discuss the relative difficulty of the attacks and also highlight defenses that have attempted to mitigate these attacks and assess their effectiveness.