Search NASA⌕ Search

Engineering topics

Casalnuovo, Casey

Publications and source records attributed to Casalnuovo, Casey.

A Framework for Evaluating the Implementation Cost of Attacks on Large Language Models

Large Language Models (LLMs) have been increasingly proposed as a method to enhance productivity in tasks that involve language and code. However, these models are large, complex, and their capabilities are not easily understood and controlled, meaning that their adoption opens many possibilities for new cyberattacks and misuse. Numerous attacks on LLMs have been reported and summarized in literature reviews, but we found existing reviews lacking in understanding the implementation cost of the attacks - i.e., how much effort would an attacker need in terms of coding, expertise, and resources to adopt attacks presented in the literature. Therefore, we divide existing attacks on LLMs into a taxonomy, and define a cost evaluation framework to determine the cost of the attack. An attack’s cost can be 1) estimated from reading the publication about the attack or 2) determined by implementing the attack from that publication. We provide an example evaluation of a couple jailbreaking frameworks based on experiments, and then apply the more lightweight cost estimate to a representative selection of attacks across the taxonomy we define. We discuss the relative difficulty of the attacks and also highlight defenses that have attempted to mitigate these attacks and assess their effectiveness.

97 MATHEMATICS AND COMPUTING↗

Evaluation of the Self Retrieval Augmented Generation Technique on Common Security Advisory Framework Data

This small experimental report evaluates a variation of Retrieval Augmented Generation (RAG), called Self-RAG. This method uses a generative language model that incorporates retrieved facts into its generation and is explicitly trained to be able to determine whether retrieved information is enough to answer the input query, with a user-defined threshold for confidence. We performed an experiment using data from the publicly available CISA Common Security Advisory Framework (CSAF) repository (https://github.com/cisagov/CSAF) as the database of facts to be used in retrieval. Qualitative results from the experiment demonstrate that the Self-RAG method has some ability to provide reasonable answers to queries that are in the dataset and will often ignore irrelevant information when asked outside of domain questions (e.g., general facts). In settings with deliberately confusing questions (the question is within domain, but asks about a fabricated advisory), it was able to refuse 40% of the time without further adjustments to the original framework. While this performance is not sufficient for current practical use, further improvements to data formatting, disambiguating results, and leveraging threshold values could improve performance significantly. However, evaluating this will require more extensive evaluations on larger datasets and potentially better models.

97 MATHEMATICS AND COMPUTING↗