Search NASASearch

Engineering topics

Berman, Brandon

Publications and source records attributed to Berman, Brandon.

Predictive Indicators of the Performance of Large Language Models

In several mission contexts, it is desirable to estimate the performance of large language models (LLMs) on tasks that we cannot run directly. In light of published “scaling laws” our hypothesis is that some tasks should be consistently more challenging than others based on characteristics of the task. The goal of this project was to begin quantifying how much information about LLM performance can be gained from the features of a model and a task. Two of our statistical models struggled to converge. Pass/fail test results may provide limited information for inference beyond model quality and task difficulty, but we see no evidence at this time for significant feature interaction effect sizes, arguing for simple models. Future work extending the models to capitalize on perplexity of ground truth answers is suggested. This project also introduces “Depth of Knowledge Variant Testing” as a strategy for more finely assessing language models on open domain question and answer tasks. We developed sets of questions that ask a language model to produce similar information while demonstrating increasing depth of knowledge, and also relabeled existing Q&A test questions with their depth of knowledge. Our results suggest further consideration of Bloom’s taxonomy and further refinement of prompts to properly elicit information at varying depths. In the course of this work, we set up a basic infrastructure for standardizing tasks and testing many language models on these tasks. In addition to testing the predictive quality of model features and performance across test suites, with this project we have introduced two new task features to contextualize each test question: the Dewey Classification main category of information covered, and the Bloom’s taxonomy level that corresponds to the depth of knowledge probed by the question. Splits across these and other features produced over five hundred task subtypes with distinct feature vectors, which we tested on half a dozen models.

97 MATHEMATICS AND COMPUTING

Non-conformity Scores for High-Quality Uncertainty Quantification from Conformal Prediction

High-quality uncertainty quantification (UQ) is a critical component of enabling trust in deep learning (DL) models and is especially important if DL models are to be deployed in high-consequence applications. Conformal prediction (CP) methods represent an emerging nonparametric approach for producing UQ that is easily interpretable and, under weak assumptions, provides a guarantee regarding UQ quality. This report describes the research outputs of an Exploratory Express Laboratory Directed Research and Development (LDRD) project at Sandia National Laboratories. This project focused on how best to implement CP methods for DL models. This report introduces new methodology for obtaining high-quality UQ from DL models using CP methods, describes a novel system of assessing UQ quality, and provides experimental results that demonstrate the quality of the new methodology and utility of the UQ quality assessment system. Avenues for future research and discussion of potential impacts at Sandia and in the wider research community are also given.

97 MATHEMATICS AND COMPUTING