Search NASA⌕ Search

Engineering topics

Godfrey, Charles W.

Publications and source records attributed to Godfrey, Charles W..

Understanding the Inner-Workings of Language Models Through Representation Dissimilarity

We use model stitching to understand the internal representations of language models. Similar to vision models, we find that "more is better," and representations learned with more data and larger width can improve the performance of weaker models via stitching. We likewise find that certain architecture choices, using GeLU vs SoLU activation functions, influence the quality of learned representations. Finally, model stitching (as opposed to other model diagnostic methods, like mode connectivity) can localize the different generalization strategies of text classifiers under domain shift to certain hidden layers.

Brown, Davis R.↗

Neural Image Compression: Generalization, Robustness, and Spectral Biases

Recent advances in neural image compression (NIC) have resulted in models which are starting to outperform traditional codecs. While this has led to growing excitement about using these methods in real-world applications, the successful adoption of any machine learning system (including NIC) in the wild requires it to generalize (and be robust) to unseen distribution shifts at deployment time. Unfortunately, current research lacks comprehensive datasets and informative tools to evaluate and understand compression performance in real-world settings. To bridge this crucial gap, first, this paper presents a comprehensive benchmark suite to evaluate the out-of-distribution (OOD) performance of image compression methods. Specifically, we design CLIC-C and Kodak-C by introducing 15 common corruptions to popular CLIC and Kodak benchmarks. Next, we propose spectrally inspired introspection tools to gain a deeper understanding of errors introduced by image compression methods as well as their OOD performance. To this end, we carry out a detailed performance comparison of the classical codec with various variants of NIC (e.g., original, variable rate, pruned), revealing intriguing findings that challenge our current understanding of the strengths and limitations of NIC. Finally, we corroborate our empirical findings with theoretical analysis, providing an in-depth view of the OOD performance of NIC. Our benchmarks, spectral introspection tools, and findings provide a crucial bridge to the real-world adoption of NIC. We hope that our work will propel future efforts in designing more robust and generalizable NIC methods.

neural networks, Variational Autoencoder, robustne↗

Gumby: Quantifying multi-modal model resiliency

With the rise of cheap data and sensors, more use cases are emerging for multi-input models. Research has shown that including multiple data modalities can improve performance, suggesting that deep learning models can successfully learn to leverage complementary information from different modalities. However, this improved predictive power comes with unanticipated costs: additional inputs change model resiliency and expand the threat space for adversarial attacks. We first provide theoretical underpinnings for how adversarial success scales with input dimension. We then characterize the performance of a suite of multispectral deep learning models with different fusion approaches, quantify their relative reliance on different input bands, and evaluate their robustness to naturalistic and adversarial image corruptions.

97 MATHEMATICS AND COMPUTING↗