Search NASASearch

DOE OSTI · 2585973

Arbitrary Autoencoder Injection for Interpretability Experimentation

Abstract

Goes over a simple software library (Python) for utilizing sparse autoencoders for more models than just language models.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Campos, Marco Vinicio, Cauthen, Katherine Regina, Krofcheck, Daniel Joseph, Naugle, Asmeret Bier, Simpson, Sarah Elizabeth, Doyle, Casey Lane, Sweitzer, Matthew Donald, Xi, Michael. 2024-09-01. Arbitrary Autoencoder Injection for Interpretability Experimentation. https://doi.org/10.2172/2585973

Cite the original work for its findings. Save a collection to share your selection of sources.