Search NASA⌕ Search

SEARCH · Search NASA

Results for “data lake”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Antarctic lake viromes reveal potential virus associated influences on nutrient cycling in ice-covered lakes

The McMurdo Dry Valleys (MDVs) of Antarctica are a mosaic of extreme habitats which are dominated by microbial life. The MDVs include glacial melt holes, streams, lakes, and soils, which are interconnected through the transfer of energy and flux of inorganic and organic material via wind and hydrology. For the first time, we provide new data on the viral community structure and function in the MDVs through metagenomics of the planktonic and benthic mat communities of Lakes Bonney and Fryxell. Viral taxonomic diversity was compared across lakes and ecological function was investigated by characterizing auxiliary metabolic genes (AMGs) and predicting viral hosts. Our data suggest that viral communities differed between the lakes and among sites: these differences were connected to microbial host communities. AMGs were associated with the potential augmentation of multiple biogeochemical processes in host, most notably with phosphorus acquisition, organic nitrogen acquisition, sulfur oxidation, and photosynthesis. Viral genome abundances containing AMGs differed between the lakes and microbial mats, indicating site specialization. Using procrustes analysis, we also identified significant coupling between viral and bacterial communities (p = 0.001). Finally, host predictions indicate viral host preference among the assembled viromes. Collectively, our data show that: (i) viruses are uniquely distributed through the McMurdo Dry Valley lakes, (ii) their AMGs can contribute to overcoming host nutrient limitation and, (iii) viral and bacterial MDV communities are tightly coupled.

Microbiology↗

WTK-LED: The WIND Toolkit Long-Term Ensemble Dataset

To satisfy a wide group of stakeholders across various wind energy disciplines, including but not limited to stakeholders in the distributed and utility scale wind industry, the new emerging airborne wind energy field, grid integration, power systems modeling, environmental modeling, and researchers in academia, and to close some of the gaps that current public datasets have, we aimed at developing an updated version of the meteorological WIND Toolkit, named WIND Toolkit Long-term Ensemble Dataset (WTK-LED), which is a meteorological dataset providing time series every 5 min and 2 km, including model uncertainty of wind speed at every modeling grid point so that users are provided with a range of possible wind speeds every 2 km. The data were produced using the Weather Research and Forecasting Model (WRF). The vertical grid used in WTK-LED includes many vertical layers in the atmospheric boundary layer to provide information of atmospheric quantities across the rotor layer of utility scale and distributed wind turbines. The WTK-LED includes: 1) Numerical simulations covering the continental United States, Alaska, and Hawaii, with high-resolution data being available for 3 years (2018-2020). 2) Climate simulations from Argonne National Laboratories covering the North American continent, including Alaska, Canada, and most of Mexico and the Caribbean Islands. These simulations complement the new WTK-LED to offer a 4-km dataset covering 20 years, from 2001-2020. 3) Specific long-term,high-resolution offshore simulations have been conducted separately for the US coasts, Hawaii, and the Great Lakes, leading to the 2023 National Offshore Wind data set. This report focuses on a description of the land-based WTK-LED for CONUS, Hawaii, and Alaska, for the 3-year 2-km/5-min dataset and the 20-year 4-km/hourly dataset, as well as the uncertainty quantification method. We also provide limited validation results. Based on our results to date, we suggest use cases and applications for each dataset of the WTK-LED.

17 WIND ENERGY↗

A Performant, Scalable Processing Pipeline for High‐Quality and FAIR Environmental Sensor Data

High-resolution environmental monitoring is necessary to record, understand, and predict biogeochemical and ecological changes particularly in coastal systems but brings significant challenges in processing and making rapidly available the resulting data. The COMPASS-FME project established a network of coastal observational sites across the Chesapeake Bay and western Lake Erie regions extensively instrumented with soil, vegetation, and weather sensors logging data every 15 min. Our data processing framework, written in R and completely open source, prioritizes rapid model-experiment iteration and makes biogeochemical data rapidly available for quality assurance/quality control, analysis, and model ingestion. This pipeline is distinguished by a standardized and modular approach to data curation, extensive metadata and documentation, and its high performance. These attributes combine to make biogeochemical data rapidly accessible across COMPASS-FME and the broader community. Flexible, powerful, and reproducible approaches to handling high-volume environmental data are crucial for accelerating biogeosciences research.

Pennington, Stephanie C. [Pacific Northwest Nation↗

1993 Salt Lake City Household Travel Diary Survey

The purpose of the Salt Lake City Household Travel Diary Survey was to gather detailed information about current travel habits in Utah, to serve as the basis for future travel modeling activities, and to inform regional and statewide transportation planning. The survey sampled residents from Weber, Davis, Salt Lake, and Utah counties only. Along with travel data, the survey collected demographic and socioeconomic characteristics for 3,082 households and 8,333 participants.

1Hz data↗

Reservoir Storage Capacity Change (ResCap)

Overview Storage capacity is an essential reservoir metric that is directly linked to various water management and energy objectives. Accurate reporting and tracking of change in storage over time is crucial for the safe and reliable operation of the associated dam. While storage information is available for many reservoirs through the National Inventory of Dams, additional details, e.g. water elevation levels as well as changes over time are not included. This dataset contains reservoir storage capacities based on conducted surveys in CONUS. To represent changes in a reservoir’s storage over time, the storage capacity as determined by the first and last conducted survey is listed. The level of detail of surveys can vary greatly and improved with technological advancements. Therefore, the type of survey and year when it was conducted is noted. To ensure a fair comparison of storage capacities, the water elevation level along with the corresponding operation of the dam is reported. Structural changes, e.g. heightening of a dam will have an influence on the storage capacity and are therefore also mentioned. A total of 739 different reservoir storage capacity comparisons are listed, with some reservoirs represented more than once (storage capacity comparison at different water elevation levels). Methodology Data were acquired from USBR reservoir survey reports, TWDB lake survey reports and elevation-area-capacity tables, the RSI Web Portal and the NID (USACE, 2024). Initial storage capacity along with year and type of survey record is compared to the most recent reported storage capacity, survey type and year. Comparison elevation in feet as well as comparison elevation type were either extracted from survey reports (USBR, TWDB) or the Web Portal (RSI) and in some cases cross-referenced with data from other sources (Water Management Data, USACE, Water Data for Texas, TWDB).

Chu, Antonia [ORNL] (ORCID:0009000510540427)↗

Detecting and Characterizing Fracture Zones Using a Convolutional Neural Network

This project directly supports the Geothermal Technologies Office (GTO) objectives outlined in the Multi-Year Program Plan (MYPP) by advancing two key research areas: “Exploration and Characterization” and “Data, Modeling, and Analysis.” This project has successfully demonstrated a pre-drilling ability to image and characterize the distribution and connectivity of subsurface faults and fractures, key parameters for identifying permeable pathways that enable geothermal fluids to circulate and produce energy. Specifically, we developed and implemented innovative machine learning methodologies to enhance geothermal exploration. Large-scale faults were detected using a Convolutional Neural Network (CNN), while small-scale fractures were characterized using a novel Double-Beam Neural Network (DBNN). These tools have proven both technically effective and cost-efficient by reducing reliance on expensive exploratory drilling. Through collaboration with our geothermal industry partner, this research has significantly advanced techniques for identifying hidden geothermal systems and extending the productive lifespan of existing geothermal fields. We applied our methods to two geothermal fields—Soda Lake (Nevada) and Lightning Dock (New Mexico)—to identify shallow steam-charged fracture zones and characterize deep faults at depths of 1.5-2 km. The steam zone identified at the Soda Lake geothermal field showed excellent agreement with prior drilling data, validating the effectiveness of our approaches. In addition, the analysis revealed three new prospective drilling targets for further development and verification. The outcomes of this project improve our scientific understanding of geothermal reservoir behavior, enhance exploration efficiency, extend the economic life of existing geothermal plants. Ultimately, these advancements contribute to GTO’s goal of achieving more sustainable, affordable, and data-driven geothermal energy development across the United States.

15 GEOTHERMAL ENERGY↗

Data for Stetten et al. (2025), "Biogeochemical controls on iron speciation and cycling across upland to shoreline gradients in freshwater and estuarine coastal soils (Lake Erie and Chesapeake Bay, United States)"

Coastal environments are dynamic interfaces that mediate carbon and nutrient exchanges between terrestrial landscapes and open waters, but it is unclear how biogeochemical reactions, in particular iron (Fe) redox transformations, affect the understanding and prediction of coastal ecosystem functions. This dataset includes measurements from two freshwater sites in the Western and Central basins of Lake Erie (Ohio, United States) and two estuarine sites in the Chesapeake Bay (Maryland, United States); the analytical results were reported by Stetten et al. (2025) in Science of the Total Environment. It was produced as part of the COMPASS-FME project, which seeks to advance a scalable, predictive understanding of the fundamental biogeochemical processes, ecological structure, and ecosystem dynamics that distinguish coastal terrestrial-aquatic interfaces from the purely terrestrial or aquatic systems to which they are coupled. The sites were sampled in November 2022 (CRC), December 2022 (MSM), February 2023 (GCW), and March 2023 (OWC); site codes follow those used by Pennington et al. (2025).The dataset consists of the following soil data:- Solid data (Fe concentration, etc.)- Porewater data (sulfate, sulfide, etc.)- Linear combination fitting results of X-ray absorption near edge structure (XANES) spectra; i.e., quantitative results of the oxidation state of Fe, indicated as a proportion of pure Fe(III) and Fe(II) model compounds- Linear combination fitting results of EXAFS (extended X-ray absorption fine structure) spectra, indicated as proportion of of Fe-model compounds (illite, smectite, etc.)Each data type has a single file in comma-separated value (CSV) format. No special software is required to read it.

54 ENVIRONMENTAL SCIENCES↗

AmeriFlux CA-GL3 Long Point

This is the AmeriFlux version of the carbon flux data for the site CA-GL3 Long Point. Site Description - Long Point Lighthouse is located at the end of Long Point on Lake Erie. The eddy covariance instrumentation is located on the historic lighthouse, completed in 1916, and instrumented with eddy covariance data in 2012 by a network of scientists from both US and Canada (eventually to be called the Great Lakes Evaporation Network (GLEN)). The intent of GLEN has been to provide observations of over-lake meteorology and evaporation, improve forecasting of Great Lakes water levels, and support a wide variety of stakeholders, including the National Weather Service (NWS), Environment and Climate Change Canada, National Oceanic and Atmospheric Administration, U.S. Coast Guard, recreational boaters and commercial shipping, emergency management officials, and the Great Lakes research community.

Spence, Chris [Environment and Climate Change Cana↗

AmeriFlux CA-GL4 Nine Mile Lighthouse

This is the AmeriFlux version of the carbon flux data for the site CA-GL4 Nine Mile Lighthouse. Site Description - Nine Mile Lighthouse is located at the south end of Simcoe Island on Lake Ontario. The eddy covariance instrumentation is located on the historic lighthouse, built in 1833, and instrumented with eddy covariance data in 2016 by a network of scientists from both US and Canada (eventually to be called the Great Lakes Evaporation Network (GLEN)). The intent of GLEN has been to provide observations of over-lake meteorology and evaporation, improve forecasting of Great Lakes water levels, and support a wide variety of stakeholders, including the National Weather Service (NWS), Environment and Climate Change Canada, National Oceanic and Atmospheric Administration, U.S. Coast Guard, recreational boaters and commercial shipping, emergency management officials, and the Great Lakes research community.

Spence, Chris [Environment and Climate Change Cana↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗

Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces

These data are from Bandopadhyay et al., "Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces". This study aims to understand the soil microbial ecology along terrestrial-aquatic interfaces of a freshwater and estuarine region and how it relates to organic matter. We analyzed soil microbial (16S rRNA gene) and organic matter (Fourier-transform ion cyclotron resonance mass spectrometry, FTICR-MS) composition from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie (freshwater) and Chesapeake Bay (estuarine) regions. This dataset includes 16S rRNA gene amplicon data (only processed file types included here) and organic matter composition from FTICR-MS data (raw and processed files included here) from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie and Chesapeake Bay regions. These sites are part of the COMPASS-FME project (https://compass.pnnl.gov/FME/COMPASSFME). File formats and software needed to access files: 16S rRNA gene amplicon data: These files follow the format reported here https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format#updates-in-v1.0.1. As per this format, there are four file types reported: 1. Taxon tables (also called sequence-by-sample or OTU (operational taxonomic unit)/ESV (exact sequence variant) tables) : available in a .txt file format and accessible using TextEdit or MS Excel. 2. Representative sequences (also called consensus sequences) : available in a .fasta format and accessible using TextEdit. 3. Sequencing metadata : available in a MS Excel workbook file format and CSV file format 4. Bioinformatic metadata : available in a MS Excel workbook file format and CSV file format FTICR-MS data: 1. Raw data converted to a processed file with intensities of the peaks in the given samples : available in a MS Excel CSV file format 2. Processed file used in analyses and visualizations (appended as icr_long_) : available in a MS Excel CSV file format 3. Metadata file for ICR features (appended as icr_meta) : available in a MS Excel CSV file format

54 ENVIRONMENTAL SCIENCES↗

Financial Analysis of the High Flow Experiment conducted at the Glen Canyon Dam during Water Year 2023

The Glen Canyon Dam (GCD) is a Colorado River Storage Project (CRSP) power resource that is a component of the Salt Lake City Area Integrated Projects (SLCA/IP). The 2016 record of decision (ROD) for the GCD long-term experimental and management plan (LTEMP) final Environmental Impact Statement (EIS) specified criteria for GCD monthly water releases, daily and hourly operating limits, and experimental releases. This report examines the financial implications of the high flow experiment (HFE) conducted at GCD during the spring of Water Year (WY) 2023 as required by the LTEMP HFE Protocol. This report is part of a series of reports that describe the financial costs of LTEMP experimental releases since the 2016 ROD was adopted in January 2017. Previous reports analyzed the impact of several past HFEs and Bug Flow Experiments. This report focuses on the HFE conducted in April 2023. For this experimental release, financial costs of approximately $1.33 million were incurred because the HFE required sustained water releases exceeding the power plant’s maximum turbine flow rate. In addition, during the experiment, operators were not allowed to shape GCD power production, either to follow Firm Electric Service (FES) customer day-ahead energy deliveries or to respond to market prices. This study identifies the main factors contributing to the HFE costs and examines the interdependencies among these factors. It applies an integrated set of tools to estimate Western Area Power Administration (WAPA) financial impacts by simulating GCD under two types of cases; namely, (1) a “With Experiment” case that mimics the operations that actually occurred and (2) a “Without Experiment” case that simulates operations under the assumption that the HFE did not occur. The “With Experiment” case mimics operations during the HFE and the entire month the HFE occurred. It complies with LTEMP hourly and daily operating criteria. The “Without Experiment” case assumes that the HFE did not occur. The monthly water release volume is assumed to be identical under both cases. The Colorado River Storage Project Python-based model (CRiSPPy) model was the main modeling tool used to simulate the dispatch of the GCD hydropower plant and associated water releases from Lake Powell. In the modeling process, the research team used extensive data sets and historical information on SLCA/IP power plant characteristics, hydrologic conditions, and WAPA’s power purchases and sales prices. In addition to estimating the financial impact of the HFE, the team used the CRiSPPy model to gain insights into the interplay among ROD operating criteria, exceptions made to criteria to accommodate the HFE, and WAPA operating practices.

13 HYDRO ENERGY↗

Financial Analysis of the Smallmouth Bass Flows implemented at the Glen Canyon Dam during Water Year 2024

The Glen Canyon Dam (GCD) is a Colorado River Storage Project (CRSP) power resource that is a component of the Salt Lake City Area Integrated Projects (SLCA/IP). The 2016 record of decision (ROD) for the GCD long-term experimental and management plan (LTEMP) final Environmental Impact Statement (EIS) specifies criteria for GCD monthly water releases, daily and hourly operating limits, and experimental releases. This report presents a financial analysis of the Smallmouth bass (Micropterus dolomieu) (SMB) flows implemented at GCD during Water Year (WY) 2024. These bypass flows were introduced by the U.S. Bureau of Reclamation (USBR) as an emergency response to the growing threat posed by invasive SMB in the Colorado River ecosystem downstream of the dam. SMB are a non-native predatory species that pose a significant threat to native fish populations, including the endangered humpback chub (Gila cypha). The thermal regime below GCD, typically cold due to hypolimnetic releases from Lake Powell, has historically served as a thermal barrier limiting SMB establishment. However, persistently low reservoir levels in recent years have reduced stratification in Lake Powell, allowing warmer water to be released downstream. This has enabled SMB to spawn successfully below the dam, prompting urgent ecological concerns. To mitigate the risk of SMB proliferation, the USBR implemented a series of bypass flows in WY 2024. Drawn from a lower elevation than the penstocks, the bypass structures released cooler water downstream. These short-duration bypass flows aimed to keep temperatures cool enough to prevent SMB from spawning, thereby reducing the ecological threat posed by this invasive species. Although motivated by ecological objectives, these bypass flows came with financial tradeoffs. Releasing water through the bypass structures instead of the turbines at GCD reduced hydropower generation, resulting in a significantly lower financial position for Western Area Power Administration (WAPA), which is responsible for marketing the electricity produced by the GCD Powerplant. This report analyzes the financial impact of the SMB flows implemented from July to November 2024. These experimental releases led to an estimated financial cost of approximately $18.9 million, primarily driven by the substantial volume of water diverted through the bypass structures. This study applies an integrated set of tools to estimate WAPA financial impacts by simulating GCD under two types of cases; namely, (1) a “With Experiment” case that mimics the water operations that actually occurred, including the SMB bypass flows, and (2) a “Without Experiment” case that simulates operations under the assumption that the SMB flows did not occur. Both cases comply with LTEMP hourly and daily operating criteria, and the monthly water release volumes are assumed to be identical under both cases. The Colorado River Storage Project Python-based model (CRiSPPy) model was the main modeling tool used to simulate the dispatch of the GCD hydropower plant and associated water releases from Lake Powell. In the modeling process, the research team used extensive data sets and historical information on SLCA/IP power plant characteristics, hydrologic conditions, and WAPA’s power purchases and sales prices.

13 HYDRO ENERGY↗

Dual‐Transformer Deep Learning Framework for Seasonal Forecasting of Great Lakes Water Levels

Abstract The Great Lakes of North America form one of the largest freshwater systems on Earth, and their lake‐wide average water levels (lake levels) can fluctuate by more than 0.5 m on a seasonal scale. These fluctuations pose substantial challenges for coastal resilience, flood risk management, and navigation planning. Accurate seasonal forecasting of lake levels using traditional mechanistic models is challenging due to the complex physical mechanisms and coupled hydroclimatic processes involved. Recently, deep learning has gained prominence in geoscience applications for its ability to recognize intricate patterns within multiphysical data sets. Here, we introduce a novel Dual‐Transformer deep learning framework, tested on the Great Lakes. This architecture integrates two modified Transformer models: the Prophet, which predicts underlying trends, and the Critic, which refines the Prophet's predictions. The final lake level prediction is derived by weighting the outputs of both models through a multi‐layer perceptron, jointly trained with the Prophet and Critic to enhance overall accuracy. Our results demonstrate that the innovative learning framework achieves the highest prediction accuracy compared to established deep learning models when using identical input features. It attains a root mean square error of 4–7 cm in predicting lake levels up to 6 months in advance across the lakes. Additionally, the Dual‐Transformer model runs six orders of magnitude faster than conventional mechanistic models, producing results in less than one second on a typical personal computer. These findings suggest that our deep learning framework has strong potential to advance lake level prediction and carries important implications for water management and disaster mitigation, thereby enhancing the quality of life in coastal regions.

Chen, Yi [Great Lakes Research Center Michigan Tec↗

Compounding effects of Lake and urbanization on summer precipitation in the Greater Chicago area

Here, this study explores the impacts of Lake Michigan and Chicago's urbanization on precipitation patterns over the Greater Chicago Area, using 22 years of observational data and Weather Research and Forecasting (WRF) model simulations focused on an early summer rain event. Observational analysis reveals that urban areas consistently experience more precipitation than the adjacent southern Lake Michigan region throughout the year, particularly before 2015. However, this disparity has narrowed since 2016 due to a more rapid increase in heavy precipitation over the lake compared to urban areas. Specifically, lake precipitation has risen by 25 mm per year, compared to 15 mm per year over urban areas. Additionally, the number of days with precipitation exceeding 5 mm per day has been rising at a rate of 1.34 days per year over the lake and 0.84 days per year over urban areas. Modeling experiments reveal that both urbanization and lake effects, including lake breezes, enhance precipitation over urban areas, primarily through convergence induced by interactions between land and lake breezes. In contrast, these same factors suppress precipitation over the lake. The suppression results from Lake Michigan's stable environment, characterized by cooler surface temperatures, limited evaporation in early summer, and a high-pressure anomaly over the lake driven by urban heating, which creates upward motion over urban areas and downward motion over the lake, further influencing precipitation patterns.

Coastal urban↗

Towards continual machine learning for particle accelerators

This talk covers our work on errant beam prognostics at the Spallation Neutron Source (SNS), focusing on the end-to-end process from data collection to the development and deployment of predictive models in specific. A short overview of AIML work done for accelerators and current trends will be presented. We will walk through key steps involved in creating robust Machine Learning (ML) models, including model training, validation, and deployment in an operational setting. In addition to presenting our technical approach, we will share valuable lessons learned, emphasizing the importance of infrastructure to support the continuous adaptation of models to evolving data and system behaviors. This talk will provide insights into the challenges and solutions involved in applying ML to real-world operational environments, with a particular focus on managing data drift and changes in accelerator setup while ensuring model resilience over time.

Accelerator Physics↗

Enhancing fire emissions inventories for acute health effects studies: integrating high spatial and temporal resolution data

Daily fire progression information is crucial for public health studies that examine the relationship between population-level smoke exposures and subsequent health events. Issues with remote sensing used in fire emissions inventories (FEI) lead to the possibility of missed exposures that impact the results of acute health effects studies. This paper provides a method for improving an FEI dataset with readily available information to create a more robust dataset with daily fire progression. High temporal and spatial resolution burned area information from two FEI products are combined into a single dataset, and a linear regression model fills gaps in daily fire progression. The combined dataset provides up to 71% more PM 2.5 emissions, 69% more burned area, and 367% more fire days per year than using a single source of burned area information. The FEI combination method results in improved FEI information with no gaps in daily fire emissions estimates. The combined dataset provides a functional improvement to FEI data that can be achieved with currently available data.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

“Understanding Robustness Lottery”: A Geometric Visual Comparative Analysis of Neural Network Pruning Approaches

Deep learning approaches have provided state-of-the-art performance in many applications by relying on large and overparameterized neural networks. However, such networks are very brittle and are difficult to deploy on resource-limited platforms. Model pruning, i.e., reducing the size of the network, is a widely adopted strategy that can lead to a more robust and compact model. Many heuristics exist for model pruning, but our understanding of the pruning process remains limited due to the black-box nature of a neural network model. Empirical studies show that some heuristics improve performance whereas others can make models more brittle. Here, this work aims to shed light on how different pruning methods alter the network’s internal feature representation and the corresponding impact on model performance. To facilitate a comprehensive comparison and characterization of the high-dimensional model feature space, we introduce a visual geometric analysis of feature representations. We evaluated a set of critical geometric concepts decomposed from the commonly adopted classification loss and used them to design a visualization system to compare and highlight the impact of pruning on model performance and feature representation. The proposed tool provides an environment for an in-depth comparison of pruning methods and a comprehensive understanding of how the model responds to common data corruption. By leveraging the proposed visualization, machine learning researchers can reveal the similarities between pruning methods and redundancy in robustness evaluation benchmarks, obtain geometric insights about the differences between pruned models that achieve superior robustness performance, and identify samples that are robust or fragile to model pruning and common data corruption.

Li, Zhimin [Univ. of Utah, Salt Lake City, UT (Uni↗