Search NASASearch

SEARCH · Search NASA

Results for “Data Summarization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING

BLOC site - NREL Scanning Lidar / Derived Data

This dataset contains daily csv files summarizing data from 10-min wind statistics from ground-based Doppler lidar at the BLOC site for the WFIP3 event log. See https://a2e.energy.gov/ds/wfip3/bloc.lidar.10min.z01.c1/summary. This lidar was Halo XR #216 through February 24, 2025, Halo XR #217 from February 24, 2025, through April 17, 2025, and again Halo XR #216 after that.

17 WIND ENERGY

BARG site - NREL Scanning Lidar / Derived Data

This dataset contains daily csv files summarizing data from 10-min wind statistics from Doppler lidar at the BARG site for the WFIP3 event log. See https://a2e.energy.gov/ds/wfip3/bloc.lidar.10min.z01.c1/summary

17 WIND ENERGY

RHOD site - NREL Scanning Lidar / Derived Data

This dataset contains daily csv files summarizing data from 10-min wind statistics from ground-based Doppler lidar at the RHOD site for the WFIP3 event log. See https://a2e.energy.gov/ds/wfip3/rhod.lidar.10min.z01.c1/summary

17 WIND ENERGY

NANT site - Doppler Lidar / Derived Data

This dataset contains daily csv files summarizing data from 10-min wind statistics from ground-based Doppler lidar at the NANT site for the WFIP3 event log. See https://a2e.energy.gov/ds/wfip3/nant.lidar.10min.z02.c1/summary

17 WIND ENERGY

NANT site - Doppler Lidar / Derived Data

This dataset contains daily csv files summarizing data from 10-min wind statistics from ground-based Doppler lidar at the NANT site for the WFIP3 event log. See https://a2e.energy.gov/ds/wfip3/nant.lidar.10min.z01.c1/summary

17 WIND ENERGY

Idaho National Laboratory Quality Of Service Dataset

The code is designed to run tests to generate and collect data from a Wi-Fi network using OPENWRT or a simulated a 5G network using Open5gs and UERANSIM. The tests simulate the network performing downloads or uploads of various files with a varying number of concurrent users. The tests use tcpdump to collect the network traffic but only stores the summarized data. The summarized datasets will be included.

Krome, Cameron [Idaho National Laboratory (INL), I

Regional Oil and gas Aerial Methane Synthesis model (ROAMS) v2.0

The Regional Oil and gas Aerial Methane Synthesis model is a tool to convert the results of wide-area, source-resolved aerial methane remote sensing surveys of oil and natural gas infrastructure in a given region into methane emissions inventories (estimates of the magnitude and breakdown of methane emissions from the surveyed infrastructure). The tool leverages databases of source-resolved methane emissions detected in aerial surveys, aerial survey coverage information (which areas were measured and when), data summarizing surveyed oil and natural gas infrastructure and production (derived from third-party databases), as well as state-of-the-art mechanistic emissions simulation tools to characterize emissions too small for the aerial system to see. The regional methane emissions estimates produced by this tool are much more granular in both space and asset type than common satellite- or flux tower-based regional estimates. Unlike other tools for converting site-level measurements into regional emissions estimates, our unique geostatistical approach integrates aerially measured emissions with limited need for statistical extrapolation, which can be highly sensitive to modeler assumptions. As a result, ROAMS-based estimates of regional methane emissions from oil and gas activity are widely viewed as highly credible, as evidenced by the success of Dr. Sherwin's recent paper in Nature.

Sherwin, Evan [Lawrence Berkeley National Laborato

Regional Oil and gas Aerial Methane Synthesis model (Analytica) (ROAMS Analytica) v1.5.2

The Regional Oil and gas Aerial Methane Synthesis model (Analytica) is a tool to convert the results of wide-area, source-resolved aerial methane remote sensing surveys of oil and natural gas infrastructure in a given region into methane emissions inventories (estimates of the magnitude and breakdown of methane emissions from the surveyed infrastructure). This version is written in the Analytica programming language, and this version accompanies a correction in preparation for submission to Sherwin et al. 2024 (Nature). The tool leverages databases of source-resolved methane emissions detected in aerial surveys, aerial survey coverage information (which areas were measured and when), data summarizing surveyed oil and natural gas infrastructure and production (derived from third-party databases), as well as state-of-the-art mechanistic emissions simulation tools to characterize emissions too small for the aerial system to see. The regional methane emissions estimates produced by this tool are much more granular in both space and asset type than common satellite- or flux tower-based regional estimates. Unlike other tools for converting site-level measurements into regional emissions estimates, our unique geostatistical approach integrates aerially measured emissions with limited need for statistical extrapolation, which can be highly sensitive to modeler assumptions. As a result, ROAMS-based estimates of regional methane emissions from oil and gas activity are widely viewed as highly credible, as evidenced by the success of Dr. Sherwin's recent paper in Nature.

Sherwin, Evan [Lawrence Berkeley National Laborato

Evaluation of Technologies to Mitigate the Presence of Gaseous Elemental Mercury in Waste Disposal Containers

A study was conducted to evaluate sorbent technologies that can mitigate the presence of elemental mercury (Hg⁰) in waste containers for mercury-contaminated debris (MCD). Decontamination and demolition (D&D) activities at the Y-12 National Security Complex (Y-12) and other U.S. Department of Energy (DOE) Oak Ridge Reservation (ORR) facilities generate MCD requiring offsite disposal. The debris is packaged in appropriate waste containers and may be temporarily stored onsite prior to transport for treatment and/or disposal. During transportation of loads that had no visible liquid Hg at the point of origin, temperature changes can cause Hg⁰ to evaporate, condense, and form droplets on container walls. Furthermore, vibration during transportation could cause beads of Hg to be released from the debris, container walls, and ceiling, resulting in pools of liquid Hg⁰ on the container floor. Waste acceptance criteria (WAC) limitations for commercial disposal facilities, the Nevada National Security Site (NNSS), and ORR mixed low-level waste landfills prohibit the presence of any free liquids in containers identified as a solid waste form. Potential solutions to mitigate the presence of residual liquids that could be formed through vapor condensation include the use of sorbents or similar materials to capture and stabilize volatile Hg⁰ vapors and thus ensure compliance with landfill WAC requirements. This report summarizes data from small-scale laboratory experiments conducted to evaluate sorbent materials for Hg⁰ vapor suppression and sorption of liquid Hg⁰ that could form under relevant transportation and disposal conditions. A series of experiments was conducted to evaluate commercial sorbent materials and their effectiveness for Hg⁰ sorption across a temperature range from 19.4°C to 60°C. The impact of residual moisture on sorption was investigated under relevant conditions, and leaching tests were performed to assess the stability of Hg⁰ captured sorbent materials. The results provide estimates for sorbent quantities needed for a given Hg⁰ mass loading based on experimental results exposing sorbents to gaseous and liquid Hg⁰ at various mass ratios. Overall, sorbents that were most effective for Hg⁰ vapor suppression were brominated activated carbons, mackinawite-based sorbents coated on vermiculite, and sulfur-modified granular activated carbon. Elevated temperatures and moisture conditions did not result in significant increases of Hg⁰ headspace concentrations, and the materials also demonstrated high sorption capacities for the sorption of liquid Hg⁰.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Differences in urban plant community compositions across an urban-rural gradient in Knoxville, TN

Urban forests, or vegetation in areas under heavy human influence, provide many ecosystem services to urban residents such as localized cooling via evapotranspiration, shade, filtering of air pollution, and the associated health benefits of natural spaces. In order to quantify the magnitude of localized cooling by trees growing in varying levels of urbanization (based on % impervious surfaces, e.g., buildings, pavement), urban forest species composition, tree size, and tree density must be characterized. As a part of Oak Ridge National Laboratory’s (ORNL) urban forest temperature study, we conducted tree censuses in five Knoxville city parks where ORNL meteorological stations are deployed. Moreover, we measured every woody plant ≥ 5 cm diameter at breast height (DBH) within a 50 m radius of each site’s meteorological station for its DBH and species identification. When possible, individuals were identified down to species. Certain genera (Quercus spp., Carya spp., Pinus spp.) were identified down to genera in interest of time. Individual and total site basal area were calculated from measured DBH data. Results show notable differences in urban plant community compositions and total woody plant basal area across sites, with more urban sites closer to downtown (West View and SEEED) having lower tree basal area than the more suburban sites (West Hills, Cumberland Estates, and Victor Ashe). We identified 54 species across all sites, with West Hills and Victor Ashe having the highest species diversity. Our results show differences in forest compositions and sizes across Knoxville, which are currently informing ORNL’s evapotranspiration estimates for each site. Data Summary: Census data for West Hills (WH), Cumberland Estates (CE), Victor Ashe (VA), West View (WV), and Socially Equal Energy Efficient Development or SEEED (SD) urban forests in Knoxville, TN, USA, including tree size based on diameter at breast height (DBH; 1.3 m), species identification (Latin and common names), and basal area per stem (BA=π×[.5*DBH]^2). Field data are summarized in this file: “Community_Composition_Data.CSV”. Site-specific data detailing each site’s coordinates, number of stems measured at DBH, average tree DBH, α-diversity (number of species present), and total site basal area (sum of individual basal areas per site) are in this file: “Site_Comparisons.CSV”.

Warren, Jeffrey [ORNL] (ORCID:0000000206804697)

Hyper Spectral Anomaly Detection

The HSA is a statistics based anomaly detection model. The model performs unsupervised anomaly detection, based on a datapoint's density and similarity within a dataset. Density and similarity data are encoded into an affinity matrix. The affinity matrix is evolved to summarize the data's structure on greater topographical scales within the data's function space. The set of evolved affinity matrices and an anomaly score vector are passed to a user defined penalized objective function. The penalized objective function of anomaly scores is then minimized. Data points where the absolute value of the z-scores of anomaly scores greater than a specified threshold are predicted as anomalies. A novel multi-filter feature has also been implemented. To reduce false positive rates, the multi-filter records the indexes of the HSA predictions. A new dataset and data loader are instantiated consisting of all the initial HSA predictions and non-anomalous data points in a 10% and 90% split respectively. The HSA is then run through this data set and a count of number of times a data point is predicted is kept. In this way the initial predictions may be compared with data spanning the entire dataset. After the multi-filter is complete, all datapoints will have an associated anomaly score, as well as a multi-filter prediction count to further filter the anomalous predictions.

Rogers, DempseyD [Idaho National Laboratory (INL),

FleetREDI Dashboard Fleet DNA Data Summaries

Developing daily duty cycle summaries for every vehicle-day within NLR’s Fleet DNA database was a key output of the FleetREDI project. This project captured second-by-second GPS and controller area network (CAN) data on in-use medium- and heavy-duty fleet vehicles and then summarized the data to provide an overview of vehicle operation throughout the United States. These data summaries were then displayed in aggregated formats on the FleetREDI dashboard, where users can explore the data within Fleet DNA. Fleet DNA’s clearinghouse of commercial fleet vehicle operating data helps vehicle manufacturers and developers optimize vehicle designs and helps fleet managers choose advanced technologies for their fleets. This online tool, which provides data summaries and visualizations similar to real-world "genetics" for medium- and heavy-duty fleet vehicles, helps users understand the broad operational range of commercial vehicles across vocations and weight classes.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Expansion of the Direct Feed High-Level Waste Glass Composition in the High Al Range

Baseline glass compositions have been developed and demonstrated for successful immobilization of Hanford high-level waste (HLW) prepared through a pretreatment process. Recent enhanced waste glass formulations have shown promise to increase the waste loading of pretreated sludge compositions from a broader range of HLW feeds. This project proposes to increase the loading of minimally pretreated Hanford HLW in glass by expanding the existing database and glass property-composition models. Estimated direct-feed high level waste (DFHLW) compositions were generated by the Hanford Tank Operations Contractor and used by Pacific Northwest National Laboratory to determine target glass compositions. Gaps in existing data were identified including one high-priority gap in the high Al compositional region. This report summarizes the data collected during the characterization of the DFHLW High Al Glass Matrix. These glasses were intentionally designed with high aluminum concentrations (15 to 30 wt%) and a high likelihood of nepheline formation, which is known to negatively affect glass durability. Some glasses were expected to either fail or approach property constraints to fill data gaps in poorly understood regions of the compositional space due to lack of data. Out of the 50 glasses tested, 14 glasses formed nepheline, while the model predicted nepheline formation in 20 glasses. All quenched glasses met the product consistency test durability constraint; however, 8 glasses failed this constraint after undergoing the canister centerline cooling treatment. Additionally, 17 glasses did not meet the viscosity constraints, 4 failed the EC constraints, and 2 exceeded the allowable T2% for spinel crystal formation. All glasses satisfied the SO 3 solubility limit. The resulting dataset provides valuable information to improve model accuracy and reduce prediction uncertainty. These insights will ultimately support the development of more robust glass formulation strategies, enabling higher waste loadings, reducing operational risks, and expanding the processing envelope.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Testing of High S Matrix Glasses to Expand DFHLW Glass Compositional Ranges (Rev.1)

Gaps in glass composition-property data for direct-feed high-level waste (DFHLW) have recently been identified. One such gap is the region of high sulfur solubility since previous, pretreated, high-level wastes contained very little sulfur. Filling this data gap will significantly broaden the range of process flowsheet options including minimal washing and will allow for optimized waste loading in DFHLW glasses. This report summarizes the data collected during the characterization of the DFHLW High S Glass Matrix (HS24). A glass matrix of 50 glass compositions was developed to evenly cover the DFHLW composition region for high sulfur glass. Matrix glasses were designed to expand the composition region outside the current component concentration and property limits so as to reduce uncertainties at the limits. The 50 matrix glasses were fabricated and tested for properties important to the success of DFHLW vitrification including: compositions, canister centerline cooling (CCC) crystallinity and isothermal crystallinity, density, viscosity, electrical conductivity (EC), product consistency test (PCT) response, toxicity, and sulfate solubility. Melter materials corrosion testing is reported elsewhere. These glasses were intentionally designed to have high SO 3 solubilities (0.7 to 2.2 SO 3 wt%) in compositional regions that had not been previously explored. Forty-eight glasses showed the measured SO 3 content retained >80% of the target SO 3 and the densities of all the glasses ranged from 2.49 g·cm -3 to 2.74 g·cm -3 . While the model predicted nepheline formation in 5 glasses, one of the tested 50 CCC glasses formed nepheline, and 35 glasses formed Cr-containing phases such as spinels and eskolaite. Only five glasses were amorphous after CCC treatment where 44 glasses with detectable crystals contained =10 wt% crystals and only one glass had > 10 wt% crystals. None of the glasses exceeded the allowable T 2% for spinel crystal formation at 950 ºC (i.e., no glasses had >2 wt% spinel at 950 ºC) during isothermal crystal fraction tests where 10 glasses showed no crystalline phases at or below 950 ºC. All the glasses (except one which failed being slightly lower than the target) satisfied the SO 3 constraint while 98 glasses did not meet the viscosity constraints and 4 failed the EC constraints. Six quenched (Q) and six CCC glasses failed the Defense Waste Processing Facility (DWPF) Environmental Assessment (EA) glass PCT threshold and 3 Q and 4 CCC failed the PCT design constraint. One glass exceeded the WTP delisting limits for Cr via EPA Method 1311 (i.e., Toxicity Characteristic Leaching Procedure, TCLP). It should be emphasized that some of these glasses were specifically designed to approach or even exceed certain property constraints, as filling data gaps in these regions will provide the greatest benefit for future model development by improving accuracy and reducing uncertainties. These insights will ultimately support the development of more robust glass formulation strategies, enabling higher waste loading, reducing operational risks, and expanding the processing envelope.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

NGEE Arctic Phase 4 Plant Functional Type Framework for Pan-Arctic Vegetation

The NGEE-Arctic research team identified a common set of hierarchical plant functional types (PFTs) for pan-arctic vegetation that we will use across our research activities. Interdisciplinary work within a large team requires agreement regarding levels of functional organization so that knowledge, data, and technologies can be shared and combined effectively. The team has identified plant functional types as a crucial area where such interoperability is needed. PFTs are used to represent plant pools and fluxes within models, summarize observational data, and map vegetation across the landscape. Within each of these applications, varying levels of PFT specificity are needed according to the specific scientific research goal, computational limitations, and data availability. By agreeing on a specific hierarchical framework for grouping variables in our vegetation data, we ensure the resulting research products will be robust, flexible, and scalable. In this document, we lay out the agreed upon PFT framework with definitions and references to existing literature. Table 1 included in the "NGA700_Phase4PFTFramework_about*" file outlines the relationship between NGEE-Arctic Phase 4, Tier 1 PFTs and the PFTs used within prominent arctic literature as well as publications by the NGEE-Arctic team during phases 1-3.This dataset consists of a table detailing a hierarchical PFT framework that spans 4 tiers with the most granular PFTs listed in tier 1 and the most general PFTs in tier 4. The PFTs within each tier has a single column in the dataset where the PFTs are named and a separate column where the characteristics used to define that PFT are listed. Grey fill of the cells is used to indicate where a given PFT starts to “lose” tier 1 details as you look from left to right. Note the excel file has merged cells to indicate grouping of PFTs across the Tiers- it will not translate into a delimited filetype (.csv, .txt, etc) without modification thus the hierarchical PFT framework table is available in three different file formats: 1) NGA700_Phase4PTS.xlsx – maintains the merged cells and grey fill; 2) NGA700_Phase4PTS.csv – merged cells are split, and grey fill is removed; 3) NGA700_Phase4PTS.pdf – image of the table with merged cells and grey fill. Metadata document included as a *.pdf and file-level metadata and data dictionary as *.csv files.

54 ENVIRONMENTAL SCIENCES

Dataset_for_Conserved_macromolecular_architecture_of_Poplar_secondary_cell_walls_revealed_by_ssNMR_and_atomistic_modeling

This dataset contains solid-state 13C NMR data and atomistic molecular dynamics simulation files supporting the study of nanoscale secondary cell wall architecture across 13 genetically diverse Populus trichocarpa genotypes grown under uniform greenhouse conditions in 13C-enriched CO2 atmospheres (~89% 13C enrichment).The dataset contains two collections of solid-state 13C NMR data. (1) 200 MHz data (Bruker Avance III HD, 4 mm HX probe, 10 kHz MAS): raw Bruker TopSpin experiment folders and DMFIT-exported ascii spectra for selective and non-selective 1D 13C-13C spin diffusion experiments (3000 ms mixing) used to quantify inter-polymer spatial proximities, and short-mixing (1 ms) reference spectra used for polymeric abundance quantification by spectral deconvolution. (2) 600 MHz data (Bruker Avance III, 1.6 mm PhoenixNMR HXY probe, 30 kHz MAS): raw Bruker TopSpin experiment folders containing 2D CORD, 2D CP-INADEQUATE, and 13C/1H relaxation (T1, T1rho) experiments for all 13 genotypes, with processed Excel workbooks per experiment type. Molecular dynamics simulation code, coordinate files, and analysis scripts (NAMD/CHARMM/Python) for six atomistic cell wall models are included. Summarized ssNMR data are compiled into a single excel file and subjected to statistical analysis. Multivariate analysis code (PCA, Pearson correlation) and summary data are provided as excel worksheets and Jupyter notebooks (Python 3).

09 BIOMASS FUELS

FleetREDI Insight: Beverage Delivery in New York City

Capturing real-world data is critical to improving efficiency and supporting technology advancements in commercial vehicles. FleetREDI’s insights provide detailed duty cycle information and highlight unique aspects of the given dataset. Each insight delivers a quick look at the collected data by summarizing the operation and identifying key findings of the initial analysis. This FleetREDI insight explores beverage delivery tractors operating in New York City. Last-mile beverage delivery supports local bars and restaurants throughout Manhattan and the broader New York City area. Manhattan Beer Distributors is a beverage delivery company operating in Manhattan and the Bronx. Logging devices were installed in 17 vehicles, and operational data were collected between August and October 2022. Two types of vehicles were included in data collection: 7 tractors and 10 bay trucks. Using NLR’s FleetREDI data platform, this dataset provides a summary of daily operation to help understand duty cycle characteristics. This includes daily distance, fuel use, and estimated engine-produced energy consumption for 17 bay trucks and tractors that operated more than 7,500 miles in slow-speed urban operation. ![FleetREDI beverage delivery](FleetREDI-beverage-delivery-nyc.jpg)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI