Search NASASearch

SEARCH · Search NASA

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Statistical data analysis of x-ray spectroscopy data enabled by neural network accelerated Bayesian inference

Bayesian inference applied to x-ray spectroscopy data analysis enables uncertainty quantification necessary to rigorously test theoretical models. However, when comparing to data, detailed atomic physics and radiation transfer calculations of x-ray emission from non-uniform plasma conditions are typically too slow to be performed in line with statistical sampling methods, such as Markov Chain Monte Carlo sampling. Furthermore, differences in transition energies and x-ray opacities often make direct comparisons between simulated and measured spectra unreliable. Here, we present a spectral decomposition method that allows for corrections to line positions and bound–bound opacities to best fit experimental data, with the goal of providing quantitative feedback to improve the underlying theoretical models and guide future experiments. In this work, we use a neural network (NN) surrogate model to replace spectral calculations of isobaric hot-spots created in Kr-doped implosions at the National Ignition Facility. The NN was trained on calculations of x-ray spectra using an isobaric hot-spot model post-processed with Cretin, a multi-species atomic kinetics and radiation code. The speedup provided by the NN model to generate x-ray emission spectra enables statistical analysis of parameterized models with sufficient detail to accurately represent the physical system and extract the plasma parameters of interest.

47 OTHER INSTRUMENTATION

Scalable Data Center Capacity for DOE's AI Prototype: A Rapidly Available Gigawatt Data Center for DOE

The multilaboratory Gigawatt Data Center working group was commissioned to identify approaches to rapidly establish federal data centers with scalable capacities up to 1,000 MW. These state-of-the-art facilities will serve as hubs for interdisciplinary collaboration, industry partnerships, and transformative applications of artificial intelligence. The proposed strategic shift includes facilitating multilaboratory collaboration, prioritizing operational efficiency, expanding public–private partnerships, optimizing investments, ensuring long-term contractual flexibility, supporting open science and secure data enclaves, and exploiting high-speed national networks. Owing to their extensive experience and best practices, the US Department of Energy national laboratories are uniquely positioned to lead this initiative. We recommend conducting a feasibility analysis to rapidly identify the optimal sites for this initiative, and the effort will likely involve private industry for design, construction, financing, and operational integration. We also propose establishing multiple geographically diverse sites to ensure energy resilience, high operational reliability, and a diverse user base, thereby effectively addressing the nation’s critical needs.

42 ENGINEERING

Community Choice Aggregation(CCA) Data Collection Webinar for Status and Trends in the Voluntary Market Report (2024 Data) [Slides]

We have subcontracted LEAN Energy US, to help us improve our CCA data collection effort for the Annual Voluntary Energy Markets Data Report. LEAN Energy US (Local Energy Aggregation Network) is a national 501(c)3 non-profit organization dedicated to accelerating the country's transition to clean and renewable power, supporting competition and customer choice in the energy sector, and maintaining affordable electricity rates. We work in partnership with a range of organizations to actively support the formation and operational success of Community Choice Aggregation (CCA) programs around the country. This webinar, hosted in partnership with LEAN Energy US, is intended to introduce their members to our data collection effort and encourage CCAs in their network to participate.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Microbial community data from throughfall exclusion experiment: Metadata, SI, community composition, LefSe, and FunGuilR data tables from PARCHED Panama tropical forest soils, 2024-2025

Soil contains more carbon (C) than terrestrial vegetation and the atmosphere combined, with some of the largest terrestrial C stocks in tropical rainforests. Soil microbes decompose organic matter, playing a vital role in the storage or loss of soil C. With climate change, drought conditions are predicted to increase in many tropical regions, including both chronic drying and extended drought, potentially influencing these processes. This project explored the effects of chronic and seasonal drying on soil microbial communities across four distinct tropical forests in a long-term drying experiment. We investigated the effects of a chronic drying manipulation on soil microbial community abundance and variation across different forests and seasons. We also compared findings with previously published data from these forests after short-term drying. This project used soils from a long-term drying experiment established in 2018 across four seasonal lowland forests in Panama. Soils were collected from 0 – 10 cm depths during three seasonal periods in control and drying plots in 2024 and 2025 from a total of 32 plots (n = 4 per forest per treatment). The forests varied in baseline rainfall and soil fertility. We calculated alpha and beta diversity indices and compared taxonomic community composition. We found significant biogeographic variation in microbial diversity and taxonomy, with significant differences across the forests and significant effects of the drying treatment. Metadata and sample IDs are within Metadata_16S.csv and Metadata_ITS.csv. Relative abundance tables of every sample at every season are shown in the Excel workbooks 16S Relative Abundance.xlsx and ITS Relative Abundance.xlsx. They are then also shown in CSV files by each taxonomic level. Linear discriminant analysis effect size (LefSe) tables are shown for the full 16S and ITS datasets (n = 96), subsets for every site at every season (n = 8), and then for the forests with each plot merged by season (n = 8). FunGuildR data table of ITS data is uploaded.

Bacteria

MINERvA s Open Data Product: A First for Neutrino Data Preservation

Access to information on neutrino nucleus interactions is critical to the success of all neutrino oscillation experiments. MINERvA's rich dataset covers a range of energies and nuclei unique amongst experiments, and as such is critical to the community in building the important shared knowledge needed to unravel the mysteries of the neutrino. In particular, its dataset provides the greatest statistical coverage in in the range of neutrino energies pertinent for DUNE until DUNE's near detector begins operation. Historically, such significant datasets in neutrino physics have been preserved primarily through their published results. While meaningful and useful, this limits the ability to explore the data to its fullest extent as new perspectives continue to form. MINERvA has undertaken a major effort to break this trend and preserve its data in a format to be as analyzable as possible from outside the collaboration. This has culminated in the officially-released MINERvA Open Data Product for the community to take advantage of and utilize. Maintaining direct access to the dataset in an analyzable form will allow new insights to continue to be extracted indefinitely. This talk will cover the contents of this product, the information included (and excluded), the tools provided to utilize the product effectively, the support MINERvA intends to provide in its use, and some lessons learned through the process.

Last, David [Rochester U.] (ORCID:0000000245147183

Multi-omics data resource: Data package 22 (Pck022)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines. The data package consists of isolated pancreatic islets from adult male C57BL6/J mice treated with IL-1β, IFNγ or IL-1β + IFNγ for 6 h and submitted for scRNA-seq. This study focused on understanding the heterogeneity of the cytokine-mediated response. Data contributors: Jennifer S Stancill & John A Corbett: Department of Biochemistry, Medical College of Wisconsin, Milwaukee, WI, USA Data repository: GSE156175 Publication: 10.26508/lsa.202000949

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data resource: Data package 23 (Pck023)

The data package consists of isolated pancreatic islets from adult male C57BL6/J mice treated with IL-1β + IFNγ, IL-1β + IFNγ + NMMA, or NMMA alone for 18 h and submitted for scRNA-seq. This study focused on the cell-type-specific effects of nitric oxide signaling in islets and characterized the heterogeneity of responses. Data contributors: Jennifer S Stancill & John A Corbett: Department of Biochemistry, Medical College of Wisconsin, Milwaukee, WI, USA Data repository: GSE183010 Publication: 10.1093/function/zqab063

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data resource: Data package 24 (Pck024)

The data package consists of isolated pancreatic islets from 3 human donors treated with IL-1β, IFNγ or IL-1β + IFNγ for 6 h and IL-1β, IFNγ, IL-1β + IFNγ, IL-1β + IFNγ + NMMA or NMMA for 18 h and submitted for scRNA-seq. This study examines cytokine-stimulated changes in gene expression in human islets using single-cell RNA sequencing. Data contributors: Jennifer S Stancill & John A Corbett: Department of Biochemistry, Medical College of Wisconsin, Milwaukee, WI, USA Data repository: GSE251730 Publication: 10.1093/function/zqae015

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data resource: Data package 25 (Pck025)

This data package comprises omics datasets from human pancreatic islets treated with IL-1β + IFNγ or with estrogen (E2) for 18 h. Two RNA-seq datasets are available: the first is a discovery dataset involving human islets treated with or without IL-1β + IFNγ for 18 hours; the second is a validation dataset, where human islets are treated with or without IL-1β + IFNγ or E2 for 18 hours. DIA proteomic analysis was performed on the same validation dataset samples. Data contributors: Kiersten L. Webster, Sarah Tersey & Raghavendra G. Mirmir: Kovler Diabetes Center and Department of Medicine, The University of Chicago, Chicago, IL, 60637, USA. Soumyadeep Sarkar, Raghavendra Mirmira, Ernesto S. Nakayasu: Biological Sciences Division, Pacific Northwest National Laboratory, Richland, WA, 99354, USA. Data repository: RNA-seq: GSE310965 Proteomics: MSV000101892 Publication: PMID 41279069

Sarkar, Soumyadeep [Pacific Northwest National Lab

Fracture Intersections under Stress: Laboratory Data and Code [Data set]

The connectivity of natural and induced fractures governs the injection and withdrawal of fluids from subsurface reservoirs. Connectivity depends on intersections that control how fluids mix and move through the entire system. Here, we present data sets from 3D X-ray microscopy measurements of simple fracture networks under stress. 3D printing was used to create prismatic blocks that formed fracture networks composed of 2 orthogonal fractures. The network orientation was either "x" or "+" relative to an applied vertical stress. 3D data sets were collected for normal loads of 25, 100 and 200 Newtons for samples with fracture surfaces with either correlated or uncorrelated asperity distributions. The file contains data from the 12 samples analyzed along with an example code used to extract the intersection geometry. Additional experimental details can be found in the manuscript "Geologic Stress Modulates Fluid Mixing at Fracture Intersections" (10.1038/s43247-026-03525-9)and supplemental information to appear in Communications Earth & Environment in 2026.

02 PETROLEUM

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij

Data‐Driven Insights into Rare Earth Mineralization: Machine Learning Applications Using Functional Material Synthesis Data

Understanding rare‐earth element (REE) mineralization mechanisms is essential for developing efficient separation strategies. Although the geochemical pathways that generate REE deposits are qualitatively known, quantitative links between specific conditions and mineralization outcomes remain limited. Herein, the repurpose laboratory REE hydrothermal synthesis data—originally collected for functional‐materials fabrication—as a surrogate for studying mineralization with data‐driven methods. The compiled 1,200+ hydrothermal reaction records and trained three machine‐learning models—K‐nearest neighbors (KNN), random forest (RF), and extreme gradient boosting (XGB)—to predict product elements and phases from precursors, additives, reaction conditions, and engineered features. Validation shows XGB achieves the highest accuracy. Feature importance indicates thermodynamic properties of cations and anions dominate model decisions. Correlations reveal positive relationships among precursor concentration, reaction time, pH, and temperature, consistent with classical crystallization behavior. XGB‐based regressors are built to predict crystallization temperature and pH from precursor/product attributes. Performance is strongest when similar training examples exist, while accuracy declines for underrepresented reactions, notably REE carbonates and heavy‐REE systems. Overall, the study shows that functional‐materials datasets can illuminate REE mineralization and provide priors for exploration and processing. Expanding datasets with less‐studied chemistries and conditions will improve generality and support deposit discovery and more efficient REE recovery.

feature importance analysis

Utilization of Data Augmentation Techniques in Automated Inspection Systems for Defect Detection in Metals With Limited Data

Accurate identification of defects on metal surfaces is of great interest to many industry sectors, such as the automotive and aerospace industries. In contrast to conventional manual inspection techniques, recent automated inspection systems employ deep learning models trained to detect defects rapidly and precisely. The development of these models often requires a substantial image dataset to acquire adequate knowledge of defect features and enhance their predictive accuracy. When data is limited, augmentation techniques are often used to improve the precision and accuracy of defect detection systems. This study examined the prediction performance of two object detection models, namely Faster Region‐based Convolutional Neural Network (Faster R‐CNN) and You Only Look Once version 8 (YOLOv8), to identify dent defects in limited images of cast iron cylinder head surfaces. The original image set contains 46 images with 563 dents. To overcome limited data availability, common image augmentation techniques along with a copy‐paste method were applied. Results show that standard augmentation improved YOLOv8 accuracy by 8.00% and average precision (AP) by 3.00%. On the other hand, the copy‐paste technique achieved a 20.00% increase in accuracy and a 1% increase in AP with just 200 synthetic dents. Furthermore, these results provide support for using the copy‐paste augmentation strategy to enhance defect detection performance, with a limited dataset, contributing to more accurate defect identification in remanufacturing processes.

36 MATERIALS SCIENCE

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han