Search NASA⌕ Search

SEARCH · Search NASA

Results for “data distributions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Graph-Learning-Assisted State and Event Tracking for Solar-Penetrated Power Grids with Heterogeneous Data Sources

Unlike transmission systems, distribution systems do not typically contain sufficient metering to enable real-time state estimation. The lack of sufficient real-time measurements prohibits accurate and timely monitoring of the state of distribution systems. As a result, control and optimal operation of distribution systems, especially those containing large numbers of renewable generation units are not possible without proper data and information about the current state of the system. The main motivation of this project is to address this shortcoming by developing an approach which provides “predicted” real-time measurements so that they can be used to execute a distribution system state estimator. Thus, the objective of the project is to make the distribution systems fully observable, such that the hosting capacity for solar generation can be accurately estimated, and unnecessary solar curtailments can be avoided. In order to accomplish this goal, the project investigated the use of a grid-model-informed machine learning (ML) tool which integrates heterogeneous data streams obtained from AMI meters, SCADA as well as PMU measurements and created synchronous measurement snapshots for the state estimator (SE); and developed a hybrid robust SE which provides not only accurate state estimates but also real-time feedback for the ML model refinement.

14 SOLAR ENERGY↗

Distributed Resilience in High-Energy Physics Data Acquisition

Historical experience in the High-Performance Computing community teaches us that as computing systems grow, the instance of failures goes from rare to a regular occurrence. A survey of the growth in the size and complexity of Data AcQuisition (DAQ) networks in High-Energy Physics (HEP) experiments reveals that these networks are scaling exponentially, trending to a point where automated fault handling should be considered over the current manual practice, especially given the rarity of data such as in DUNE's mission to observe core-collapse supernovae. We propose a general system, DiDAQt, which is designed to provide fault detection and handling in HEP DAQs specifically, through MPI-like primitives that allow it to be added easily to existing systems. We evaluate the scalability and response time of a prototype on the FABRIC national testbed, with results indicating sufficient scalability for current and near-future DAQs as well as practical response times (under 1 microsecond decision time).

Wolosewicz, A. [IIT, Chicago]↗

Counter Data Paucity through Adversarial Invariance Encoding: A Case Study on Modeling Battery Thermal Runaway

Lithium-ion batteries, widely used for their durability and high energy storage, face the risk of internal short circuits leading to catastrophic thermal runaway events. These events, triggered by external stimuli like mechanical loads, pose safety concerns in applications such as electric vehicles. Detecting and understanding thermal runaway events is crucial, but physics-driven models struggle to explain the non-linear evolution of battery temperature during these events, considering factors like material composition and state-of-charge. Due to the rarity of these events and the cost of data collection, we propose a deep learning (DL) model to predict battery temperature responses during thermal runaway. The challenge lies in the scarcity of data, making traditional DL models prone to overfitting and learning low-quality representations of the complex process.Our approach introduces a novel few-shot architecture that incorporates an adversarially governed invariant encoding process. This architecture aims to distill "invariant" relationships by addressing distributional shifts in data across various battery properties, facilitating the detection of thermal runaway events. Specifically, our results demonstrate that deep learning models conditioned on these "invariant" representations outperform state-of-the-art baselines, achieving a remarkable 96.8% performance improvement in terms of the popular metric MAPE. This framework presents a promising direction for enhancing battery safety modeling, particularly in the context of rare and complex events like thermal runaway. Our code and code and dataset used for the paper are public1.

Tabassum, Anika [ORNL] (ORCID:0000000254600955)↗

Geometric Measures of Trustworthiness for Machine Learning Predictions

his report details the findings from the research and investigation of Geometric Measures of Trustworthiness for Machine Learning Predictions. We explored the trustworthiness of machine learning (ML) models’ predictions using geometric measures to quantify the similarity of a query point with the training data. Predictive uncertainty in ML can originate from at least three sources: (1) Model uncertainty, which represents the uncertainty in model form (e.g. decision tree, vs neural network) and estimating the model parameters from the training data, (2) Data uncertainty, which represents the natural complexities of the data such as class overlap and inherent noise, and (3) Distributional uncertainty, which represents the mismatch between the training and operational distributions. The proposed measures focus on measuring and explaining the data and distributional uncertainties by measuring the relationships of operational data with the training data.

97 MATHEMATICS AND COMPUTING↗

Detecting Unclassified Electromagnetic Signals for Secure Wireless Communication Using Open Set Recognition

We developed multiple machine learning methods for the detection and classification of new wireless communication waveforms, which is critical for targeted attacks in wireless networks and electronic warfare. Our machine learning models are capable of dynamically detecting security threats in near real time through our advanced open set recognition (OSR) approach. This model has demonstrated significant improvements in the detection of unknown waveforms, thereby enhancing the security and reliability of mission critical communications. Our approach to detecting uncertain security threats is novel; we advanced OSR techniques by incorporating domain knowledge of wireless signals. Specifically, we combined time and frequency domain model features to enhance the model’s performance. Utilizing an OSR approach eliminates the need for training data to be distributed similarly to the deployment environment and removes the requirement for the training set to contains all possible threat classes. This is crucial because it is often infeasible to determine and characterize all potential security threats in advance. Our model were trained on simulated data, generated in partnership with the University at Albany, State of New York. The data set contained a diverse array of wireless signals, including those with additive white Gaussian noise and multipath signals, with and without line of sight. This comprehensive training set allowed us to optimize our models to detect unknown waveforms under various challenging scenarios, such as low signal-to-noise ratios. By training on various waveforms, varying signal-to-noise ratio, and different sample sizes under normal conditions, our models were fine tuned to perform effectively in challenging environments.

99 - GENERAL AND MISCELLANEOUS↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Conditional distribution estimation of building characteristics with diffusion models for urban energy modeling

Understanding current energy consumption behavior in communities is critical for informing future energy use decisions and enabling efficient energy management. Urban energy models, which are used to simulate these energy use patterns, require large datasets with detailed building characteristics for accurate outcomes. However, such detailed characteristics at the individual building level are often unknown and costly to acquire, or unavailable. Through this work, we propose using a generative modeling approach to generate realistic building attributes to fill in the data gaps and finally provide complete characteristics as inputs to energy models. Our model learns complex, building-level patterns from training on a large-scale residential building stock model containing 2.2 million buildings. We employ a tabular diffusion-based framework that is designed to handle heterogeneous (discrete and continuous) features in tabular building data, such as occupancy, floor area, heating, cooling, and other equipment details. We develop a capability for conditional diffusion, enabling the imputation of missing building characteristics conditioned on known attributes. We conduct a comprehensive validation of our conditional diffusion model, firstly by comparing the generated conditional distributions against the underlying data distribution, and secondly, by performing a case study for a Baltimore residential region, showing the practical utility of our approach. Our work is one of the first to demonstrate the potential of generative modeling to accelerate building energy modeling workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Deconvoluting thermomechanical effects in X-ray diffraction data using machine learning

X-ray diffraction is ideal for probing the sub-surface state during complex or rapid thermomechanical loading of crystalline materials. However, challenges arise as the size of diffraction volumes increases due to spatial broadening and because of the inability to deconvolute the effects of different lattice deformation mechanisms. Here, we present a novel approach that uses combinations of physics-based modeling and machine learning to deconvolve thermal and mechanical elastic strains for diffraction data analysis. The method builds on a previous effort to extract thermal strain distribution information from diffraction data. The new approach is applied to extract the evolution of the thermomechanical state during laser melting of an Inconel 625 wall specimen which produces significant residual stress upon cooling. A combination of heat transfer and fluid flow, elasto-plasticity and X-ray diffraction simulations is used to generate training data for machine-learning (Gaussian process regression, GPR) models that map diffracted intensity distributions to underlying thermomechanical strain fields. First-principles density functional theory is used to determine accurate temperature-dependent thermal expansion and elastic stiffness used for elasto-plasticity modeling. The trained GPR models are found to be capable of deconvoluting the effects of thermal and mechanical strains, in addition to providing information about underlying strain distributions, even from complex diffraction patterns with irregularly shaped peaks.

36 MATERIALS SCIENCE↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

Modeling Plutonium Decorporation in a Female Nuclear Worker Treated with Ca-DTPA after Inhalation Intake

The present work models plutonium (Pu) biokinetics in a female former nuclear worker. Her bioassay measurements are available at the US Transuranium and Uranium Registries. The worker was internally exposed to a plutonium-americium mixture via acute inhalation at a nuclear weapons facility. She was medically treated with injections of 1 g Ca-DTPA on days 0, 5, and 14 after the intake. Between days 0 and 20, fecal and urine samples were collected and analyzed for 239 Pu and 241 Am. Subsequently, she was followed up for bioassay monitoring over 14 y, with additional post-treatment urine samples collected and analyzed for 239 Pu. The uniqueness of this dataset is due to the availability of: (1) both early and long-term bioassay data from a female with plutonium intake; (2) data on chelation therapy for a female; and (3) fecal measurement results. Chelation therapy with Ca- and/or Zn-salts of DTPA is known to aid in reducing the internal radiation dose by enhancing the excretion of plutonium and americium from the body. Such enhancement affects plutonium biokinetics in the human body, posing a challenge to the internal dose assessment. The current radiation dose assessment practice is to exclude the data affected by Ca-DTPA from the analysis. The present analysis is the first to explicitly model the chelation-affected bioassay data in a female by using a newly developed chelation model. Thus, the bioassay data collected during and after the Ca-DTPA administrations were used for biokinetic modeling and dose assessment. The Markov Chain Monte Carlo method was used to investigate model parameter uncertainty, based on the bioassay data and assumed prior probability distributions. A χ 2 /nData (number of data points) ≈ 1 was observed in this study, which indicates self-consistency of the data with the model. Results of this study show that the worker’s 239 Pu intake was 12 Bq, with a committed effective dose to the whole-body of 1.2 mSv and a committed equivalent dose to the bone surfaces, liver, and lungs of 37.8, 9.1, and 0.8 mSv, respectively. This study also discusses the worker’s dose reduction due to chelation treatment.

61 RADIATION PROTECTION AND DOSIMETRY↗

Fast Response Temperature by airborne measurements over BNF

The original data were collected during the AAF Engineering Flights (AEF2025) in the vicinity of the ARM Bankhead National Forest (BNF) Atmospheric Observatory (https://www.arm.gov/capabilities/observatories/bnf ) in northwestern Alabama in March 2025. The ARM Aerial Facility ArcticShark uncrewed aerial system (UAS, https://www.arm.gov/capabilities/observatories/aaf/uas) was based at the public-use airport of Posey Field, Alabama (FAA LID: 1M4, 34.28027778° N, 87.60055556° W, 283m MSL) from March 9 through March 24, 2025. The ArcticShark UAS performed nine flights, including eight research flights over the AMF3 (BNF Main Site) and Supplemental Facilities to measure atmospheric state, turbulence, surface IR temperature and imagery, and aerosol number concentration and size distribution. The current data set presents fast response temperature in the atmospheric boundary layer and lower free troposphere measured on the airborne platform throughout the field campaign. The primary instruments used to create the current data set were the fine wire thermocouple probe, the Aircraft Integrated Meteorological Measurement System (AIMMS-30), the Pitot-static system (part of UAS flight control), and the infrared gas analyzer sensor for H2O and CO2 (LI-840). All parameters used in temperature calculations (static pressure, True Air Speed, and absolute humidity in form of dew point temperature) were included in the data set. For user convenience, one additional parameter was also included: the type of flight flag (level, up, down, turn, and combination of thereof).

Air temperature, fast response↗

Turbulent Parameters by airborne measurements over BNF in March 2025

The original data were collected during the AAF Engineering Flights (AEF2025) in the vicinity of the ARM Bankhead National Forest (BNF) Atmospheric Observatory (https://www.arm.gov/capabilities/observatories/bnf ) in northwestern Alabama in March 2025. The ARM Aerial Facility ArcticShark uncrewed aerial system (UAS, https://www.arm.gov/capabilities/observatories/aaf/uas) was based at the public-use airport of Posey Field, Alabama (FAA LID: 1M4, 34.28027778° N, 87.60055556° W, 283m MSL) from March 10 through March 24, 2025. The ArcticShark UAS performed nine flights, including eight research flights over the AMF3 (BNF Main Site) and Supplemental Facilities to measure atmospheric state, turbulence, surface IR temperature and imagery, and aerosol number concentration and size distribution. The current data set presents a collection of turbulent parameters in the atmospheric boundary layer or lower free troposphere based on airborne measurement throughout the field campaign. The primary instruments used to create the current data set were the Aircraft Integrated Meteorological Measurement System (AIMMS-30) and the fine-wire thermocouple probe.

Atmosphere↗

CACTI: Fast Liquid Water Content

These data were collected during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI; https://www.arm.gov/research/campaigns/amf2018cacti ) field campaign in the Sierras de Córdoba mountain range of north-central Argentina as part of the ARM Aerial Facility (AAF) deployment. The ARM Aerial Facility Gulfstream-1 was operated from Las Higueras Airport (IATA: RCU, ICAO: SAOC), Río Cuarto, Córdoba, Argentina, for the Intensive Observation Period (IOP) from Nov. 1 through Dec. 15, 2018. The G-1 aircraft performed 22 research flights over the first ARM Mobile Facility (AMF1) location in the Sierras de Córdoba mountain range to measure atmospheric state and turbulence, cloud water content and droplet size distributions, aerosol precursor gases, and aerosol chemical composition and size distributions. The current data set presents re-processed Particle Volume Monitor PVM-100A (aka Gerber probe) data: Liquid Water Content (LWC), Particle Surface Area (PSA), and the effective droplet radius (re) averaged to 50Hz, 10Hz, and 1Hz.

54 ENVIRONMENTAL SCIENCES↗

Geospatial Analysis of Built Infrastructure and Modeled Household Driving Patterns

The level of access to opportunities for a location can be quantified in the amount of time it takes to travel from a departure point to the destinations that somebody would want or need to visit. Isochrone maps are geometric representations of the area accessible from a departure point within a set amount of time. Informed in part by the National Household Travel Survey, this report merges location data for amenities and opportunities across six frequent destination categories – employment, education, health, food, community, and transportation – with isochrone maps generated by the TravelTime API, whose departure points are census tract population-weighted centroids. Using a “Points-In-Polygon” analysis, destinations that fall within a census tract’s isochrone are tallied as accessible from the region within one of three time thresholds: 15-, 30-, and 45-minutes by the walking, cycling, public transit, and driving modes of travel. We find that access to a high number of jobs within a typical commute duration is negatively correlated with annual household vehicle miles traveled (VMT). The spatial distribution of our data suggests that the high household VMT frequently seen surrounding the edges of major cities may be related to worker commutes into the city core, and that the high household VMT frequently seen in rural tracts may be related to the longer travel distances required to access a variety of key opportunities from these areas.

99 GENERAL AND MISCELLANEOUS↗

ACE-ENA: Fast Liquid Water Content

These data were collected during the Aerosol and Cloud Experiments in the Eastern North Atlantic field campaign as part of ARM Aerial Facility deployment (ACE-ENA, https://www.arm.gov/research/campaigns/aaf2017ace-ena). The ARM Aerial Facility Gulfstream-1 was deployed at Lajes Air Base (IATA: TER, ICAO: LPLA), on Terceira Island in the Azores, Portugal, for the two Intensive Observation Periods from June 20 through July 22, 2017 (IOP#1) and from January 11 through February 22, 2018 (IOP#2). The G-1 aircraft performed 20+19 research flights over the ARM Eastern North Atlantic (ENA) site and Atlantic Ocean to measure atmospheric turbulence, cloud water content and drop size distributions, aerosol precursor gases, aerosol chemical composition and size distributions. The current data set presents re-processed Particle Volume Monitor PVM-100A (aka Gerber probe) data: Liquid Water Content (LWC), Particle Surface Area (PSA), and the effective droplet radius (re) averaged to 50 Hz, 10 Hz, and 1 Hz.

54 ENVIRONMENTAL SCIENCES↗

CHELAX-BNF: Fast Response Temperature by airborne measurements

The original data were collected on board the ARM Aerial Facility ArcticShark uncrewed aerial system (UAS; https://www.arm.gov/capabilities/observatories/aaf/uas ) during the “Characterizing HEterogeneous Land-Atmosphere eXchanges at BNF” field campaign (CHEAX-BNF; https://arm.gov/research/campaigns/aaf2025CHELAX-BNF ). The ARM Aerial Facility ArcticShark UAS was based at the public-use airport of Posey Field, AL (FAA LID: 1M4, 34.28027778° N, 87.60055556° W, 283m MSL) from May 28 to June 23, 2025. The ArcticShark UAS performed 5 flights, including 4 research flights over the BNF Main Site (ARM Mobile Facility 3, https://arm.gov/capabilities/observatories/amf ) and Supplemental Facilities to measure atmospheric state, turbulence, surface IR temperature and imagery, aerosol number concentration, and aerosol size distribution. The current data set presents fast response temperature in the atmospheric boundary layer and lower free troposphere measured on the airborne platform throughout the field campaign. The primary instruments used to create the current data set were the fine wire thermocouple probe, the Aircraft Integrated Meteorological Measurement System (AIMMS-30), the Pitot-static system (part of UAS flight control), and the infrared gas analyzer sensor for H2O and CO2 (LI-840). All parameters used in temperature calculations (static pressure, True Air Speed, and absolute humidity in the form of dew point temperature) were included in the data set. For user convenience, one additional parameter was also included: the type of flight flag (level, up, down, turn, and combination of thereof).

Air temperature, fast response↗