Search NASA⌕ Search

SEARCH · Search NASA

Results for “heterogeneous data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence↗

Roadmap for transforming heterogeneous catalysis with artificial intelligence

Artificial intelligence (AI) is poised to transform heterogeneous catalysis, opening avenues for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental and chemical sectors. This promise, however, hinges on overcoming fundamental barriers, including limitations in data availability and quality, challenges in the generalizability and interpretability of data-augmented decisions, and the persistent gap between in silico predictions and experiments. Furthermore, we outline a forward-looking roadmap for deeply integrating AI into heterogeneous catalysis with an AI-ready data ecosystem, multimodal foundation models, and ultimately autonomous laboratories to accelerate the development of next-generation catalytic technologies via AI-empowered human–machine collaboration.

Computational methods↗

Data-driven analysis to understand GPU hardware resource usage of optimizations

With heterogeneous systems, the number of GPUs per chip increases to provide computational capabilities for solving science at a nanoscopic scale. However, low utilization for single GPUs defies the need to invest more money in expensive accelerators. Although related work develops optimizations to improve application performance, none studies how these optimizations impact hardware resource usage or average GPU utilization. Here, this paper takes a data-driven analysis approach in addressing this gap by (1) characterizing how hardware resource usage affects device utilization, execution time, or both, (2) presenting a multiobjective metric to identify important application-device interactions that can be optimized to improve device utilization and application performance jointly, (3) studying hardware resource usage behaviors of several optimizations for a benchmark application, and finally (4) identifying optimization opportunities for several scientific proxy applications based on their hardware resource usage behaviors. Furthermore, we demonstrate the applicability of our methodology by applying the identified optimizations to a proxy application, which improves the execution time, device utilization, and power consumption by up to 29.6%, 5.3% and 26.5% respectively.

Computer science↗

Descriptor: Infrastructure Perception and Control: Multi-Sensor Object Tracking Dataset (IPC-MSOT)

Traffic intersections are crucial and challenging nodes in transportation networks where multiple lanes of vehicles and pedestrians converge. Traffic accidents often occur at traffic intersections, including a large proportion of traffic fatalities and about one-half of all traffic injuries in the United States. Object detection data were collected in 2024 across three intersections in Colorado Springs, CO, USA, over the course of multiple days and various times to induce a heterogeneous mix of traffic conditions and behaviors. The purpose of the data collection exercises was to learn various attributes about infrastructure sensors and to build a repository of high-resolution, object-level data that can be used for research and development (e.g., to develop multisensor data fusion algorithms). The Infrastructure Perception and Control:Multi-Sensor Object tracking (IPC-MSOT) dataset was collected as part of the U.S. Department of Transportation's Strengthening Mobility and Revolutionizing Transportation (SMART) project, where the city of Colorado Springs, Colorado, and the National Renewable Energy Laboratory collaborated to collect object-level trajectory data from road users using multiple types of infrastructure sensors deployed at different intersections. This dataset allows for testing of late-stage sensor fusion algorithms and their ability to ingest multimodal sensor data, and it can be utilized by traffic engineers to design and evaluate trajectory-based signal control strategies.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

U-net architected deep material network training with microstructure local field information

The Deep Material Network (DMN) has recently emerged as a powerful reduced-order modeling framework for simulating the mechanical response of heterogeneous materials such as composites. Unlike most data-driven approaches that directly learn a material’s response under prescribed loading, the DMN acts as a homogenization operator, learning the kinematic constraints and mechanical interactions of the underlying microstructure. However, traditional DMN training relies exclusively on homogenized effective properties derived from Direct Numerical Simulations (DNS), discarding the rich local field data that govern microstructural interactions. In this work, we extend the DMN framework to incorporate such local field information into the offline training process. Utilizing a U-Net architecture, we augment the DMN training objective to include the first and second statistical moments of the local stress fields obtained from linear DNS. This ensures that the learned network topology not only fits the effective stiffness but also accurately reflects the internal local stress and strain partitioning of the microstructure. The results confirm that supervising the localization process during training yields a superior surrogate model, reducing local prediction errors by an order of magnitude and significantly improving generalization to unseen nonlinear constitutive behaviors compared to traditional DMNs.

36 MATERIALS SCIENCE↗

Metrology for femtosecond pulsed x-ray heating in diamond anvil cell experiments at the European XFEL: Revisiting the iron phase diagram up to 150 GPa

The development of pulsed intense x-ray sources, such as free electron laser, offers new avenues for high pressure experiments. Here, we study the feasibility and metrology of x-ray heating in diamond anvil cells at the European x-ray free electron laser. This method enables one to volumetrically heat the sample while inhibiting chemical migration and probing the crystallographic structure of the sample throughout the heating with a high repetition rate. We focus our study on iron, whose phase diagram is well established up to 100 GPa, to explore the possibilities and limitations of this technique. We volumetrically heat iron samples at starting pressures ranging from 10 to 138 GPa, using the x-ray beam pulsed at 4.5 MHz in a serial pump-and-probe experimental design. Experimental challenges arise from temperature gradients within the sample, changes in temperature at the 100 ns timescale, the difficulty of direct temperature estimates, the effect of thermal pressure, and the presence of metastable crystallites due to rapid cycles of heating and cooling. Hence, we develop a multi-crystal-like data processing method that allows us to account for sample heterogeneity in probed conditions. We then calibrate our measurements using known physical properties of iron under pressure. Thermal pressure in our experiments increases from 4% of the isochoric prediction at 10 GPa to 23% at 138 GPa, and we show that our data are in agreement with most previous observations of iron in this pressure range. The method can now be implemented at higher pressures and temperatures and on materials with unknown phase diagrams.

Materials science↗

A Multi-Region SEIR Model Incorporating Inter-County Mobility and Time-Dependent Transmission Dynamics: Application to COVID-19 Disease Outbreak Data in North Carolina.

Classical infectious disease compartmental models typically do not incorporate spatial heterogeneity or mobility. We develop a multi-region susceptible-exposed-infected-recovered (SEIR) model in which disease dynamics are coupled via inter-region mobility and the transmission rate is both region and time dependent. We calibrate the model using rolling averages of daily COVID-19 data in all 100 North Carolina counties. Mobility parameters are prescribed using daily inter-county commuter data. The number of transmission rate parameters is substantially reduced by hypothesizing that the dynamics correlate with county-level population density. Parameter estimation is carried out using several objective functions with error terms at different scales. An additive combination of least squares error at the county-level and the state-level, along with a quadratic transmission rate polynomial, yields the lowest overall error at both spatial scales. The calibrated model is used to simulate regional effects of perturbing disease transmission rates in adjacent counties and to illustrate effects of the state’s mobility infrastructure on disease dynamics and spread for a new disease outbreak.

COVID-19 modeling↗

Monitoring spatiotemporal evolution of fractures during hydraulic stimulations at the first EGS collab testbed using anisotropic elastic-waveform inversion

The EGS Collab project acquired continuous active-source seismic monitoring (CASSM) data before, during, and after hydraulic stimulations at the first testbed at the depth of 4850 ft (1478 m) at the Sanford Underground Research Facility in Lead, South Dakota, for monitoring fracture creation and evolution. CASSM acquisition was conducted using 24 hydrophones, 18 accelerometers, and 17 piezoelectric sources within four fracture-parallel wells and two orthogonal wells. 3D anisotropic traveltime tomography and anisotropic elastic-waveform inversion of the campaign cross-borehole seismic data show that the rock within the stimulation region is a heterogeneous horizontal transverse isotropic medium. Here we use these inversion results as the initial models and apply 3D anisotropic first-arrival traveltime tomography and 3D anisotropic elastic-waveform inversion to the CASSM data acquired after each stimulation in May, 2018 and December, 2018. We observe the spatiotemporal evolution of seismic velocities and anisotropic parameters caused by hydraulic fracture stimulations, showing the regions of rock alternation caused by hydraulic fracture stimulation.

15 GEOTHERMAL ENERGY↗

Characterizing the GD-1 Stream with DESI DR2 Data: Thin Stream and Hot Cocoon

GD-1 is among the longest, coldest stellar streams in the Milky Way, making it an ideal target for probing dark matter substructure through dynamical heating. We present a catalog of 608 spectroscopically confirmed GD-1 members from the first three years of Dark Energy Spectroscopic Instrument (DESI) observations. This constitutes the largest homogeneous spectroscopic sample of GD-1, doubling the number of members previously available only through heterogeneous compilations combining multiple surveys with different systematics. Using these data, we derive updated stream tracks in sky position, proper motion, and radial velocity that extend over $100^\circ$ of the stream. We apply a Gaussian mixture model to decompose the stream into a dynamically cold thin component ($σ_V = 2.49\pm 0.28$ km s$^{-1}$, width $= 0.23\pm0.01^\circ$) and a kinematically hot cocoon ($σ_V = 6.13\pm0.75$ km s$^{-1}$, width $= 2.18\pm0.17^\circ$). The cocoon contains $\sim30\%$ of members and its velocity dispersion is consistent with $\sim11$ Gyr of heating by cold dark matter subhalos. We also detect a large proper motion dispersion ($41.36\pm4.98$ km s$^{-1}$) along the stream direction in the cocoon component. This feature indicates a significant line-of-sight distance spread in the cocoon, and its origin will be further explored in a forthcoming paper. These measurements demonstrate the power of DESI spectroscopy for characterizing the multi-component phase-space structure of stellar streams and constraining small-scale dark matter substructure.

Jarvis, Emma [Toronto U.] (ORCID:0009000656127336)↗

Evolution of the ATLAS event data model for the HL-LHC

The upcoming high-luminosity run of the CERN Large Hadron Collider (HL-LHC) will yield an unprecedented volume of data. In order to process this data, the ATLAS collaboration is evolving its offline software to be able to use heterogeneous resources such as graphical processing units (GPUs) and field-programmable gate arrays (FPGAs). To reduce conversion overheads, the event data model (EDM) should be compatible with the requirements of these resources. While the ATLAS EDM has long allowed representing data as a structure of arrays, further evolution of the EDM can enable more efficient sharing of data between CPU and GPU resources. Some of this work will be summarized here, including extensions to allow controlling how memory for event data is allocated and the implementation of jagged vectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]↗

Using Machine Learning to Understand Electric and Hybrid Vehicles Ownership in Burdened and Nonburdened Communities

Transitioning to electric and hybrid vehicles (EHVs) for all communities is a pivotal step toward sustainable transportation and environmental conservation. This paper aims to understand the adoption of EHVs, focusing on burdened communities (BCs) in the United States. The EHV ownership-based analysis combines two datasets—behavioral data from the Puget Sound Regional Travel Survey integrated with BCs (Justice40) data covering transportation insecurity, environmental burden, social vulnerability, health vulnerability, and climate and disaster risk burden. After creating this unique database, descriptive analysis and modeling are used to analyze the data and predict EHV ownership in the future. Specifically, we use a new method that combines particle swarm optimization (PSO) with a stacking model named PSO-Stacking, which incorporates heterogeneous base learners of machine learning and deep learning. PSO applies a customized objective function to select the optimal hyperparameters for heterogeneous learners within the stacking model, effectively addressing challenges such as multicollinearity, data imbalance, nonlinearity, and overfitting. The proposed solution covers more accurate results than standard benchmark models for EHV ownership in BCs and non-BCs. In addition, the results of the PSO-Stacking method are explained using the local interpretable model-agnostic explanations technique. Results show a negative correlation between the BCs indicators, that is, higher transportation insecurity associated with lower EHV ownership. Furthermore, BCs have higher future climate risk scores, diesel particulate matter levels, and PM2.5 in the air than non-BCs because of higher conventional vehicle ownership. These communities are at higher risk and can benefit from electrification, EV infrastructure, and EV policies to address environmental challenges.

Aslam, Zeeshan [ORNL]↗

MatRIS: Addressing the Challenges for Portability and Heterogeneity Using Tasking for Matrix Decomposition (Cholesky)

The ubiquitous in-node heterogeneity of HPC and cloud computing platforms makes software portability and performance optimization extremely challenging. Described here, the MatRIS multilevel math library abstraction framework employs tasking to alleviate these difficulties. MatRIS includes the IRIS task-based runtime on the bottom level and exposes different layers of abstraction to render algorithms architecturally agnostic. MatRIS ensures the decomposition and creation of tasks that represent the necessary encapsulation of the optimized kernels from both vendor and open-source math libraries. Once built, MatRIS can select different combinations of accelerators at runtime, making it portable even on diverse heterogeneous architectures. By leveraging the IRIS runtime’s features for managing heterogeneity, MatRIS deploys algorithms that remove the need to specify orchestration and data transfer. This study describes how the serial task abstraction of a tiled Cholesky factorization is made portable and scalable in the case of multi-device and multi-vendor heterogeneity on a node with NVIDIA and AMD GPUs by using MatRIS. First, we demonstrate that Cholesky in MatRIS provides multi-GPU scalability that offers competitive performance versus cuSolverMG. Then, we present the challenges and opportunities for heterogeneous execution.

Monil, M. A. H.↗

CAFE AU LAIT: Compute-Aware Federated Augmented Low-Rank AI Training

Federated finetuning is crucial for unlocking the knowledge embedded in pretrained Large Language Models (LLMs) when data are geographically distributed across clients. Unlike finetuning with data from a single institution, federated finetuning allows collaboration across multiple institutions, enabling the utilization of diverse and decentralized datasets while preserving data privacy. Given the high computing costs of LLM training and the emphasis on energy efficiency in Federated Learning (FL), Low-Rank Adaptation (LoRA) has emerged as a widely adopted algorithm due to its significantly reduced number of trainable parameters. However, this assumes that all data silos have the necessary computing resources to compute local updates of LLMs. Nevertheless, in practice, the computing resources across clients are highly heterogeneous: while some may have access to hundreds of GPUs, others might have limited or no GPU access. Recently, federated finetuning using synthetic data has been proposed, allowing clients to participate in a collaborative training run without training LLMs locally. However, our experimental results reveal a performance gap between models trained using synthetic data and those trained using local updates. Motivated by the observed heterogeneity in computing resources and the performance gap, we propose a novel two-stage algorithm that leverages the storage and computing capabilities of a strong server. In the first stage, under the coordination of the strong server, clients with limited computing resources collaborate to generate synthetic data, which is transferred to and stored on the strong server. In the second stage, the strong server uses this synthetic data on behalf of the resource-constrained clients to perform federated LoRA finetuning alongside clients with sufficient computing resources. This approach ensures that all clients can participate in the finetuning process. Experimental results demonstrate that incorporating local updates from even a small fraction of clients improves performance compared to using synthetic data for all clients. Furthermore, we incorporate the Gaussian mechanism in both stages to guarantee client-level differential privacy.

Wang, Jiayi [ORNL]↗

A dynamic solvent chamber propagation estimation framework using RNN for warm solvent injection in heterogeneous reservoirs

Warm solvent injection (WSI), injecting low-temperature solvent into formations to reduce the viscosity of heavy oil, is a clean technology for heavy oil production through reducing greenhouse gas emissions and water usage. The success of WSI operation depends on the uniform development and propagation of solvent chambers in reservoirs. However, reservoir heterogeneity stemming from shale barriers plays a detrimental role in the conformance of solvent chamber development and oil production rate. In this work, we developed a novel recurrent neural network (RNN)-based framework with the capability of efficiently tracking and estimating the solvent chamber positions in heterogeneous reservoirs based on only production time-series data. The developed estimation model utilizes the “sequence-to-sequence" mapping methodology to correlate observed production time-series sequence and solvent chamber edge sequence via a long short-term memory (LSTM) algorithm. The trained RNN models exhibit high accuracy, evidenced by the predicted dynamic solvent chamber locations match the corresponding true locations from numerical simulation, with a high coefficient of determination (R 2 ) and a low mean squared error. Specifically, the achieved R 2 values exceed 0.98 on both the training and testing data. The developed RNN-based workflow was tested via several cases from both regularly- and irregularly-shaped shale barriers, and the results were promising. The predicted solvent chambers showed strong agreement with those obtained from numerical simulations. The major benefits of this workflow include reducing computational time and saving overall monitoring and tracking costs for conventional techniques. In conclusion, the present work would provide a good demonstration of the capability of practical integration of machine learning methods in solving engineering problems.

58 GEOSCIENCES↗

Remote Sensing of Live Fuel Moisture for Wildfires Using SMAP Satellite Observations

Live Fuel Moisture (LFM) is a critical parameter for wildfire risk assessment, traditionally measured by labor-intensive field sampling. However, sampled LFM data are influenced by site-specific factors, such as local vegetation types and plant traits, and are often collected retrospectively after wildfire events, making it difficult to obtain pre-fire data for predictive applications. Here, we evaluate the relationship between LFM and Vegetation Water Content (VWC) and Soil Moisture (SM) retrieved from SMAP L-band brightness temperature using the Maximum Entropy Production (MEP) approach. The MEP-retrieved VWC exhibited strong correlation with in situ measurements of LFM ( r > 0.6) in the Western U.S. The integration of high-resolution vegetation coverage data enhances the detection of sub-grid vegetation heterogeneity. This study demonstrates the operational potential of remote sensing derived VWC as a scalable proxy of LFM, supporting its application in regional assessment of wildfire risk.

Cho, Kyeungwoo [Georgia Institute of Technology, A↗

4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer’s Diagnosis

Multimodal neuroimaging provides complementary structural and functional insights into both human brain organization and disease-related dynamics. Recent studies demonstrate enhanced diagnostic sensitivity for Alzheimer’s disease (AD) through synergistic integration of neuroimaging data (e.g., sMRI, fMRI) with tabular data (e.g., behavioral and cognitive tests). However, the intrinsic heterogeneity across modalities (e.g., 4D spatiotemporal fMRI dynamics vs. 3D anatomical sMRI structure) presents critical challenges for discriminative feature fusion, often leading to information loss or biased fusion. To bridge this gap, we propose M2M-AlignNet: a multimodal co-attention network with latent alignment for early AD diagnosis using sMRI and fMRI. At the core of our approach is a multi-patch-to-multi-patch (M2M) contrastive loss function that quantifies and reduces representational discrepancies via weighted patch correspondence, explicitly aligning fMRI components across brain regions with their sMRI structural substrates without one-to-one constraints. Additionally, we propose a latent-as-query co-attention module to autonomously discover fusion patterns, circumventing modality prioritization biases while minimizing feature redundancy. We conduct extensive experiments to confirm the effectiveness of our method and highlight the correspondence between fMRI and sMRI as AD biomarkers.

Wei, Yuxiang [Georgia Institute of Technology]↗