Search NASA⌕ Search

SEARCH · Search NASA

Results for “sparse data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Antiviral discovery using sparse datasets by integrating experiments, molecular simulations, and machine learning

Computational methods have demonstrated success in identifying virucidal agents, effectively contributing to the discovery of novel virucidal molecules. In this study, we developed a machine learning (ML) model, trained on a small dataset, to predict inhibitors of human enterovirus 71 (EV71), a pathological agent that causes severe disease in children and immunocompromised adults. Despite the dataset’s limitation, comprising of only 36 compounds tested, our ML framework demonstrated significant predictive capability. Notably, experimental validation revealed that five out of the eight compounds predicted by our model from the Chinese cosmetic material list exhibited virucidal activity. The inhibitor effects displayed by the main active compounds were further confirmed by molecular dynamics simulation. This underscores the potential of our AI-driven approach to bypass data constraints in identifying active molecules against viral pathogens.

60 APPLIED LIFE SCIENCES↗

Structural response reconstruction using a system-equivalent singular vector basis

Here, this paper develops a novel method for reconstructing the full-field response of structural dynamic systems using sparse measurements. The singular value decomposition is applied to a frequency response matrix relating the structural response to physical loads, base motion, or modal loads. The left singular vectors form a non-physical reduced basis that can be used for response reconstruction with far fewer sensors than existing methods. The contributions of the singular vectors to measured response are termed singular-vector loads (SVLs) and are used in a regularized Bayesian framework to generate full-field response estimates and confidence intervals. The reconstruction framework is applicable to the estimation of single data records and power spectral densities from multiple records. Reconstruction is successfully performed in configurations where the number of SVLs to identify is less than, equal to, and greater than the number of sensors used for reconstruction. In a simulation featuring a seismically excited shear structure, SVL reconstruction significantly outperforms modal FRF-based reconstruction and successfully estimates full-field responses with as few as two uniaxial accelerometers. SVL reconstruction is further verified in a simulation featuring an acoustically excited cylinder. Finally, response reconstruction and uncertainty quantification are performed on an experimental structure with three shaker inputs and 27 triaxial accelerometer outputs.

42 ENGINEERING↗

AmeriFlux FLUXNET-1F CA-LU2 Lutose

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site CA-LU2 Lutose. This is the FLUXNET version of the carbon flux data for the site CA-LU2 Lutose produced by applying the standard ONEFlux (1F) software. Site Description - Lutose is a peat plateau that burned from a moderate forest fire in June 2007. Prior to the fire the site was likley an open canopy of stunted black spruce and a ground layer of Labrador tea shurbs and lichen or sphagnum. No black spruce survived the fire and lichens were still completely absent during this study. By 2019, most charred tree boles has fallen over and vegetation recovery was dominated by dense Labrador tea shrubs and sparse regenerating black spruce with around >150cm of peat.

Schulze, Christopher [University of Alberta]↗

Novel CHI3L1 ‐Associated Angiogenic Phenotypes Define Glioma Microenvironments: Insights From Multi‐Omics Integration

ABSTRACT The CHI3L1 signaling pathway significantly influences glioma angiogenesis, but its role in the tumor microenvironment (TME) remains elusive. We propose a novelCHI3L1‐associated vascular phenotype classification for glioma through integrative analyses of multiple datasets with bulk and single‐cell transcriptome, genomics, digital pathology, and clinical data. We investigated the biological characteristics, genomic alterations, therapeutic vulnerabilities, and immune profiles within these phenotypes through a comprehensive multi‐omics approach. We constructed the vascular‐related risk (VR) score based onCHI3L1‐associated vascular signatures (CAVS) identified by machine learning algorithms. Utilizing unsupervised consensus clustering, gliomas were stratified into three distinct vascular phenotypes: Cluster A, marked by high vascularization and stromal activation with a relatively low levels of tumor‐infiltrating lymphocytes (TILs); Cluster B, characterized by moderate vascularization and stromal activity, coupled with a high density of TILs; and Cluster C, defined by low vascularization and sparse immune cell infiltration. We observed that the CAVS effectively indicated glioma‐associated angiogenesis and immune suppression by single‐cell RNA‐seq analysis. Moreover, the high‐VR‐score group exhibited enhanced angiogenic activity, reduced immune response, resistance to immunotherapy, and poorer clinical outcomes. The VR score independently predicted glioma prognosis and, combined with a nomogram, provided a robust clinical decision‐making tool. Potential drug prediction based on transcription factors for high‐risk patients was also performed. Our study reveals thatCHI3L1‐associated vascular phenotypes shape distinct immune landscapes in gliomas, offering insights for optimizing therapeutic strategies to improve patient outcomes.

Oncology↗

Data-Driven Surrogate Modeling with Microstructure-Sensitivity of Viscoplastic Creep in Grade 91 Steel

Abstract To support the development of advanced steel alloys tailored to withstand extreme conditions, it is imperative to account for the mechanical performance of components, while considering the influence of local microstructure on the macroscopic response. To this end, this study focuses on the development of microstructure-sensitive constitutive models for the mechanical response of Grade 91 steel exposed to extreme thermo-mechanical environments. Polynomial chaos expansion (PCE) surrogates are used to emulate high-fidelity polycrystal simulations of the viscoplastic response of Grade 91 steel as a function of the microstructure fingerprint (e.g., dislocations and precipitates). To cover a wide temperature–stress domain, two separate PCE surrogates—one that captures softening and the other that captures hardening behavior—are combined using another (sparse) Gaussian process regression model. The resulting constitutive creep surrogate model is integrated within the MOOSE finite element framework to simulate the intricate effects of microstructure, in particular MX-phase precipitates, on a component with a graded microstructure. Surrogate sensitivity analysis is applied to quantify the relevant impact of spatially varying microstructure on the creep response in a test-case involving a Grade 91 alloy with a prototypical weld.

36 MATERIALS SCIENCE↗

Proteome-wide analysis of protein stability in Escherichia coli under acid stress

Knowledge of protein acid sensitivity remains sparse and is largely derived from low-throughput, enzyme-specific assays. We used a scalable framework to map acid stability across the Escherichia coli proteome to assess the acid stability of 1,675 unique proteins, estimating pH 50 values for over 90% of them. The parameter pH50 was defined as the pH value at which only 50% of the initial protein remains in solution following acid treatment. Proteome-wide pH 50 values ranged from 2.28 to 6.33 (median 5.11). Approximately 9% of detected proteins remained stable across all tested pH conditions. Our results align with published data and the assay of citrate synthase (GltA) performed here. Protein acid stability differed significantly by subcellular localization: periplasmic proteins were relatively more abundant in the acid-stable group, cytoplasmic proteins were abundant at pH 50 values 4.5–5.5, and inner membrane proteins at higher pH 50 between 5.5 and 6.0. Outer membrane proteins were too few to draw strong conclusions regarding enrichment within specific pH 50 groups. Notably, the periplasmic binding protein of the molybdate ABC transporter (ModA), was enriched after incubation at low pH. Estimated pH 50 values showed no correlation with protein isoelectric point and molecular weight. Together, this work provides the first proteome-wide map of protein acid stability and establishes a general framework for studying different chemical stressors.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Variance Preserving Spectral Subsampling

Generating statistically faithful short-duration gamma-ray spectra from a single long measurement is essential in nuclear safeguards, supporting tasks such as algorithm development and machine-learning applications, especially when list-mode data are unavailable. Existing subsampling methods often distort the statistical characteristics of genuine short-duration measurements, leading to biased or unreliable analytical outcomes and thereby undermining downstream tasks. In this work, we compare five subsampling approaches using a benchmark set of 156 genuine replicate spectra collected with a high-purity germanium detector. We evaluate each method with respect to run-to-run variance, channel-to-channel variance, and preservation of total counts (losslessness). Across a wide range of subsampling ratios, only binomial subsampling without replacement consistently reproduces the statistical properties of genuine short-duration spectra, maintaining proper dispersion even in sparse spectral regions and perfectly preserving total counts. These results provide a mathematically principled and practically validated framework for generating synthetically shortened spectra when true short-duration measurements are unavailable.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

Evidence for a Single Holocene Paleoseismic Event on the Pajarito Fault, Northern New Mexico

Low-slip rate fault systems tend to be less studied than their high-slip rate counterparts, and paleoseismic techniques used to study them may pose challenges in interpretation that differ from high-slip rate systems. A good example of this is the Pajarito fault system (PFS), a normal fault complex within the Rio Grande rift. Despite numerous previous paleoseismic trenching studies conducted on the PFS between 1990 and 2003, considerable uncertainty remains regarding its Holocene paleoseismic history, particularly for the primary Pajarito fault (PF). To further clarify the PF paleoseismic history, we present data from paleoseismic investigations of 6 trenches at 3 distinct locations along the PF. Though the totality of the age and structural data obtained in this study is complex and not entirely consistent with any one interpretation, a single Holocene paleoearthquake occurring younger than ∼1,600 to 2,300 kcal yr BP is the simplest interpretation. It is possible that the PF records two Holocene events, with a penultimate event 6.9–2.4 kcal yr BP event and the aforementioned most recent event (MRE) between 2.3 and 1.6 kcal yr BP. However, only a single wall of one trench, out of a total of 12 walls in our 6 trenches, provides evidence supporting that interpretation. This study finds evidence of a single late Holocene paleoseismic event on the PF and sparse evidence for 2 Holocene paleoseismic events on the PF and highlights the benefits of logging multiple trench walls to better understand the complexity that results from this low-slip rate, low-deposition-rate fault system.

58 GEOSCIENCES↗

A projection method for particle resampling

Particle discretizations of partial differential equations are advantageous for high-dimensional kinetic models in phase-space due to their better scalability than continuum approaches with respect to dimension. Complex processes collectively referred to as particle noise hamper long time simulations with particle methods. One approach to address this problem is particle mesh adaptivity, or remapping, known as particle resampling and remeshing. Here, this work introduces a resampling method that projects particles to and from a (finite element) function space. The method is simple, using standard sparse linear algebra and finite element techniques, and it preserves all moments up to the order of a polynomial represented exactly by the continuum function space. It is distinguished from most other mesh-based methods in that new particle positions and number are decoupled from the mesh, allowing particle and continuum meshes to be adapted relatively independently. While this work is developed with structured particle and continuum phase-space grids on 1X + 1V Vlasov-Poisson models of Landau damping and two-stream instability, the method is well-suited to unstructured grids. Stable long time dynamics are demonstrated up to time T = 500. Reproducibility artifacts and data are publicly available.

Kinetic methods↗

Accurate and uncertainty-aware multi-task prediction of HEA properties using prior-guided deep Gaussian processes

Surrogate modeling techniques have become indispensable in accelerating the discovery and optimization of high-entropy alloys (HEAs), especially when integrating computational predictions with sparse experimental observations. This study systematically evaluates the training and testing performance of four prominent surrogate models—conventional Gaussian processes (cGP), Deep Gaussian processes (DGP), encoder-decoder neural networks for multi-output regression and eXtreme Gradient Boosting (XGBoost)—applied to a hybrid dataset of experimental and computational properties of the 8-component HEA system Al-Co-Cr-Cu-Fe-Mn-Ni-V. We specifically assess their capabilities in predicting correlated material properties, including yield strength, hardness, modulus, ultimate tensile strength, elongation, and average hardness under dynamic/quasi-static conditions, alongside auxiliary computational properties. The comparison highlights the strengths of hierarchical deep modeling approaches in handling heteroscedastic, heterotopic, and incomplete data commonly encountered in materials science. Our findings illustrate that combined surrogate models such as DGPs infused with machine-learned priors outperform other surrogates by effectively capturing inter-property correlations and by assimilating prior knowledge. This enhanced predictive accuracy positions the combined surrogate models as powerful tools for robust and data-efficient materials design.

36 MATERIALS SCIENCE↗

Radiation induced athermal diffusivity in uranium mononitride

Uranium mononitride (UN) is one of the ceramic nuclear fuel alternatives to oxide fuel considered for light water reactors and advanced reactor designs. Properties like self- and fission gas diffusivity need to be better understood, given that they influence key fuel performance phenomena such as fission gas swelling and release. In particular, the radiation induced athermal (D 3 ) diffusivity remains challenging to accurately predict and has only been sparsely characterized in UN, despite its importance as it likely governs diffusion at the low temperatures this high-thermal-conductivity fuel form may operate. Molecular Dynamics simulations are used to estimate the mean square displacement induced by a primary knock-on atom (PKA) with a given kinetic energy. These results are combined with the PKA energy distributions obtained from binary collision approximation calculations to obtain the displacement due to a particular fission fragment. Finally, this is combined with experimental fission fragment yields to determine the displacement due to an average fission event and, thus, express the athermal diffusivity as a function of the fission rate density. These results are in excellent agreement with available experimental data. In conclusion, a particular importance is given to the understanding and the quantification of the variability of these results.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

AmeriFlux US-TLR Timberlake Observatory for Wetland Restoration (TOWeR)

This is the AmeriFlux version of the carbon flux data for the site US-TLR Timberlake Observatory for Wetland Restoration (TOWeR). Site Description - This tower is located in Timberlake Forest, a restored forested wetland located 4 km from Albermarle Sound on the North Carolina coast. The forest to the south of the tower had historically been ditched, drained and converted to agricultural land, and subsequently reforested with Black Gum and Cypress trees and naturally re-wetted in 2006. This southern region has a bottomland forest ecosystem which saw a large-scale succession event of pine trees following the reforestation efforts. The northern part of the forest had been ditched, but never drained or converted for other land uses. It is much wetter and more sparsely vegetated when compared to the southern region, and has a mixed bottomland forest and swamp ecosystem.

Rey-Sanchez, Camilo [North Carolina State Universi↗

Deep Learning Reconstruction of Daily Soil CO 2 Efflux Reveals Biogeochemical Insights and Reduces Annual Estimate Uncertainty Despite Limited Daily Predictability

Soil CO 2 efflux is commonly measured monthly or seasonally, leaving daily dynamics poorly resolved and contributing to global estimation uncertainty. We trained a single Long Short-Term Memory (LSTM) model to predict daily soil CO 2 efflux across 82 globally distributed sites in COSORE, with 0.2%–46.9% daily data coverage from 2003 to 2020. Despite using far fewer sites than are typically used to train a single deep learning model, with observations biased toward temperate mesic sites, the LSTM model performed well at approximately one-third of sites, reconstructed nearly 2 decades of daily efflux, and outperformed commonly used approaches for estimating daily efflux when applied to the same data set. Performance was weakest at pronounced peaks and troughs and at non-temperate sites with <1.5 years of observations and irregular data patterns. Nevertheless, annual efflux from reconstructed daily data had <40% error even at underperforming sites, substantially improving estimates derived from monthly and seasonal sampling (maximum errors of 95% and 136%, respectively). Temperature sensitivity (Q 10 ) estimated from reconstructed daily predictions closely matched estimates from daily observations, whereas Q 10 values derived from monthly or seasonal observations deviated substantially, suggesting that coarse temporal sampling may contribute to uncertainty in reported Q 10 values. Consistent daily reconstructions further enabled trend analyses for well-performing, predominantly temperate sites and showed increasing soil CO 2 efflux at most sites from 2003 to 2020, with more variable summer trends. Despite limitations, these results demonstrate the potential of LSTM models to reconstruct daily soil CO 2 efflux and reduce estimation uncertainties from sparse observations.

Smykalov, Valerie [Pennsylvania State University, ↗

Solving high-dimensional inverse problems using amortized likelihood-free inference with noisy and incomplete data

Here, we present a likelihood-free probabilistic inversion method based on normalizing flows for high-dimensional inverse problems. The proposed method is composed of two complementary networks: a summary network for data compression and an inference network for parameter estimation. The summary network encodes raw observations into a fixed-size vector of summary features, while the inference network generates samples of the approximate posterior distribution of the model parameters based on these summary features. The posterior samples are produced in a deep generative fashion by sampling from a latent Gaussian distribution and passing these samples through an invertible transformation. We construct this invertible transformation by sequentially alternating conditional invertible neural network and conditional neural spline flow layers. The summary and inference networks are trained simultaneously. We apply the proposed method to an inversion problem in groundwater hydrology to estimate the posterior distribution of the log-conductivity field conditioned on spatially sparse time-series observations of the system’s hydraulic head responses. The conductivity field is represented with 706 degrees of freedom in the considered problem. Comparison with the likelihood-based iterative ensemble smoother PEST-IES method demonstrates that the proposed method accurately estimates the parameter posterior distribution and the observations’ predictive posterior distribution at a fraction of the inference time of PEST-IES.

conditional invertible neural network↗

Advancing Asset Management in Water Infrastructure Systems

Aging water system infrastructure, including drinking water, wastewater, and stormwater, poses a growing challenge for utilities and municipalities. These water systems have well documented challenges with respect to their age, condition, and level of service. ASCE annual report cards consistently rate these infrastructure systems in the United States as underfunded, overcapacity, or past service life (ASCE 2025 Report Card). For example, Chini and Stillwell (2017) estimated that the mean water loss in drinking water systems, i.e., non-revenue water, is approximately 16% across the United States. These concerns are not just relegated to the United States, with Courtenay, British Columbia, identifying 17% of their water main pipes as in a ‘poor’ condition state, defined as a category condition 5 out of 5 (City of Courtenay, 2024). These cases illustrate the challenges utilities are facing to manage extensive networks of infrastructure to deliver a consistent and high level of service. For buried infrastructure such as water systems, studies suggest that preventative interventions can lead to lower maintenance costs and fewer service disruptions (Mazumder et al, 2018; Li et al, 2014). The demonstrated need and benefit of appropriately applied asset management is juxtaposed against the relatively sparse literature that evaluates water systems within an asset management construct. Since 2020, just 37 papers specifically reference asset management in the Journal of Water Resources Planning and Management. Of those, only a few specifically look to develop strategies for improved asset management. Therefore, we highlight four key research areas that represent opportunities for advancement of asset management research for water systems. First, advances in condition assessment and forecasting are needed to better estimate asset deterioration using diverse datasets. Second, machine learning (ML) and artificial intelligence (AI) hold promise for predictive maintenance and investment prioritization, though questions of generalizability and model transparency remain. Third, applying a value of information framework can guide utilities in making cost-effective sensor deployment and data collection decisions, to direct monitoring strategies towards data-informed asset management decisions. Finally, integrated infrastructure management is critical, requiring coordinated planning with other infrastructure systems and stakeholder engagement to reduce costs and enhance service delivery.

Chini, Christopher M.↗

AmeriFlux FLUXNET-1F CA-SCC Scotty Creek Landscape

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site CA-SCC Scotty Creek Landscape. This is the FLUXNET version of the carbon flux data for the site CA-SCC Scotty Creek Landscape produced by applying the standard ONEFlux (1F) software. Site Description - The Scotty Creek flux tower is located in an organic-rich boreal forest-wetland landscape about 50 km south of Fort Simpson in the Taiga Plains of the Mackenzie watershed. The tower was installed in 2013 and operates an open-path EC system year-round running on solar power only. Flux footprints contain about 50 % forested peat plateaus and 50 % wetlands (i.e., collapse-scar bogs). The forests are underlain by permafrost, while the treeless wetlands are permafrost-free. The tower itself is located on a forested peat plateau. Black spruce tree density on plateaus is sparse and the mean canopy height is ca. 5 m.

Sonnentag, Oliver↗

High-resolution national mapping of natural gas composition substantially updates methane leakage impacts

Methane is emitted from oil and gas operations alongside heavier hydrocarbons and non-hydrocarbon gases, shaping emissions management decision-making, including air quality impacts. Yet, most assessments assume fixed gas composition, overlooking significant spatial and temporal variations. Here, we generate a high-resolution, data-driven map of natural gas composition across the United States, reconstructing methane, heavier hydrocarbons, and non-hydrocarbon species using spatio-temporal interpolation and oil-and-gas production patterns. Our approach is able to reduce composition prediction errors by 39% in terms of Mean Absolute Error (MAE) compared to standard techniques and reveals that methane loss rates have been underestimated by more than 50% in some regions. Beyond methane, we uncover substantial variability in co-emitted gases, exposing blind spots in current emissions inventories and emissions management frameworks. Our work enables more accurate emissions assessments, guides targeted measurement strategies, and informs emissions management decision-making. It also provides a general framework for prediction in environmental applications that integrate sparse measurements with auxiliary variables.

03 NATURAL GAS↗