Search NASA⌕ Search

SEARCH · Search NASA

Results for “Scarce data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data-scarce surrogate modeling of shock-induced pore collapse process

Understanding the mechanisms of shock-induced pore collapse is of great interest in various disciplines in sciences and engineering, including materials science, biological sciences, and geophysics. However, numerical modeling of the complex pore collapse processes can be costly. To this end, a strong need exists to develop surrogate models for generating economic predictions of pore collapse processes. Here, in this work, we study the use of a data-driven reduced-order model, namely dynamic mode decomposition, and a deep generative model, namely conditional generative adversarial networks, to resemble the numerical simulations of the pore collapse process at representative training shock pressures. Since the simulations are expensive, the training data are scarce, which makes training an accurate surrogate model challenging. To overcome the difficulties posed by the complex physics phenomena, we make several crucial treatments to the plain original form of the methods to increase the capability of approximating and predicting the dynamics. In particular, physics information is used as indicators or conditional inputs to guide the prediction. In realizing these methods, the training of each dynamic mode composition model takes only around 30 s on CPU. In contrast, training a generative adversarial network model takes 8 h on GPU. Moreover, using dynamic mode decomposition, the final-time relative error is around 0.3% in the reproductive cases. We also demonstrate the predictive power of the methods at unseen testing shock pressures, where the error ranges from 1.3 to 5% in the interpolatory cases and 8 to 9% in extrapolatory cases.

97 MATHEMATICS AND COMPUTING↗

Stochastic modeling and statistical calibration with model error and scarce data

This paper introduces a procedure to assess the predictive accuracy of stochastic models subject to model error and sparse data. Model error is introduced as uncertainty on the coefficients of appropriate polynomial chaos expansions (PCE). The error associated with finite sample size allows us to conceive of these coefficients as statistics of the data that we describe as random variables whose influence on output quantities of interest is evaluated through the extended polynomial chaos expansion (EPCE). A Bayesian data assimilation scheme is introduced to update these expansions by considering the resulting nested chaos expansion as a hierarchical probabilistic model. Stochastic models of quantities of interest (QoI) are thus constructed and efficiently evaluated. Here, the Metropolis–Hastings Markov chain Monte Carlo procedure is used to sample the posterior. Two illustrative analytical and numerical problems are used to demonstrate the proposed approach.

Bayesian inference↗

Projection-based multifidelity linear regression for data-scarce applications

Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. Here, the proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2% – 12% improvement in median accuracy and a higher R 2 score relative to single-fidelity methods at comparable computational cost.

data augmentation↗

Hydrologic applicability of satellite-based precipitation estimates for irrigation water management in the data-scarce region

Reliable precipitation estimates are crucial for planning and managing water resources, monitoring hydrologic extremes, and fulfilling irrigation water requirements. Accurate precipitation estimates are particularly challenging in complex mountain terrains, where monitoring gauges are often sparsely distributed due to their remote locations, and high installation and long-term operation costs. Recent advances in satellite-based precipitation estimates offer promising opportunities to improve our understanding of hydrologic processes and their applications for irrigation water management. Several datasets are available varying considerably in terms of their data sources, quality control methods, estimation procedure, and spatiotemporal resolutions. Choosing the most suitable dataset for a particular application is a complex task. In this study, we (1) evaluate the performance of six satellite-based precipitation estimates (SPEs): i) CHIRPS v2.0, ii) CMORPH v1.0, iii) ERA5, iv) IMERG v6, v) MSWEP v2.8, and vi) PERSIANN-CDR against the gauge precipitation using continuous statistical and categorical indices, (2) integrate SPEs with a calibrated semi-distributed hydrologic model to predict streamflow, and (3) demonstrate practical implications of improved streamflow prediction for irrigation water management in the central Himalayan region, Nepal. Our results illustrate that satellite-based precipitation estimates have competitive performance in capturing a wide range of rainfall characteristics, with demonstrated variability across river basins and time scales. Further, there are no significant discrepancies observed in satellite-based precipitation estimates for estimating irrigation water requirements for the three major crops (maize, wheat, and paddy) during the cropping period across the selected river basins, showing a greater promise for irrigation water management planning and decision making.

54 ENVIRONMENTAL SCIENCES↗

Wildfire-Power Grid Interactions: Feedback, Impacts, Monitoring, Modeling, and Mitigation Strategies

Wildfires are increasingly interacting with electric power systems through a two-way hazard chain: fires damage grid assets and trigger cascading outages, while grid faults can ignite new fires under hot, dry, and windy conditions. This review synthesizes the state of knowledge across five domains: (i) physical impacts of flames, heat, and smoke on lines, towers, insulators, and substations; (ii) power-infrastructure-initiated ignitions via conductor clash, high-impedance faults, and corona discharge; (iii) widespread blackouts and disproportionate societal impacts; (iv) multi-scale monitoring spanning laboratory tests, in-situ and grid-integrated sensors, and Earth observation; (v) coupled modeling that links fire behavior with grid operations; and (vi) technological and strategic mitigation pathways spanning prevention, response, and recovery. We integrate these domains into a novel 'feedback-aware' socio-technical framework. Through a longitudinal analysis (2005-2025) of global incidents, we identify that while vegetation contact remains the most frequent ignition source, aging infrastructure failure has emerged as a critical driver of catastrophic 'mega-fires'. We further identify persistent gaps, including limited interoperability of high-frequency grid and environmental data, scarce real-time data assimilation, and under-developed equity metrics for outage management. We conclude by outlining a research agenda to (1) deploy interoperable sensing architectures, (2) advance feedback-coupled fire-grid simulations, and (3) evaluate mitigation portfolios through techno-economic and fairness lenses. Recognizing wildfire-grid interactions as coupled socio-technical systems is essential for protecting infrastructure and communities and for ensuring reliable, sustainable electricity in a changing world.

24 POWER TRANSMISSION AND DISTRIBUTION↗

62°S Witnesses the Transition of Boundary Layer Marine Aerosol Pattern Over the Southern Ocean (50°S–68°S, 63°E– 150°E) During the Spring and Summer: Results From MARCUS (I)

The Atmospheric Radiation Measurement Mobile Facility-2 was installed onboard the research vessel Aurora Australis to measure aerosol properties during the 2017-2018 Measurement of Aerosols, Radiation, and CloUds over the pristine Southern ocean (MARCUS) Experiment, providing unique data on aerosols latitudinal and seasonal variation, including south of 60 degrees S where previous observations are scarce. Data from a Cloud Condensation Nuclei (CCN) counter and Ultra-High-Sensitivity Aerosol Spectrometer show that both the number concentration (N-CCN) and size distribution of CCN-active aerosols, with diameters (D) between 60 nm < D < 1,000 nm are different over the North Southern Ocean (NSO) (50 degrees S-60 degrees S) and the South Southern Ocean (SSO) (62 degrees S-68 degrees S). The average NSO N-CCN at 0.2% and 0.5% supersaturation were 28% and 49% less than that over the SSO. This increase of CCN over the SSO is caused by the increase of aerosols with 60 nm < D < 200 nm, consistent with calculations of Aerosol Scattering Angstrom Exponents derived from a nephelometer. Aerosol hygroscopicity growth factor measured by the Hygroscopic Tandem Differential Mobility Analyzer stayed close to 1.41 for aerosols with 50 nm < D < 250 nm over the SSO, but increased from 1.30 to 1.67 over the NSO, indicating different chemical compositions. Both CCN and Ice Nucleating Particles (INPs) showed a stronger variation with season than with latitude. The variation of heat-labile and presumably proteinacous INPs suggests an increase of ice nucleating-active microbes in summer.

Niu, Qing↗

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES↗

A comparative study of multimodal data fusion strategies for planetary spectroscopy

Integrating heterogeneous data sources can improve scientific inference when different modalities capture complementary information, but doing so is challenging in high-dimensional, small-sample settings. In spectroscopy for planetary exploration, Laser-Induced Breakdown Spectroscopy (LIBS), Raman Spectroscopy (Raman), Visible Infrared Spectroscopy (VISIR), and Mid-Infrared Spectroscopy (MIR) each examine different aspects of composition and mineralogy, raising fundamental questions about when and how data fusion improves predictive performance. Using a Mars-relevant set of geologic standards with measurements from all four modalities, we present a rigorous systematic evaluation of four data fusion strategies: low-level (data) fusion, mid-level (feature) fusion, high-level (decision) fusion, and residual-boosting (sequential) fusion. We assess performance in predicting oxide composition via nested cross-validation and corrected significance testing to evaluate whether data fusion improves upon single-modality baselines. We show that data fusion does not uniformly improve accuracy, and that observed gains are modest, oxide-dependent, and sensitive to modality and model structure. To move beyond aggregate accuracy metrics, we use model coefficients, permutation importance, and residual gain analysis to examine how the fusion models weight individual modalities and to identify patterns of apparent complementarity or redundancy. Though focused on spectroscopy for planetary exploration, our framework for data fusion evaluation and interpretation extends to other scientific domains with heterogeneous and scarce data and provides a principled approach evaluating data fusion strategies, interpreting modality contributions, and understanding tradeoffs among data fusion strategies.

97 MATHEMATICS AND COMPUTING↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning

Modern medicine increasingly relies on multimodal data, ranging from clinical notes to imaging and genomics, to guide diagnosis and treatment. However, integrating these heterogeneous data sources in a principled and interpretable manner remains a major challenge. We present MODES (Multi-mOdal Disentangled Embedding Space), a representation fusion framework that explicitly separates shared and modality-specific factors of variation, offering a structured latent space for multimodal information that improves both prediction and interpretability. By leveraging pre-trained unimodal foundation models, MODES mitigates the dependency on extensive paired datasets, crucial in data-scarce clinical settings. We introduce a masking strategy that optimizes representation dimensionality by eliminating low-information dimensions, to achieve compact, information-rich representations. Our framework demonstrates superior performance in predicting diagnoses and phenotypes compared to unimodal and conventional fusion models. MODES also enables robust diagnostic inference in missing data scenarios, offering an opportunity toward interpretable and efficient multimodal diagnostics in personalized healthcare.

60 APPLIED LIFE SCIENCES↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations↗

Derivation of physical equations for high-speed laser welding using large language models

It is challenging to formulate complex physical phenomena that occur in a manufacturing process, particularly when the available data are limited, rendering conventional data-driven approaches ineffective. This study aims to predict humping onset in high-speed laser welding by introducing a novel framework, namely text-to-equations generative pre-trained transformer (T2EGPT). This method leverages the capabilities of large language models (LLMs), in combination with sparse experimental data and enriched literature data, to derive an interpretable and generalizable equation for predicting humping initiation. By capturing key correlations among physical parameters, T2EGPT generates a compact and dimensionless expression that accurately predicts hump formation. The equation reveals that humping arises from the interplay between inertia-driven backward melt flow and capillary-driven surface stabilization, where inertial forces drive molten metal backward and capillary forces resist surface deformation. Furthermore, compared to traditional data-driven models, T2EGPT demonstrates enhanced predictive accuracy and cross-material transferability. More broadly, this study highlights the potential of LLMs to integrate textual information with data-driven discovery, enabling the extraction of physical laws in data-scarce scientific domains.

36 MATERIALS SCIENCE↗