Search NASA⌕ Search

SEARCH · Search NASA

Results for “data scarce”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Streamflow simulation in data-scarce basins using Bayesian and physics-informed machine learning models

Hydrologic predictions at rural watersheds are important but also challenging due to data shortage. Long short-term memory (LSTM) networks are a promising machine learning approach and have demonstrated good performance in streamflow predictions. However, due to its data-hungry nature, most LSTM applications focus on well-monitored catchments with abundant and high-quality observations. In this work, we investigate predictive capabilities of LSTM in poorly monitored watersheds with short observation records. To address three main challenges of LSTM applications in data-scarce locations, i.e., overfitting, uncertainty quantification (UQ), and out-of-distribution prediction, we evaluate different regularization techniques to prevent overfitting, apply a Bayesian LSTM for UQ, and introduce a physics-informed hybrid LSTM to enhance out-of-distribution prediction. Through case studies in two diverse sets of catchments with and without snow influence, we demonstrate that 1) when hydrologic variability in the prediction period is similar to the calibration period, LSTM models can reasonably predict daily streamflow with Nash–Sutcliffe efficiency above 0.8, even with only 2 years of calibration data; 2) when the hydrologic variability in the prediction and calibration periods is dramatically different, LSTM alone does not predict well, but the hybrid model can improve the out-of-distribution prediction with acceptable generalization accuracy; 3) L2 norm penalty and dropout can mitigate overfitting, and Bayesian and hybrid LSTM have no overfitting; and 4) Bayesian LSTM provides useful uncertainty information to improve prediction understanding and credibility. In conclusion, these insights have vital implications for streamflow simulation in watersheds where data quality and availability are a critical issue.

54 ENVIRONMENTAL SCIENCES↗

Data-scarce surrogate modeling of shock-induced pore collapse process

Understanding the mechanisms of shock-induced pore collapse is of great interest in various disciplines in sciences and engineering, including materials science, biological sciences, and geophysics. However, numerical modeling of the complex pore collapse processes can be costly. To this end, a strong need exists to develop surrogate models for generating economic predictions of pore collapse processes. Here, in this work, we study the use of a data-driven reduced-order model, namely dynamic mode decomposition, and a deep generative model, namely conditional generative adversarial networks, to resemble the numerical simulations of the pore collapse process at representative training shock pressures. Since the simulations are expensive, the training data are scarce, which makes training an accurate surrogate model challenging. To overcome the difficulties posed by the complex physics phenomena, we make several crucial treatments to the plain original form of the methods to increase the capability of approximating and predicting the dynamics. In particular, physics information is used as indicators or conditional inputs to guide the prediction. In realizing these methods, the training of each dynamic mode composition model takes only around 30 s on CPU. In contrast, training a generative adversarial network model takes 8 h on GPU. Moreover, using dynamic mode decomposition, the final-time relative error is around 0.3% in the reproductive cases. We also demonstrate the predictive power of the methods at unseen testing shock pressures, where the error ranges from 1.3 to 5% in the interpolatory cases and 8 to 9% in extrapolatory cases.

97 MATHEMATICS AND COMPUTING↗

Stochastic modeling and statistical calibration with model error and scarce data

This paper introduces a procedure to assess the predictive accuracy of stochastic models subject to model error and sparse data. Model error is introduced as uncertainty on the coefficients of appropriate polynomial chaos expansions (PCE). The error associated with finite sample size allows us to conceive of these coefficients as statistics of the data that we describe as random variables whose influence on output quantities of interest is evaluated through the extended polynomial chaos expansion (EPCE). A Bayesian data assimilation scheme is introduced to update these expansions by considering the resulting nested chaos expansion as a hierarchical probabilistic model. Stochastic models of quantities of interest (QoI) are thus constructed and efficiently evaluated. Here, the Metropolis–Hastings Markov chain Monte Carlo procedure is used to sample the posterior. Two illustrative analytical and numerical problems are used to demonstrate the proposed approach.

Bayesian inference↗

Projection-based multifidelity linear regression for data-scarce applications

Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. Here, the proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2% – 12% improvement in median accuracy and a higher R 2 score relative to single-fidelity methods at comparable computational cost.

data augmentation↗

Hydrologic applicability of satellite-based precipitation estimates for irrigation water management in the data-scarce region

Reliable precipitation estimates are crucial for planning and managing water resources, monitoring hydrologic extremes, and fulfilling irrigation water requirements. Accurate precipitation estimates are particularly challenging in complex mountain terrains, where monitoring gauges are often sparsely distributed due to their remote locations, and high installation and long-term operation costs. Recent advances in satellite-based precipitation estimates offer promising opportunities to improve our understanding of hydrologic processes and their applications for irrigation water management. Several datasets are available varying considerably in terms of their data sources, quality control methods, estimation procedure, and spatiotemporal resolutions. Choosing the most suitable dataset for a particular application is a complex task. In this study, we (1) evaluate the performance of six satellite-based precipitation estimates (SPEs): i) CHIRPS v2.0, ii) CMORPH v1.0, iii) ERA5, iv) IMERG v6, v) MSWEP v2.8, and vi) PERSIANN-CDR against the gauge precipitation using continuous statistical and categorical indices, (2) integrate SPEs with a calibrated semi-distributed hydrologic model to predict streamflow, and (3) demonstrate practical implications of improved streamflow prediction for irrigation water management in the central Himalayan region, Nepal. Our results illustrate that satellite-based precipitation estimates have competitive performance in capturing a wide range of rainfall characteristics, with demonstrated variability across river basins and time scales. Further, there are no significant discrepancies observed in satellite-based precipitation estimates for estimating irrigation water requirements for the three major crops (maize, wheat, and paddy) during the cropping period across the selected river basins, showing a greater promise for irrigation water management planning and decision making.

54 ENVIRONMENTAL SCIENCES↗

Wildfire-Power Grid Interactions: Feedback, Impacts, Monitoring, Modeling, and Mitigation Strategies

Wildfires are increasingly interacting with electric power systems through a two-way hazard chain: fires damage grid assets and trigger cascading outages, while grid faults can ignite new fires under hot, dry, and windy conditions. This review synthesizes the state of knowledge across five domains: (i) physical impacts of flames, heat, and smoke on lines, towers, insulators, and substations; (ii) power-infrastructure-initiated ignitions via conductor clash, high-impedance faults, and corona discharge; (iii) widespread blackouts and disproportionate societal impacts; (iv) multi-scale monitoring spanning laboratory tests, in-situ and grid-integrated sensors, and Earth observation; (v) coupled modeling that links fire behavior with grid operations; and (vi) technological and strategic mitigation pathways spanning prevention, response, and recovery. We integrate these domains into a novel 'feedback-aware' socio-technical framework. Through a longitudinal analysis (2005-2025) of global incidents, we identify that while vegetation contact remains the most frequent ignition source, aging infrastructure failure has emerged as a critical driver of catastrophic 'mega-fires'. We further identify persistent gaps, including limited interoperability of high-frequency grid and environmental data, scarce real-time data assimilation, and under-developed equity metrics for outage management. We conclude by outlining a research agenda to (1) deploy interoperable sensing architectures, (2) advance feedback-coupled fire-grid simulations, and (3) evaluate mitigation portfolios through techno-economic and fairness lenses. Recognizing wildfire-grid interactions as coupled socio-technical systems is essential for protecting infrastructure and communities and for ensuring reliable, sustainable electricity in a changing world.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Weakly Supervised Event Classification Using Imperfect Real-world PMU Data with Scarce Labels

This paper studies event classification using imperfect real-world phasor measurement unit (PMU) data with scarce event types (labels). By investigating the real-world PMU data, it is observed that most real-world PMU data's event type is unknown, which makes it challenging to directly use such dataset to build event classifiers as existing classification techniques require high-quality training data with known event type (i.e., label). To address this challenge, a weakly supervised learning based event classification approach is developed, which can use noisy and low-quality PMU data for the training. First, data quality issues are fixed using data preprocessing techniques and then event features are constructed from the PMU data. Using these features, a series of labeling functions are learnt to generate initial estimates of the labels of large amounts of unlabeled PMU data. As the labeling functions are learnt using the same data with scarce labels, the label estimates from the labeling functions can be correlated, noisy, and bias. To enhance these initial estimates, a generative model is developed to characterize the dependencies among the estimated labels, based on which better labels are obtained for training event classifiers. Numerical experiments using the real-world dataset from the Western Interconnection of the U.S. power transmission grid show that the proposed weakly supervised event classifier trained using the dataset with only 5% labeled data can achieve 78.4% classification accuracy.

Liu, Yunchuan↗

62°S Witnesses the Transition of Boundary Layer Marine Aerosol Pattern Over the Southern Ocean (50°S–68°S, 63°E– 150°E) During the Spring and Summer: Results From MARCUS (I)

The Atmospheric Radiation Measurement Mobile Facility-2 was installed onboard the research vessel Aurora Australis to measure aerosol properties during the 2017-2018 Measurement of Aerosols, Radiation, and CloUds over the pristine Southern ocean (MARCUS) Experiment, providing unique data on aerosols latitudinal and seasonal variation, including south of 60 degrees S where previous observations are scarce. Data from a Cloud Condensation Nuclei (CCN) counter and Ultra-High-Sensitivity Aerosol Spectrometer show that both the number concentration (N-CCN) and size distribution of CCN-active aerosols, with diameters (D) between 60 nm < D < 1,000 nm are different over the North Southern Ocean (NSO) (50 degrees S-60 degrees S) and the South Southern Ocean (SSO) (62 degrees S-68 degrees S). The average NSO N-CCN at 0.2% and 0.5% supersaturation were 28% and 49% less than that over the SSO. This increase of CCN over the SSO is caused by the increase of aerosols with 60 nm < D < 200 nm, consistent with calculations of Aerosol Scattering Angstrom Exponents derived from a nephelometer. Aerosol hygroscopicity growth factor measured by the Hygroscopic Tandem Differential Mobility Analyzer stayed close to 1.41 for aerosols with 50 nm < D < 250 nm over the SSO, but increased from 1.30 to 1.67 over the NSO, indicating different chemical compositions. Both CCN and Ice Nucleating Particles (INPs) showed a stronger variation with season than with latitude. The variation of heat-labile and presumably proteinacous INPs suggests an increase of ice nucleating-active microbes in summer.

Niu, Qing↗

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES↗

A comparative study of multimodal data fusion strategies for planetary spectroscopy

Integrating heterogeneous data sources can improve scientific inference when different modalities capture complementary information, but doing so is challenging in high-dimensional, small-sample settings. In spectroscopy for planetary exploration, Laser-Induced Breakdown Spectroscopy (LIBS), Raman Spectroscopy (Raman), Visible Infrared Spectroscopy (VISIR), and Mid-Infrared Spectroscopy (MIR) each examine different aspects of composition and mineralogy, raising fundamental questions about when and how data fusion improves predictive performance. Using a Mars-relevant set of geologic standards with measurements from all four modalities, we present a rigorous systematic evaluation of four data fusion strategies: low-level (data) fusion, mid-level (feature) fusion, high-level (decision) fusion, and residual-boosting (sequential) fusion. We assess performance in predicting oxide composition via nested cross-validation and corrected significance testing to evaluate whether data fusion improves upon single-modality baselines. We show that data fusion does not uniformly improve accuracy, and that observed gains are modest, oxide-dependent, and sensitive to modality and model structure. To move beyond aggregate accuracy metrics, we use model coefficients, permutation importance, and residual gain analysis to examine how the fusion models weight individual modalities and to identify patterns of apparent complementarity or redundancy. Though focused on spectroscopy for planetary exploration, our framework for data fusion evaluation and interpretation extends to other scientific domains with heterogeneous and scarce data and provides a principled approach evaluating data fusion strategies, interpreting modality contributions, and understanding tradeoffs among data fusion strategies.

97 MATHEMATICS AND COMPUTING↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

Hydrologic Regionalization under Data Scarcity: Implications for Streamflow Prediction

Continuous streamflow prediction is crucial in many applications of water resources planning and management. However, streamflow prediction is challenging, particularly in data-scarce regions. Here, we demonstrate an approach to regionalize the flow duration curve for predicting daily streamflow in the data-scare region of the central Himalayas. We developed a regression-based model to estimate streamflow at various segments of a flow duration curve by incorporating basin characteristics and climate variables. This study analyzes the sensitivities of proximity and characteristics between the donor (gauged) and receptor (ungauged) basins for time-series streamflow prediction. Our results show that regionalization techniques perform better in low to medium flows over high flows. Our findings are significant in the central Himalayan regional context to inform operational and management decisions in water sector projects like hydropower plants, which generally rely on low-to-medium streamflow information. Although the quantitative results are region-specific, the approach and insights are generalizable to the Himalayan region.

54 ENVIRONMENTAL SCIENCES↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

Designing alloys with process-mapping AI pre-trained on empirical knowledge

<span style="font-family: Calibri, sans-serif; font-size: 12pt;">Accelerated materials design should match the recent trends in the product development cycles. Materials data analytics can be used to significantly shorten development time of specialized alloys needed for next generation energy applications. However, it faces a challenge of scarce data available for training ML models. Incorporation of the domain knowledge into deep-learning graph structure via fuzzy pre-training and causal process imitation presents a viable approach to developing accurate data-driven models and reliable alloy design tools, with limited datasets. Artificial Intelligence (AI) was used in this study to incorporate such knowledge in the domain-specific computational tool, pyroMind. The tool provides not only novel design ideas but also their interpretation via physics and engineering concepts.</span>

Romanov, Vyacheslav↗

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning

Modern medicine increasingly relies on multimodal data, ranging from clinical notes to imaging and genomics, to guide diagnosis and treatment. However, integrating these heterogeneous data sources in a principled and interpretable manner remains a major challenge. We present MODES (Multi-mOdal Disentangled Embedding Space), a representation fusion framework that explicitly separates shared and modality-specific factors of variation, offering a structured latent space for multimodal information that improves both prediction and interpretability. By leveraging pre-trained unimodal foundation models, MODES mitigates the dependency on extensive paired datasets, crucial in data-scarce clinical settings. We introduce a masking strategy that optimizes representation dimensionality by eliminating low-information dimensions, to achieve compact, information-rich representations. Our framework demonstrates superior performance in predicting diagnoses and phenotypes compared to unimodal and conventional fusion models. MODES also enables robust diagnostic inference in missing data scenarios, offering an opportunity toward interpretable and efficient multimodal diagnostics in personalized healthcare.

60 APPLIED LIFE SCIENCES↗