Search NASA⌕ Search

SEARCH · Search NASA

Results for “Scarce data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Remote Sensing-Driven Hydrodynamic Modeling in Data-Scarce Regions: Integrating ICESat-2, Sentinel-2, SWOT and Re-analysis Models for Coastal Monitoring

Hydrodynamic models in coastal and estuarine systems are typically constrained by sparse bathymetry, boundary, and validation data, especially in regions where field campaigns are costly or impractical. Here we develop and test a fully satellite-driven framework for hydrodynamic modeling in South Africa’s Langebaan Lagoon without using any local in situ measurements. Bathymetry is derived by training multispectral Sentinel-2 reflectance against ICESat-2 ATL24 photon-derived depths using an XGBoost model optimized with Bayesian search. The final satellite-derived bathymetry reproduces independent ATL24 points with RMSE = 0.45 m and R 2 = 0.97. This bathymetry was used in a depth-averaged Delft3D Flexible Mesh model driven at the open boundary by TPXO tidal harmonics and by ERA5 winds. We validate modeled water surface elevation against 16 SWOT low-rate (250 m, unsmoothed) passes in 2023. SWOT–model comparisons yield an overall RMSE of 0.11 m and R 2 = 0.61, with typical point differences <0.10 m (∼7% of the 1.5 m tidal range), and showed consistent spatial gradients in water level from the offshore boundary, through Saldanha Bay, and into the lagoon. At the offshore boundary, TPXO and SWOT sea surface heights agree closely (R 2 = 0.86). A simple phase adjustment of ∼26,min between TPXO and SWOT lowers the RMSE from 0.18,m to 0.11,m, showing that phase offset accounts for some of the discrepancy, with additional errors likely linked to non-tidal signals. Our results demonstrate that combining passive optical, photon-counting LiDAR, radar interferometry, and global tidal/atmospheric models enables robust, transferrable hydrodynamic modeling in data-scarce coastal systems, offering a cost-effective pathway for monitoring.

ICESat-2↗

Data-scarce surrogate modeling of shock-induced pore collapse process

Understanding the mechanisms of shock-induced pore collapse is of great interest in various disciplines in sciences and engineering, including materials science, biological sciences, and geophysics. However, numerical modeling of the complex pore collapse processes can be costly. To this end, a strong need exists to develop surrogate models for generating economic predictions of pore collapse processes. Here, in this work, we study the use of a data-driven reduced-order model, namely dynamic mode decomposition, and a deep generative model, namely conditional generative adversarial networks, to resemble the numerical simulations of the pore collapse process at representative training shock pressures. Since the simulations are expensive, the training data are scarce, which makes training an accurate surrogate model challenging. To overcome the difficulties posed by the complex physics phenomena, we make several crucial treatments to the plain original form of the methods to increase the capability of approximating and predicting the dynamics. In particular, physics information is used as indicators or conditional inputs to guide the prediction. In realizing these methods, the training of each dynamic mode composition model takes only around 30 s on CPU. In contrast, training a generative adversarial network model takes 8 h on GPU. Moreover, using dynamic mode decomposition, the final-time relative error is around 0.3% in the reproductive cases. We also demonstrate the predictive power of the methods at unseen testing shock pressures, where the error ranges from 1.3 to 5% in the interpolatory cases and 8 to 9% in extrapolatory cases.

97 MATHEMATICS AND COMPUTING↗

Stochastic modeling and statistical calibration with model error and scarce data

This paper introduces a procedure to assess the predictive accuracy of stochastic models subject to model error and sparse data. Model error is introduced as uncertainty on the coefficients of appropriate polynomial chaos expansions (PCE). The error associated with finite sample size allows us to conceive of these coefficients as statistics of the data that we describe as random variables whose influence on output quantities of interest is evaluated through the extended polynomial chaos expansion (EPCE). A Bayesian data assimilation scheme is introduced to update these expansions by considering the resulting nested chaos expansion as a hierarchical probabilistic model. Stochastic models of quantities of interest (QoI) are thus constructed and efficiently evaluated. Here, the Metropolis–Hastings Markov chain Monte Carlo procedure is used to sample the posterior. Two illustrative analytical and numerical problems are used to demonstrate the proposed approach.

Bayesian inference↗

Application of GRACE for Monitoring Groundwater in Data Scarce Regions

In the United States, groundwater storage is somewhat well monitored (spatial and temporal data gaps notwithstanding) and abundant data are freely and easily accessible. Outside of the U.S., groundwater often is not monitored systematically and where it is the data are rarely centralized and made available. Since 2002 the Gravity Recovery and Climate Experiment (GRACE) satellite mission has delivered gravity field observations which have been used to infer variations in total terrestrial water storage, including groundwater, at regional to continental scales. Challenges to using GRACE for groundwater monitoring include its relatively coarse spatial and temporal resolutions, its inability to differentiate groundwater from other types of water on and under the land surface, and typical 2-3 month data latency. Data assimilation can be used to overcome these challenges, but uncertainty in the results remains and is difficult to quantify without independent observations. Nevertheless, the results are preferable to the alternative - no data at all- and GRACE has already revealed groundwater variability and trends in regions where only anecdotal evidence existed previously.

Rodell, Matt↗

Projection-based multifidelity linear regression for data-scarce applications

Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. Here, the proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2% – 12% improvement in median accuracy and a higher R 2 score relative to single-fidelity methods at comparable computational cost.

data augmentation↗

Hydrologic applicability of satellite-based precipitation estimates for irrigation water management in the data-scarce region

Reliable precipitation estimates are crucial for planning and managing water resources, monitoring hydrologic extremes, and fulfilling irrigation water requirements. Accurate precipitation estimates are particularly challenging in complex mountain terrains, where monitoring gauges are often sparsely distributed due to their remote locations, and high installation and long-term operation costs. Recent advances in satellite-based precipitation estimates offer promising opportunities to improve our understanding of hydrologic processes and their applications for irrigation water management. Several datasets are available varying considerably in terms of their data sources, quality control methods, estimation procedure, and spatiotemporal resolutions. Choosing the most suitable dataset for a particular application is a complex task. In this study, we (1) evaluate the performance of six satellite-based precipitation estimates (SPEs): i) CHIRPS v2.0, ii) CMORPH v1.0, iii) ERA5, iv) IMERG v6, v) MSWEP v2.8, and vi) PERSIANN-CDR against the gauge precipitation using continuous statistical and categorical indices, (2) integrate SPEs with a calibrated semi-distributed hydrologic model to predict streamflow, and (3) demonstrate practical implications of improved streamflow prediction for irrigation water management in the central Himalayan region, Nepal. Our results illustrate that satellite-based precipitation estimates have competitive performance in capturing a wide range of rainfall characteristics, with demonstrated variability across river basins and time scales. Further, there are no significant discrepancies observed in satellite-based precipitation estimates for estimating irrigation water requirements for the three major crops (maize, wheat, and paddy) during the cropping period across the selected river basins, showing a greater promise for irrigation water management planning and decision making.

54 ENVIRONMENTAL SCIENCES↗

Wildfire-Power Grid Interactions: Feedback, Impacts, Monitoring, Modeling, and Mitigation Strategies

Wildfires are increasingly interacting with electric power systems through a two-way hazard chain: fires damage grid assets and trigger cascading outages, while grid faults can ignite new fires under hot, dry, and windy conditions. This review synthesizes the state of knowledge across five domains: (i) physical impacts of flames, heat, and smoke on lines, towers, insulators, and substations; (ii) power-infrastructure-initiated ignitions via conductor clash, high-impedance faults, and corona discharge; (iii) widespread blackouts and disproportionate societal impacts; (iv) multi-scale monitoring spanning laboratory tests, in-situ and grid-integrated sensors, and Earth observation; (v) coupled modeling that links fire behavior with grid operations; and (vi) technological and strategic mitigation pathways spanning prevention, response, and recovery. We integrate these domains into a novel 'feedback-aware' socio-technical framework. Through a longitudinal analysis (2005-2025) of global incidents, we identify that while vegetation contact remains the most frequent ignition source, aging infrastructure failure has emerged as a critical driver of catastrophic 'mega-fires'. We further identify persistent gaps, including limited interoperability of high-frequency grid and environmental data, scarce real-time data assimilation, and under-developed equity metrics for outage management. We conclude by outlining a research agenda to (1) deploy interoperable sensing architectures, (2) advance feedback-coupled fire-grid simulations, and (3) evaluate mitigation portfolios through techno-economic and fairness lenses. Recognizing wildfire-grid interactions as coupled socio-technical systems is essential for protecting infrastructure and communities and for ensuring reliable, sustainable electricity in a changing world.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An Experimental System for a Global Flood Prediction: From Satellite Precipitation Data to a Flood Inundation Map

Floods impact more people globally than any other type of natural disaster. It has been established by experience that the most effective means to reduce the property damage and life loss caused by floods is the development of flood early warning systems. However, advances for such a system have been constrained by the difficulty in estimating rainfall continuously over space (catchment-. national-, continental-. or even global-scale areas) and time (hourly to daily). Particularly, insufficient in situ data, long delay in data transmission and absence of real-time data sharing agreements in many trans-boundary basins hamper the development of a real-time system at the regional to global scale. In many countries around the world, particularly in the tropics where rainfall and flooding co-exist in abundance, satellite-based precipitation estimation may be the best source of rainfall data for those data scarce (ungauged) areas and trans-boundary basins. Satellite remote sensing data acquired and processed in real time can now provide the space-time information on rainfall fluxes needed to monitor severe flood events around the world. This can be achieved by integrating the satellite-derived forcing data with hydrological models, which can be parameterized by a tailored geospatial database. An example that is a key to this progress is NASA's contribution to the Tropical Rainfall Measuring Mission (TRMM), launched in November 1997. Hence, in an effort to evolve toward a more hydrologically-relevant flood alert system, this talk articulates a module-structured framework for quasi-global flood potential naming, that is 'up to date' with the state of the art on satellite rainfall estimation and the improved geospatial datasets. The system is modular in design with the flexibility that permits changes in the model structure and in the choice of components. Four major components included in the system are: 1) multi-satellite precipitation estimation; 2) characterization of land surface including digital elevation from NASA SRTM, topography-derived hydrologic parameters such as flood direction. flow accumulation, basin, and river network etc.; 3) spatially distributed hydrological models to infiltrate rainfall and route overland runoff; and 4) an implementation interface to relay thc input data to the models and display the flood inundation results to the users and decision-makers. Early results appear reasonable in terms of location and frequency of events. Case studies of this experimental system are evaluated with surface runoff data and other river monitoring systems. such as Dartmouth Flood Observatory's "Surface Water Watch" array of river reaches that are measured daily via other satellite remote sensing data. A major outcome of this progress will be the availability of a global overview of flood alerts that should consequently improve the performance of Decision Support System. We expect these developments in utilization of satellite remote sensing technology to offer a practical solution to the challenge of building a cost-effective early warning system for data scarce and under-developed areas.

Adler, Robert↗

62°S Witnesses the Transition of Boundary Layer Marine Aerosol Pattern Over the Southern Ocean (50°S–68°S, 63°E– 150°E) During the Spring and Summer: Results From MARCUS (I)

The Atmospheric Radiation Measurement Mobile Facility-2 was installed onboard the research vessel Aurora Australis to measure aerosol properties during the 2017-2018 Measurement of Aerosols, Radiation, and CloUds over the pristine Southern ocean (MARCUS) Experiment, providing unique data on aerosols latitudinal and seasonal variation, including south of 60 degrees S where previous observations are scarce. Data from a Cloud Condensation Nuclei (CCN) counter and Ultra-High-Sensitivity Aerosol Spectrometer show that both the number concentration (N-CCN) and size distribution of CCN-active aerosols, with diameters (D) between 60 nm < D < 1,000 nm are different over the North Southern Ocean (NSO) (50 degrees S-60 degrees S) and the South Southern Ocean (SSO) (62 degrees S-68 degrees S). The average NSO N-CCN at 0.2% and 0.5% supersaturation were 28% and 49% less than that over the SSO. This increase of CCN over the SSO is caused by the increase of aerosols with 60 nm < D < 200 nm, consistent with calculations of Aerosol Scattering Angstrom Exponents derived from a nephelometer. Aerosol hygroscopicity growth factor measured by the Hygroscopic Tandem Differential Mobility Analyzer stayed close to 1.41 for aerosols with 50 nm < D < 250 nm over the SSO, but increased from 1.30 to 1.67 over the NSO, indicating different chemical compositions. Both CCN and Ice Nucleating Particles (INPs) showed a stronger variation with season than with latitude. The variation of heat-labile and presumably proteinacous INPs suggests an increase of ice nucleating-active microbes in summer.

Niu, Qing↗

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES↗

A comparative study of multimodal data fusion strategies for planetary spectroscopy

Integrating heterogeneous data sources can improve scientific inference when different modalities capture complementary information, but doing so is challenging in high-dimensional, small-sample settings. In spectroscopy for planetary exploration, Laser-Induced Breakdown Spectroscopy (LIBS), Raman Spectroscopy (Raman), Visible Infrared Spectroscopy (VISIR), and Mid-Infrared Spectroscopy (MIR) each examine different aspects of composition and mineralogy, raising fundamental questions about when and how data fusion improves predictive performance. Using a Mars-relevant set of geologic standards with measurements from all four modalities, we present a rigorous systematic evaluation of four data fusion strategies: low-level (data) fusion, mid-level (feature) fusion, high-level (decision) fusion, and residual-boosting (sequential) fusion. We assess performance in predicting oxide composition via nested cross-validation and corrected significance testing to evaluate whether data fusion improves upon single-modality baselines. We show that data fusion does not uniformly improve accuracy, and that observed gains are modest, oxide-dependent, and sensitive to modality and model structure. To move beyond aggregate accuracy metrics, we use model coefficients, permutation importance, and residual gain analysis to examine how the fusion models weight individual modalities and to identify patterns of apparent complementarity or redundancy. Though focused on spectroscopy for planetary exploration, our framework for data fusion evaluation and interpretation extends to other scientific domains with heterogeneous and scarce data and provides a principled approach evaluating data fusion strategies, interpreting modality contributions, and understanding tradeoffs among data fusion strategies.

97 MATHEMATICS AND COMPUTING↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning

Modern medicine increasingly relies on multimodal data, ranging from clinical notes to imaging and genomics, to guide diagnosis and treatment. However, integrating these heterogeneous data sources in a principled and interpretable manner remains a major challenge. We present MODES (Multi-mOdal Disentangled Embedding Space), a representation fusion framework that explicitly separates shared and modality-specific factors of variation, offering a structured latent space for multimodal information that improves both prediction and interpretability. By leveraging pre-trained unimodal foundation models, MODES mitigates the dependency on extensive paired datasets, crucial in data-scarce clinical settings. We introduce a masking strategy that optimizes representation dimensionality by eliminating low-information dimensions, to achieve compact, information-rich representations. Our framework demonstrates superior performance in predicting diagnoses and phenotypes compared to unimodal and conventional fusion models. MODES also enables robust diagnostic inference in missing data scenarios, offering an opportunity toward interpretable and efficient multimodal diagnostics in personalized healthcare.

60 APPLIED LIFE SCIENCES↗