Search NASA⌕ Search

SEARCH · Search NASA

Results for “sparse data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

Stoichiometrically-informed symbolic regression for extracting chemical reaction mechanisms from data

A data-driven computational method is introduced to extract chemical reaction mechanisms from time series chemical concentration data. It is realized through the use of dynamic symbolic regression in which a sparse analytical form for a dynamical system is discoverable from the underlying data. We specifically develop the stoichiometrically-informed symbolic regression (SISR) method to address a standing challenge in complex chemical reaction networks: given a time-series dataset of concentrations of several components, what is the mechanism and the associated rate constants? SISR finds the optimal mechanism, kinetic equations and rate constants by combining differential optimization with a genetic optimization approach that searches a symbolic space of possible reaction mechanisms. Use of SISR in several paradigmatic examples spanning linear and nonlinear reaction schemes results in excellent agreement between true and predicted mechanisms, including when the method is applied to noisy data. The advantages of a stoichiometrically-informed approach such as SISR to address reaction discovery is illustrated through comparison with the use of generic state-of-the-art data-driven approaches.

36 MATERIALS SCIENCE↗

X-ray Computed Tomography Data of Dense Metallic Components

The data shared in here are X-ray computed tomography (XCT) scans of a hexagonal fuel nozzle in 3 sections with the Metrotom 800 system at the Manufacturing Demonstration Facility (MDF) at Oak Ridge National Laboratory. The data are used in the paper "Tomographic Sparse View Selection using the View Covariance Loss, by Lin et al. (doi:10.1109/TPAMI.2025.36000720), accepted to the international conference on computational imaging (ICCP 2025). Figures 4-7 in the paper describe the part/XCT scan. File name Descriptions: Bottom section: TCR- Single Channeled SRC L 2019-3-18 12-26-41.hdf5 Medium section: TCR- Single Channeled SRC M 2019-3-18 13-8-9.hdf5 Top section: TCR- Single Channeled SRC T 2019-3-18 13-45-39.hdf5 Each hdf5 file contains projection data, and all the relevant X-ray CT scan setting. The full list of included attributes: distance_unit: Units of all distances specified angle_unit : Units of the angles angles: Array of all angles used voxel_size_xy: Baseline recon (if any) has this voxel size in the in-plane direction voxel_size_z: Baseline recon (if any) has this voxel size in the cross-plane direction det_pixel_size_col: Size of the detector pixels in the column dimension det_pixel_size_row: Size of the detector pixels in the row dimension src_iso_dist: Source to iso-center distance iso_det_dist: Iso-center to detector distance det_angle: If the detector is rotated/tilted, this angle corresponds to that value det_row_offset: Center of rotation offset in the vertical direction det_col_offset: Center of rotation offset in the horizontal direction reconstruction: A baseline reconstruction stored as 3D array BHC params: Beam-hardening parameters - Van De Casteel Model - if it has been used to pre-process the projections We also provided a python script (hdf_io.py) that allows the user to read the relevant data from each hdf5 file.

Ziabari, Amir [Oak Ridge National Laboratory]↗

Stellar Laboratories: New GeV and Ge VI Oscillator Strengths and their Validation in the Hot White Dwarf RE0503-289

State-of-the-art spectral analysis of hot stars by means of non-LTE model-atmosphere techniques has arrived at a high level of sophistication. The analysis of high-resolution and high-S/N spectra, however, is strongly restricted by the lack of reliable atomic data for highly ionized species from intermediate-mass metals to trans-iron elements. Especially data for the latter has only been sparsely calculated. Many of their lines are identified in spectra of extremely hot, hydrogen-deficient post-AGB stars. A reliable determination of their abundances establishes crucial constraints for AGB nucleosynthesis simulations and, thus, for stellar evolutionary theory. Aims. In a previous analysis of the UV spectrum of RE 0503-289, spectral lines of highly ionized Ga, Ge, As, Se, Kr, Mo, Sn, Te, I, and Xe were identified. Individual abundance determinations are hampered by the lack of reliable oscillator strengths. Most of these identified lines stem from Ge V. In addition, we identified Ge VI lines for the first time. We calculated Ge V and Ge VI oscillator strengths in order to reproduce the observed spectrum. Methods. We newly calculated Ge V and Ge VI oscillator strengths to consider their radiative and collisional bound-bound transitions in detail in our non-LTE stellar-atmosphere models for the analysis of the Ge IV-VI spectrum exhibited in high-resolution and high-S/N FUV (FUSE) and UV (ORFEUS/BEFS, IUE) observations of RE 0503-289. Results. In the UV spectrum of RE 0503-289, we identify four Ge IV, 37 Ge V, and seven Ge VI lines. Most of these lines are identified for the first time in any star. We can reproduce almost all Ge IV, GeV, and Ge VI lines in the observed spectrum of RE 0503-289 (T(sub eff) = 70 kK, log g = 7.5) at log Ge = -3.8 +/- 0.3 (mass fraction, about 650 times solar). The Ge IV/V/VI ionization equilibrium, that is a very sensitive T(sub eff) indicator, is reproduced well. Conclusions. Reliable measurements and calculations of atomic data are a prerequisite for stellar-atmosphere modeling. Our oscillator-strength calculations have allowed, for the first time, Ge V and Ge VI lines to be successfully reproduced in a white dwarf s (RE 0503-289) spectrum and to determine its photospheric Ge abundance.

Rauch, T.↗

Sparse Superpixel Unmixing for Exploratory Analysis of CRISM Hyperspectral Images

Fast automated analysis of hyperspectral imagery can inform observation planning and tactical decisions during planetary exploration. Products such as mineralogical maps can focus analysts' attention on areas of interest and assist data mining in large hyperspectral catalogs. In this work, sparse spectral unmixing drafts mineral abundance maps with Compact Reconnaissance Imaging Spectrometer (CRISM) images from the Mars Reconnaissance Orbiter. We demonstrate a novel "superpixel" segmentation strategy enabling efficient unmixing in an interactive session. Tests correlate automatic unmixing results based on redundant spectral libraries against hand-tuned summary products currently in use by CRISM researchers.

Image Segmentation↗

Latent Twins

Over the past decade, scientific machine learning has transformed the development of mathematical and computational frameworks for analyzing, modeling, and predicting complex systems. From inverse problems to numerical partial differential equations (PDEs), dynamical systems, and model reduction, these advances have pushed the boundaries of what can be simulated. Yet they have often progressed in parallel, with representation learning and algorithmic solution methods evolving largely as separate pipelines. With Latent Twins, we propose a unifying mathematical framework that creates a hidden surrogate in latent space for the underlying equations. Whereas digital twins mirror physical systems in the digital world, Latent Twins mirror mathematical systems in a learned latent space governed by operators. Through this lens, classical modeling, inversion, model reduction, and operator approximation all emerge as special cases of a single principle. We establish the fundamental approximation properties of Latent Twins for both ordinary differential equations (ODEs) and PDEs and demonstrate the framework across three representative settings: (i) canonical ODEs, capturing diverse dynamical regimes; (ii) a PDE benchmark using the shallow-water equations, contrasting Latent Twin simulations with deep operator network and forecasts with a four-dimensional variational method baseline; and (iii) a challenging real-data geopotential reanalysis dataset, reconstructing and forecasting from sparse, noisy observations. Latent Twins provide a compact, interpretable surrogate for solution operators that evaluate across arbitrary time gaps in a single-shot, while remaining compatible with scientific pipelines such as assimilation, control, and uncertainty quantification. Looking forward, this framework offers scalable, theory-grounded surrogates that bridge data-driven representation learning and classical scientific modeling across disciplines.

Latent Twins↗

Efficient Implementation of an Optimal Interpolator for Large Spatial Data Sets

Interpolating scattered data points is a problem of wide ranging interest. A number of approaches for interpolation have been proposed both from theoretical domains such as computational geometry and in applications' fields such as geostatistics. Our motivation arises from geological and mining applications. In many instances data can be costly to compute and are available only at nonuniformly scattered positions. Because of the high cost of collecting measurements, high accuracy is required in the interpolants. One of the most popular interpolation methods in this field is called ordinary kriging. It is popular because it is a best linear unbiased estimator. The price for its statistical optimality is that the estimator is computationally very expensive. This is because the value of each interpolant is given by the solution of a large dense linear system. In practice, kriging problems have been solved approximately by restricting the domain to a small local neighborhood of points that lie near the query point. Determining the proper size for this neighborhood is a solved by ad hoc methods, and it has been shown that this approach leads to undesirable discontinuities in the interpolant. Recently a more principled approach to approximating kriging has been proposed based on a technique called covariance tapering. This process achieves its efficiency by replacing the large dense kriging system with a much sparser linear system. This technique has been applied to a restriction of our problem, called simple kriging, which is not unbiased for general data sets. In this paper we generalize these results by showing how to apply covariance tapering to the more general problem of ordinary kriging. Through experimentation we demonstrate the space and time efficiency and accuracy of approximating ordinary kriging through the use of covariance tapering combined with iterative methods for solving large sparse systems. We demonstrate our approach on large data sizes arising both from synthetic sources and from real applications.

Memarsadeghi, Nargess↗

Improved microgrid resiliency through distributionally robust optimization under a policy-mode framework

Critical energy infrastructure are constantly under stress due to the ever increasing disruptions caused by wildfires, hurricanes, other weather related extreme events and cyber-attacks. Hence it becomes important to make critical infrastructure resilient to threats from such cyber-physical events. However, such events are hard to predict and numerous in nature and type and it becomes infeasible to make a system resilient to every possible such cyber-physical event. Such an approach can make the system operation overly conservative and impractical to operate. Furthermore, distributions of such events are hard to predict and historical data available on such events can be very sparse, making the problem even harder to solve. To deal with these issues, in this paper we present a policy-mode framework that enumerates and predicts the probability of various cyber-physical events and then a distributionally robust optimization (DRO) formulation that is robust to the sparsity of the available historical data. The proposed algorithm is illustrated on an islanded microgrid example: a modified IEEE 123-node feeder with distributed energy resources (DERs) and energy storage. Simulations are carried to validate the resiliency metrics under the sampled disruption events.

Nazir, Mohammad Nawaf↗

Determination of the Venus flyby orbits of the Soviet Vega probes using VLBI techniques

In December 1984, the Soviet Union launched two identical Vega spacecraft with the dual objective of exploring Venus and continuing to rendezvous with the comet Halley. The two Vega spacecraft encountered Venus in mid-June 1985 and successfully deployed entry probes and wind-measuring balloons into the Venus atmosphere. An objective of the Venus Balloon experiment was to measure the Venus winds using differential VLBI from the balloon and the flyby bus. NASA's Deep Space 64 meter subnet was part of a world wide network organized to collect data from the Vega probes and balloons. A critical element of this experiment was an accurate determination of the Venus relative flyby orbits of the Vega spacecraft during the 46 hour balloon lifetime. Venus flyby solutions were independently determined by the Soviets using two-way range and Doppler from Soviet stations and by JPL using one-way Doppler and VLBI data collected from the DSN. The Vega flyby solutions determined by the Soviets using a sparse two-way tracking strategy with JPL solutions using the DSN VLBI data to complement the Soviet data and with solutions using only one-way data collected by the DSN were compared.

Ellis, J.↗

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES↗

Early Inference of Nuclear Technology-Directed Research Activities of Authors from Scientific Publications

Nuclear research articles can provide information about early nuclear proliferation indicators such as influential research entities and technology capability levels of a country, but detection of nuclear activities typically occurs after they have started. We investigate the extent to which nuclear research articles can be used to infer whether a research entity will acquire or develop a nuclear technology before it happens. Early detection of nuclear proliferation or technology development indicators from data is challenging due to partial observability, sparse and unlabeled information, and confounding signals from multiple concurrent activities. This paper presents the early detection problem as a sequential decision-making, goal inference problem, where the objective is to characterize and predict an individual’s, organization’s, or a country’s intent (unobserved goal-directed behavior) towards developing a nuclear capability from partially observed sequences of their research publications, using inverse reinforcement learning and Bayesian goal inference methods. A computational framework is presented, and its application demonstrated using 29,196 Scopus records for a case study related to a civil nuclear capability. The case study results serve as a proof-of-concept demonstration for inference of technology-directed research activity of authors who publish in the nuclear domain. The inference method, combined with advanced computing, may be used to assess and monitor activities pertaining to early developmental stages of a nuclear technology or capability, which in turn can help to identify and prioritize activities with nuclear proliferation potential for further investigation.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Synoptic studies of chromospheric variability in F - K dwarfs with the IUE

Time-sequential series of IUE spectra for ten F, G and K dwarfs were obtained in 1980 and 1981 to study the rotational dependence of chromospheric flux in the ultraviolet. An interactive computational method using unbiased estimators was developed to measure emission line fluxes free of arbitrary judgement concerning the behavior of the underlying spectrum and shapes of the line profiles. Due to the limited number of observational samples per star, we have used special techniques to analyze the sparsely and anharmonically sampled emission line flux data. Two different autocorrelation measures were computed for each emission line as a function of temporal frequency. Examples and results of this analysis now in progress are given for several stars.

Hallam, K. L.↗

Effects of partitioning and scheduling sparse matrix factorization on communication and load balance

A block based, automatic partitioning and scheduling methodology is presented for sparse matrix factorization on distributed memory systems. Using experimental results, this technique is analyzed for communication and load imbalance overhead. To study the performance effects, these overheads were compared with those obtained from a straightforward 'wrap mapped' column assignment scheme. All experimental results were obtained using test sparse matrices from the Harwell-Boeing data set. The results show that there is a communication and load balance tradeoff. The block based method results in lower communication cost whereas the wrap mapped scheme gives better load balance.

Venugopal, Sesh↗

Impact of Quikscat Data on Numerical Weather Prediction

Scatterometer observations of the ocean surface wind speed and direction improve the depiction and prediction of storms at sea. These data are especially valuable where observations are otherwise sparse ---mostly in the Southern Hemisphere and tropics, but also on occasion in the North Atlantic and North Pacific. The SeaWinds scatterometer on the QuikScat satellite was launched in July 1999 and it represents a dramatic departure in design from the other scatterometer instruments launched during the past decade (ERS-1,2 and NSCAT). The NASA Data Assimilation Office (DAO) was the first data assimilation center to assimilate QuikScat SeaWinds data and evaluate their impact on numerical weather prediction. Several data impact experiments have been performed, using systems from both the DAO (GEOS-3) and from NCEP (GDAS). In general, these experiments have shown a modest impact of SeaWinds data on numerical weather prediction, the magnitude of which appears to be comparable to the magnitude of the impact of AMI scatterometer data from the ERS satellites. Some of the main results from these experiments will be presented at the meeting.

Atlas, Robert↗

Long Term Monitoring of the Io Plasma Torus During the Galileo Encounter

In the fall of 1999, the Galileo spacecraft made four passes into the Io plasma torus, obtaining the best in situ measurements ever of the particle and field environment in this densest region of the Jovian magnetosphere. Supporting observations from the ground are vital for understanding the global and temporal context of the in situ observations. We conducted a three-month-long Io plasma torus monitoring campaign centered on the time of the Galileo plasma torus passes to support this aspect of the Galileo mission. The almost-daily plasma density and temperature measurements obtained from our campaign allow the much more sparse but also much more detailed Galileo data to be used to address the issues of the structure of the Io plasma torus, the stability mechanism of the Jovian magnetosphere, the transport of material from the source region near Io, and the nature and source of persistent longitudinal variations. Combining the ground-based monitoring data with the detailed in situ data offers the only possibility for answering some of the most fundamental questions about the nature of the Io plasma torus.

Brown, Michael E.↗

Winter QPF Sensitivities to Snow Parameterizations and Comparisons to NASA CloudSat Observations

Steady increases in computing power have allowed for numerical weather prediction models to be initialized and run at high spatial resolution, permitting a transition from larger scale parameterizations of the effects of clouds and precipitation to the simulation of specific microphysical processes and hydrometeor size distributions. Although still relatively coarse in comparison to true cloud resolving models, these high resolution forecasts (on the order of 4 km or less) have demonstrated value in the prediction of severe storm mode and evolution and are being explored for use in winter weather events . Several single-moment bulk water microphysics schemes are available within the latest release of the Weather Research and Forecast (WRF) model suite, including the NASA Goddard Cumulus Ensemble, which incorporate some assumptions in the size distribution of a small number of hydrometeor classes in order to predict their evolution, advection and precipitation within the forecast domain. Although many of these schemes produce similar forecasts of events on the synoptic scale, there are often significant details regarding precipitation and cloud cover, as well as the distribution of water mass among the constituent hydrometeor classes. Unfortunately, validating data for cloud resolving model simulations are sparse. Field campaigns require in-cloud measurements of hydrometeors from aircraft in coordination with extensive and coincident ground based measurements. Radar remote sensing is utilized to detect the spatial coverage and structure of precipitation. Here, two radar systems characterize the structure of winter precipitation for comparison to equivalent features within a forecast model: a 3 GHz, Weather Surveillance Radar-1988 Doppler (WSR-88D) based in Omaha, Nebraska, and the 94 GHz NASA CloudSat Cloud Profiling Radar, a spaceborne instrument and member of the afternoon or "A-Train" of polar orbiting satellites tasked with cataloguing global cloud characteristics. Each system provides a unique perspective. The WSR-88D operates in a surveillance mode, sampling cloud volumes of Rayleigh scatterers where reflectivity is proportional to the sixth moment of the size distribution of equivalent spheres. The CloudSat radar provides enhanced sensitivity to smaller cloud ice crystals aloft, as well as consistent vertical profiles along each orbit. However, CloudSat reflectivity signatures are complicated somewhat by resonant Mie scattering effects and significant attenuation in the presence of cloud or rain water. Here, both radar systems are applied to a case of light to moderate snowfall within the warm frontal zone of a cold season, synoptic scale storm. Radars allow for an evaluation of the accuracy of a single-moment scheme in replicating precipitation structures, based on the bulk statistical properties of precipitation as suggested by reflectivity signatures.

Molthan, Andrew↗

An Investigation of Widespread Ozone Damage to the Soybean Crop in the Upper Midwest Determined From Ground-Based and Satellite Measurements

Elevated concentrations of ground-level ozone (O3) are frequently measured over farmland regions in many parts of the world. While numerous experimental studies show that O3 can significantly decrease crop productivity, independent verifications of yield losses at current ambient O3 concentrations in rural locations are sparse. In this study, soybean crop yield data during a 5-year period over the Midwest of the United States were combined with ground and satellite O3 measurements to provide evidence that yield losses on the order of 10% could be estimated through the use of a multiple linear regression model. Yield loss trends based on both conventional ground-based instrumentation and satellite-derived tropospheric O3 measurements were statistically significant and were consistent with results obtained from open-top chamber experiments and an open-air experimental facility (SoyFACE, Soybean Free Air Concentration Enrichment) in central Illinois. Our analysis suggests that such losses are a relatively new phenomenon due to the increase in background tropospheric O3 levels over recent decades. Extrapolation of these findings supports previous studies that estimate the global economic loss to the farming community of more than $10 billion annually.

Fishman, Jack↗

Patterns of Canopy and Surface Layer Consumption in a Boreal Forest Fire from Repeat Airborne Lidar

Fire in the boreal region is the dominant agent of forest disturbance with direct impacts on ecosystem structure, carbon cycling, and global climate. Global and biome-scale impacts are mediated by burn severity, measured as loss of forest canopy and consumption of the soil organic layer. To date, knowledge of the spatial variability in burn severity has been limited by sparse field sampling and moderate resolution satellite data. Here, we used pre- and post-fire airborne lidar data to directly estimate changes in canopy vertical structure and surface elevation for a 2005 boreal forest fire on Alaskas Kenai Peninsula. We found that both canopy and surface losses were strongly linked to pre-fire species composition and exhibited important fine-scale spatial variability at sub-30m resolution. The fractional reduction in canopy volume ranged from 0.61 in lowland black spruce stands to 0.27 in mixed white spruce and broad leaf forest. Residual structure largely reflects standing dead trees, highlighting the influence of pre-fire forest structure on delayed carbon losses from above ground biomass, post-fire albedo, and variability in understory light environments. Median loss of surface elevation was highest in lowland black spruce stands (0.18 m) but much lower in mixed stands (0.02 m), consistent with differences in pre-fire organic layer accumulation. Spatially continuous depth-of-burn estimates from repeat lidar measurements provide novel information to constrain carbon emissions from the surface organic layer and may inform related research on post-fire successional trajectories. Spectral measures of burn severity from Landsat were correlated with canopy (r = 0.76) and surface (r = -0.71) removal in black spruce stands but captured less of the spatial variability in fire effects for mixed stands (canopy r = 0.56, surface r = -0.26), underscoring the difficulty in capturing fire effects in heterogeneous boreal forest landscapes using proxy measures of burn severity from Landsat.

Alonzo, Michael↗

Utilizing Multiple Datasets for Snow Cover Mapping

Snow-cover maps generated from surface data are based on direct measurements, however they are prone to interpolation errors where climate stations are sparsely distributed. Snow cover is clearly discernable using satellite-attained optical data because of the high albedo of snow, yet the surface is often obscured by cloud cover. Passive microwave (PM) data is unaffected by clouds, however, the snow-cover signature is significantly affected by melting snow and the microwaves may be transparent to thin snow (less than 3cm). Both optical and microwave sensors have problems discerning snow beneath forest canopies. This paper describes a method that combines ground and satellite data to produce a Multiple-Dataset Snow-Cover Product (MDSCP). Comparisons with current snow-cover products show that the MDSCP draws together the advantages of each of its component products while minimizing their potential errors. Improved estimates of the snow-covered area are derived through the addition of two snow-cover classes ("thin or patchy" and "high elevation" snow cover) and from the analysis of the climate station data within each class. The compatibility of this method for use with Moderate Resolution Imaging Spectroradiometer (MODIS) data, which will be available in 2000, is also discussed. With the assimilation of these data, the resolution of the MDSCP would be improved both spatially and temporally and the analysis would become completely automated.

Tait, Andrew B.↗