Search NASA⌕ Search

SEARCH · Search NASA

Results for “sparse data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Validation of the Pulmonary Function System for Use on the International Space Station

Aerobic deconditioning occurs during long duration space flight despite the use of exercise countermeasures (Convertino, 1996). As a part of International Space Station (ISS) medical operations, periodic tests designed to estimate aerobic capacity are performed to track changes in aerobic fitness and to determine the effectiveness of exercise countermeasures. These tests are performed prior to, during, and after missions of greater than 30 days in duration. Crewmembers selected for missions aboard the ISS perform a graded exercise test on a cycle ergometer approximately 270 days prior to their scheduled launch date in order to measure peak oxygen consumption (VO2PK) and peak heart rate (HRpk). Approximately 30 to 45 days prior to launch, crewmembers perform a submaximal cycle ergometer test at work rates set to elicit 25, 50 and 75% of their pre-flight VO2PK. This test, known as the Periodic Fitness Evaluation (PFE), serves as a baseline measure to which subsequent in-and post-flight exercise tests are compared. While onboard the ISS, crewmembers are normally scheduled to perform the PFE beginning with flight day (FD) 14 and every 30 days thereafter. The PFE is also conducted 5 and 30 days following flight. Using PFE data, aerobic fitness is estimated by quantifying the VO2 vs. HR relationship using linear regression and calculating the VO2 that would occur at the crewmember s previously measured HRpk. Currently, for data collected during flight, this technique assumes that the pre- vs. in-flight oxygen consumption per given cycle workload is similar. However, the validity of this assumption is based upon a sparse amount of data collected during the Skylab era (Michel, et al. 1977). The method of using heart rate and cycle ergometer work rates has been used to estimate aerobic fitness in normal gravity (Astrand and Ryhming, 1954; Lee, 1993). Due to spaceflight induced physiological alterations, such as shifts in extracellular fluid (e.g. plasma) volume, this method may not be valid during space flight. In addition, the ergometer onboard ISS is vibration-isolated and moves with the astronaut s application of force into the pedals. The effect of this movement on the VO2 of cycle exercise on ISS has not been quantified.

McCleary, Frank A.↗

Application of Spectral Analysis Techniques in the Intercomparison of Aerosol Data. Part II: Using Maximum Covariance Analysis to Effectively Compare Spatiotemporal Variability of Satellite and AERONET Measured Aerosol Optical Depth

Moderate Resolution Imaging SpectroRadiometer (MODIS) and Multi-angle Imaging Spectroradiomater (MISR) provide regular aerosol observations with global coverage. It is essential to examine the coherency between space- and ground-measured aerosol parameters in representing aerosol spatial and temporal variability, especially in the climate forcing and model validation context. In this paper, we introduce Maximum Covariance Analysis (MCA), also known as Singular Value Decomposition analysis as an effective way to compare correlated aerosol spatial and temporal patterns between satellite measurements and AERONET data. This technique not only successfully extracts the variability of major aerosol regimes but also allows the simultaneous examination of the aerosol variability both spatially and temporally. More importantly, it well accommodates the sparsely distributed AERONET data, for which other spectral decomposition methods, such as Principal Component Analysis, do not yield satisfactory results. The comparison shows overall good agreement between MODIS/MISR and AERONET AOD variability. The correlations between the first three modes of MCA results for both MODIS/AERONET and MISR/ AERONET are above 0.8 for the full data set and above 0.75 for the AOD anomaly data. The correlations between MODIS and MISR modes are also quite high (greater than 0.9). We also examine the extent of spatial agreement between satellite and AERONET AOD data at the selected stations. Some sites with disagreements in the MCA results, such as Kanpur, also have low spatial coherency. This should be associated partly with high AOD spatial variability and partly with uncertainties in satellite retrievals due to the seasonally varying aerosol types and surface properties.

aerosols↗

Application of Compressive Sensing to Gravitational Microlensing Data and Implications for Miniaturized Space Observatories

Compressive Sensing is a technique for simultaneous acquisition and compression of data that is sparse or can be made sparse in some domain. It is currently under intense development and has been profitably employed for industrial and medical applications. We here describe the use of this technique for the processing of astronomical data. We outline the procedure as applied to exoplanet gravitational microlensing and analyze measurement results and uncertainty values. We describe implications for on-spacecraft data processing for space observatories. Our findings suggest that application of these techniques may yield significant, enabling benefits especially for power and volume-limited space applications such as miniaturized or micro-constellation satellites.

Gravitational Microlensing↗

Application of Compressive Sensing to Gravitational Microlensing Data and Implications for Miniaturized Space Observatories

Compressive Sensing is a technique for simultaneous acquisition and compression of data that is sparse or can be made sparse in some domain. It is currently under intense development and has been profitably employed for industrial and medical applications. We here describe the use of this technique for the processing of astronomical data. We outline the procedure as applied to exoplanet gravitational microlensing and analyze measurement results and uncertainty values. We describe implications for on-spacecraft data processing for space observatories. Our findings suggest that application of these techniques may yield significant, enabling benefits especially for power and volume-limited space applications such as miniaturized or micro-constellation satellites.

Compressive Sensing↗

Predicting Drug Effects from High-dimensional Asymmetric Drug Data Sets using Graph Neural Networks: A Comprehensive Analysis of Multi-target Drug Effect Prediction

Graph neural networks (GNNs) have emerged as one of the most effective Machine learning (ML) techniques for drug effect prediction from drug molecular graphs. Despite having immense potential, GNN models lack performance when using data sets that contain high dimensional asymmetrically co-occurrent drug effects as targets with complex correlations between them. Training individual learning models for each drug effect and incorporating every prediction result for a wide spectrum of drug effects is beyond practicality. Such an implication provides a testbed to address this challenge as multi-target prediction problems, aiming to predict all drug effects at a time. We develop standard and hybrid graph neural networks (GNNs)to perform two separate tasks that are multi-regression for continuous values and multi-label classification for categorical values contained in our data sets. Since this step makes the target data even more sparse and introduces asymmetric label co-occurrence, the learning of multi-label classification models becomes difficult and heavily impacts the GNN's performance. To address these challenges, we propose a new data oversampling technique to improve multi-label classification performances on all the given imbalanced molecular graph data sets. Using the technique, we improve the data imbalance ratio of the drug effects better than before while protecting the data set's integrity. Finally, we evaluate multi-label classification performance using the best-performant hybrid GNN model on all the oversampled data sets obtained from the proposed oversampling technique. These results outperform those of other ML models including GNN models when they are trained on the original data sets or oversampled data sets using MLSMOTE (a well-known oversampling technique) in all evaluation metrics precision, recall, and F1 score by a significant margin.

Bose, Avishek [ORNL]↗

Using data tagging to improve the performance of Kanerva's sparse distributed memory

The standard formulation of Kanerva's sparse distributed memory (SDM) involves the selection of a large number of data storage locations, followed by averaging the data contained in those locations to reconstruct the stored data. A variant of this model is discussed, in which the predominant pattern is the focus of reconstruction. First, one architecture is proposed which returns the predominant pattern rather than the average pattern. However, this model will require too much storage for most uses. Next, a hybrid model is proposed, called tagged SDM, which approximates the results of the predominant pattern machine, but is nearly as efficient as Kanerva's original formulation. Finally, some experimental results are shown which confirm that significant improvements in the recall capability of SDM can be achieved using the tagged architecture.

Rogers, David↗

Classifying Forest Type in the National Forest Inventory Context with Airborne Hyperspectral and Lidar Data

Forest structure and composition regulate a range of ecosystem services, including biodiversity, water and nutrient cycling, and wood volume for resource extraction. Forest type is an important metric measured in the US Forest Service Forest Inventory and Analysis (FIA) program, the national forest inventory of the USA. Forest type information can be used to quantify carbon and other forest resources within specific domains to support ecological analysis and forest management decisions, such as managing for disease and pests. In this study, we developed a methodology that uses a combination of airborne hyperspectral and lidar data to map FIA-defined forest type between sparsely sampled FIA plot data collected in interior Alaska. To determine the best classification algorithm and remote sensing data for this task, five classification algorithms were tested with six different combinations of raw hyperspectral data, hyperspectral vegetation indices, and lidar-derived canopy and topography metrics. Models were trained using forest type information from 632 FIA subplots collected in interior Alaska. Of the thirty model and input combinations tested, the random forest classification algorithm with hyperspectral vegetation indices and lidar-derived topography and canopy height metrics had the highest accuracy (78% overall accuracy). This study supports random forest as a powerful classifier for natural resource data. It also demonstrates the benefits from combining both structural (lidar) and spectral (imagery) data for forest type classification.

random forest↗

Subspace-Driven Learning for Anomaly Detection in Process Transients

Nuclear power plant (NPP) monitoring and diagnostic centers are actively investigating and implementing automated anomaly detection algorithms to help plants catch anomalies sooner, thereby preventing or reducing the duration of unexpected shutdowns. Current machine learning-based anomaly detection methods are expected to be highly effective during stable, full-power operations because NPPs typically operate as baseload power generators, meaning there are extensive operating data available from plant equipment. However, it is expected that anomaly detection methods will face significant challenges during transient conditions (i.e., when power output falls below full power) because plants only occasionally operate at these lower power levels, generating sparse transient operational data, and resulting in false alarms or missed detections. Here, to address this issue, transfer learning is used, which for this problem leverages knowledge (in the form of learned features) from stable, full-power operations to improve detection accuracy during transient conditions, even with limited data. In this effort, a novel subspace approach is developed to transfer a subset of the data features from full power operation to transients. This approach is validated through experiments using synthetic data and was found to outperform two baseline transfer learning approaches in anomaly detection performance across a range of amounts of transient data used in the training process.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Informing forest carbon inventories under the Paris Agreement using ground-based forest monitoring data

Human interactions with forests have shaped Earth's climate for millennia and will continue to do so as we target net-zero emission goals. Accurately characterizing these climate impacts requires making reliable forest carbon data available for forest monitoring and planning. Here, we develop a semi-automated process for submitting forest carbon measurements from the largest relevant scientific database to the International Panel on Climate Change's Emission Factor Database, which currently has sparse forest carbon data. Building this bridge from scientific research to international policy is an important step towards managing forests in a net-zero motivated future. Humans have been influencing Earth's climate via transformative impacts on forests for millennia, and forests are now recognized as critical to climate change mitigation under the Paris Agreement. The efficacy of climate change mitigation planning and reporting depends on quality data on forest carbon (C) stocks and changes. The Emission Factor Database (EFDB) of the International Panel on Climate Change (IPCC) is intended to be a definitive source for such data, but needs comprehensive and well-documented data to be so. To facilitate submission of forest C estimates from scientific studies to EFDB, we develop and document a process for semi-automated data submission from the Global Forest C database (ForC v4.0), which is the largest compilation of ground-based forest C estimates. We then assess the data currently available through ForC and provide recommendations for improving forest data collection, analysis, and reporting. As of September 2024, ForC contained ~19,286 records potentially relevant to EFDB, 1068 of which had been submitted and posted to EFDB. These represented 19% of the total EFDB records for forest land. Records were unevenly distributed across variables and geographic regions. ForC records (37%) reviewed could not be submitted because the original publication lacked required information. In the future, ground-based forest C estimates should target gaps in the record, and studies should ensure that they report all information necessary for inclusion in EFDB. Given that climate change is rapidly impacting the world's forests, timely reporting of recent estimates will be critical to accurate forest C inventories.

54 ENVIRONMENTAL SCIENCES↗

Predicting Open Quantum Dynamics with Data-Informed Quantum-Classical Dynamics

We introduce a data-informed quantum-classical dynamics (DIQCD) approach for predicting the evolution of an open quantum system. The equation of motion in DIQCD is a Lindblad equation with a flexible, time-dependent Hamiltonian that can be optimized to fit sparse and noisy data from local observations of an extensive open quantum system. We demonstrate the accuracy and efficiency of DIQCD for both experimental and simulated quantum devices. We show that DIQCD can predict entanglement dynamics of ultracold molecules (calcium fluoride) in optical tweezer arrays. DIQCD also successfully predicts carrier mobility in organic semiconductors (rubrene) with accuracy comparable to nearly exact numerical methods.

Lindblad equation↗

Exploring the Whole Set of Accurate Sparse Interpretable Models

In data science applications, there are often many models that fit the data well. This phenomenon was called the Rashomon Effect by Leo Breiman. The set of good models is called the Rashomon Set, and the goal of this project is to locate, store, and study the Rashomon sets for classes of interpretable models, including decision trees and generalized additive models.

97 MATHEMATICS AND COMPUTING↗

Uncertainty-Guided Prediction Horizon of Phase-Resolved Ocean Wave Forecasting Under Data Sparsity: Experimental and Numerical Evaluation

Accurate short-term wave forecasting is critical for the safe and efficient operation of marine structures that rely on real-time, phase-resolved ocean wave information for control and monitoring purposes (e.g., digital twins). These systems often depend on environmental sensors (e.g., waverider buoys, wave-sensing LIDAR). Challenges arise when upstream sensor data are missing, sparse, or phase-shifted due to drift. This study investigates the performance of two machine learning models, time-series dense encoder (TiDE) and long short-term memory (LSTM), for forecasting phase-resolved ocean surface elevations under varying degrees of data degradation. We introduce the τ-trimming algorithm, which adapts the prediction horizon based on uncertainty thresholds derived from historical forecasts. Numerical wave tank (NWT) and wave basin experiments are used to benchmark model performance under short- and long-term data masking, spatially coarse sensor grids, and upstream phase shifts. Results show under a 50% probability of upstream data loss, the τ-trimmed TiDE model achieves a 46% reduction in error at the most upstream target, compared to 22% for LSTM. Furthermore, phase misalignment in upstream data introduces a near-linear increase in forecast error. Under moderate model settings, a ±3 s misalignment increases the mean absolute error by approximately 0.5 m, while the same error is accumulated at ±4 s using the more conservative approach. These findings inform the design of resilient, uncertainty-aware wave forecasting systems suited for realistic offshore sensing environments.

42 ENGINEERING↗

Grid Topology Discovery Algorithm Evaluation of Suitability for Utility Deployment (CRADA 606 Final Report)

This work presents the results of a field-informed demonstration aimed at evaluating the practical suitability of a topology discovery algorithm for utility environments. We demonstrated an algorithm that uses a graph-theory-informed state estimation approach for model selection. In collaboration with Survalent and Peninsula Light Co., the algorithm was applied to real feeder models and field measurements from supervisory control and data acquisition (SCADA) and advanced metering infrastructure (AMI) systems to identify the operational topology of a power distribution system. The demonstration assessed the algorithm’s performance under realistic data conditions, including sparse and noisy measurements, and examined its ability to identify the most likely network configurations. The results confirmed that the approach can effectively narrow down feasible topologies, providing operators with improved situational awareness of network status. Key lessons learned emphasize the need for systematic data validation and strategic sensor placement to enhance observability. These insights inform future deployment strategies and guide refinements for broader adoption in utility operations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reduction and analysis of seasons 15 and 16 (1991 - 1992) Pioneer Venus radio occulation data and correlative studies with observations of the near infrared emission of Venus

In this study, we sought to characterize variations in the abundance and distribution of subcloud H2SO4(g) in the Venus atmosphere by using a number of 13cm radio occultation measurements conducted with the Pioneer Venus Orbiter near the inferior conjunction of 1991. A total of ten data sets were examined and analyzed, producing vertical profiles of temperature and pressure in the neutral atmosphere, and sulfuric acid vapor abundance below the main cloud layer. Two of the vertical profiles of the abundance of H2SO4(g) were correlated with NIR images of the night side of Venus made during the same period of time by Boris Ragent (under a separate PVO Guest Investigator Grant). Initially, we had hoped that the combination of these two different types of data would make it possible to constrain or identify the composition of the large particles causing the features observed in the NIR images. However, the sparseness of the radio occultation data set, combined with the sparseness of the NIR data set (one image per day over an 8 day period) made it impossible to draw strong conclusions. Considered on their own, however, the parameters retrieved from the radio occultation experiments are valuable science products.

Jenkins, Jon M.↗

A ModEx Framework for Watershed Subsurface Investigation With Limited Geophysical Data Using Machine Learning and Hydrologic Modeling

Abstract Subsurface heterogeneity influences watershed hydrology strongly but remains difficult to characterize at catchment scales with sparse and costly field data. Geophysical surveys such as electromagnetic induction (EMI) provide local spatial subsurface images yet scaling them to watershed scales and converting EMI‐derived resistivity into hydraulic properties remains a challenge. We present a Model–Experiment (ModEx) framework that integrates limited EMI data with machine learning (ML) and hydrologic modeling to improve process representation and guide field investigations. Sparse EMI surveys were scaled to the catchment scale using a Random Forest model, and the resulting resistivity fields were combined with nearby borehole constraints to parameterize a hydrologic model. The EMI‐informed hydrological simulations improved predictions of streamflow sustained by subsurface flow and shallow saturation patterns. By combining EMI data and ML with hydrologic modeling, the ModEx framework guides future subsurface surveys, providing a transferable and efficient strategy for data–model integration across diverse watersheds. Plain Language Summary Mapping the underground network of soil and rock that controls water is essential for predicting floods and droughts, but seeing underground is difficult and expensive. We cannot drill everywhere, so scientists use geophysical tools to scan broad areas. There are two key challenges: these geophysical scans are often sparse across the whole watershed, and the geophysical data is hard to translate into water‐related properties. We used artificial intelligence to solve these problems. We taught a computer to find patterns linking the limited geophysical data to the land surface properties. This allowed it to fill in the gaps and create a complete, useful subsurface map for the entire watershed. This new map improves hydrologic simulations, leading to more accurate predictions of water movement in the watershed. It also helps scientists build better models with less data and generates a priority map showing where to measure next, making future investigations more efficient. Key Points Limited EMI scaled with ML improves catchment‐scale subsurface parameterization for hydrologic models The framework integrates hydrologic modeling with limited geophysical data to support subsurface investigation design ModEx framework offers a transferable data–model integration strategy that quantifies and reduces uncertainty guiding watershed studies

Chen, Hang↗

The Parameterization of Top-Hat Particle Sensors with Microchannel-Plate-Based Detection Systems and its Application to the Fast Plasma Investigation on NASA's Magnetospheric MultiScale Mission

The most common instrument for low energy plasmas consists of a top-hat electrostatic analyzer geometry coupled with a microchannel-plate (MCP)-based detection system. While the electrostatic optics for such sensors are readily simulated and parameterized during the laboratory calibration process, the detection system is often less well characterized. Furthermore, due to finite resources, for large sensor suites such as the Fast Plasma Investigation (FPI) on NASA's Magnetospheric Multiscale (MMS) mission, calibration data are increasingly sparse. Measurements must be interpolated and extrapolated to understand instrument behavior for untestable operating modes and yet sensor inter-calibration is critical to mission success. To characterize instruments from a minimal set of parameters we have developed the first comprehensive mathematical description of both sensor electrostatic optics and particle detection systems. We include effects of MCP efficiency, gain, scattering, capacitive crosstalk, and charge cloud spreading at the detector output. Our parameterization enables the interpolation and extrapolation of instrument response to all relevant particle energies, detector high voltage settings, and polar angles from a small set of calibration data. We apply this model to the 32 sensor heads in the Dual Electron Sensor (DES) and 32 sensor heads in the Dual Ion Sensor (DIS) instruments on the 4 MMS observatories and use least squares fitting of calibration data to extract all key instrument parameters. Parameters that will evolve in flight, namely MCP gain, will be determined daily through application of this model to specifically tailored in-flight calibration activities, providing a robust characterization of sensor suite performance throughout mission lifetime. Beyond FPI, our model provides a valuable framework for the simulation and evaluation of future detection system designs and can be used to maximize instrument understanding with minimal calibration resources.

Instrument-model↗

Temperature-dependent Absolute Refractive Index Measurements of Synthetic Fused Silica

Using the Cryogenic, High-Accuracy Refraction Measuring System (CHARMS) at NASA's Goddard Space Flight Center, we have measured the absolute refractive index of five specimens taken from a very large boule of Corning 7980 fused silica from temperatures ranging from 30 to 310 K at wavelengths from 0.4 to 2.6 microns with an absolute uncertainty of plus or minus 1 x 10 (exp -5). Statistical variations in derived values of the thermo-optic coefficient (dn/dT) are at the plus or minus 2 x 10 (exp -8)/K level. Graphical and tabulated data for absolute refractive index, dispersion, and thermo-optic coefficient are presented for selected wavelengths and temperatures along with estimates of uncertainty in index. Coefficients for temperature-dependent Sellmeier fits of measured refractive index are also presented to allow accurate interpolation of index to other wavelengths and temperatures. We compare our results to those from an independent investigation (which used an interferometric technique for measuring index changes as a function of temperature) whose samples were prepared from the same slugs of material from which our prisms were prepared in support of the Kepler mission. We also compare our results with sparse cryogenic index data from measurements of this material from the literature.

Leviton, Douglas B.↗

Monitoring global vegetation using Nimbus-7 37 GHz data - Some empirical relations

The difference of the vertically and horizontally polarized brightness temperatures observed by the 37 GHz channel of the SMMR on board the Nimbus-7 satellite are correlated temporally with three indicators of vegetation density, namely the temporal variation of the atmospheric CO2 concentration at Mauna Loa (Hawaii), rainfall over the Sahel and the normalized difference vegetation index derived from the AVHRR on board the NOAA-7 satellite. SMMR 37 GHz and AVHRR provide complementary data sets for monitoring global vegetation, the 37 GHz data being more suitable for arid and semiarid regions as these data are more sensitive to changes in sparse vegetation. The 37-GHz data might be useful for understanding desertification and indexing Co2 exchange between the biosphere and the atmosphere.

Choudhury, B. J.↗