Search NASA⌕ Search

SEARCH · Search NASA

Results for “sparse data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33

Global Estimates of Fine Particulate Matter Using a Combined Geophysical-Statistical Method with Information from Satellites, Models, and Monitors

We estimated global fine particulate matter (PM(sub 2.5)) concentrations using information from satellite-, simulation- and monitor-based sources by applying a Geographically Weighted Regression (GWR) to global geophysically-based satellite-derived PM(sub 2.5) estimates. Aerosol optical depth from multiple satellite products (MISR, MODIS Dark Target, MODIS and SeaWiFS Deep Blue, and MODIS MAIAC) was combined with simulation (GEOS-Chem) based upon their relative uncertainties as determined using ground-based sun photometer (AERONET) observations for 1998−2014. The GWR predictors included simulated aerosol composition and land use information. The resultant PM(sub 2.5) estimates were highly consistent (R(sup 2) equals 0.81) with out-of-sample cross-validated PM(sub 2.5) concentrations from monitors. The global population-weighted annual average PM(sub 2.5) concentrations were 3-fold higher than the 10 micrograms per cubic meter WHO guideline, driven by exposures in Asian and African regions. Estimates in regions with high contributions from mineral dust were associated with higher uncertainty, resulting from both sparse ground-based monitoring, and challenging conditions for retrieval and simulation. This approach demonstrates that the addition of even sparse ground-based measurements to more globally continuous PM(sub 2.5) data sources can yield valuable improvements to PM(sub 2.5) characterization on a global scale.

aerosols↗

PaleoSTeHM v1.0: a modern, scalable spatiotemporal hierarchical modeling framework for paleo-environmental data

Abstract. Geological records of past environmental change provide crucial insights into long-term climate variability, trends, non-stationarity, and nonlinear feedback mechanisms. However, reconstructing spatiotemporal fields from these records is statistically challenging due to their sparse, indirect, and noisy nature. Here, we present PaleoSTeHM, a scalable and modern framework for spatiotemporal hierarchical modeling of paleo-environmental data. This framework enables the implementation of flexible statistical models that rigorously quantify spatial and temporal variability from geological data while clearly distinguishing measurement and inferential uncertainty from process variability. We illustrate its application by reconstructing temporal and spatiotemporal paleo-sea-level changes across multiple locations. Using various modeling and analysis choices, PaleoSTeHM demonstrates the impact of different methods on inference results and computational efficiency. Our results highlight the critical role of model selection in addressing specific paleo-environmental questions, showcasing the PaleoSTeHM framework's potential to enhance the robustness and transparency of paleo-environmental reconstructions.

58 GEOSCIENCES↗

RatXcan: A framework for cross-species integration of genome-wide association and gene expression data

Genome-wide association studies (GWAS) have implicated specific alleles and genes as risk factors for numerous complex traits. However, translating GWAS results into biologically and therapeutically meaningful discoveries remains extremely challenging. Most GWAS results identify noncoding regions of the genome, suggesting that differences in gene regulation are the major driver of trait variability. To better integrate GWAS results with gene regulatory polymorphisms, we previously developed PrediXcan (also known as “transcriptome-wide association studies” orTWAS), which maps SNPs to predicted gene expression using GWAS data. In this study, we developed RatXcan, a framework that extends this methodology to outbred heterogeneous stock (HS) rats. RatXcan accounts for the close familial relationships among HS rats by modeling the relatedness with a random effect that encodes the genetic relatedness. RatXcan also corrects for polygenic-driven inflation because of the equivalence between a relatedness random effect and the infinitesimal polygenic model. To develop RatXcan, we trained transcript predictors for 8,934 genes using reference genotype and expression data from five rat brain regions. We found that the cis genetic architecture of gene expression in both rats and humans was sparse and similar across brain tissues. We tested the association between predicted expression in rats and two example traits (body length and BMI) using phenotype and genotype data from 5,401 densely genotyped HS rats and identified a significant enrichment between the genes associated with rat and human body length and BMI. Thus, RatXcan represents a valuable tool for identifying the relationship between gene expression and phenotypes across species and paves the way to explore shared biological mechanisms of complex traits.

Genetics & Heredity↗

Tropospheric-stratospheric exchange, part 1.1A

Much of the observational evidence of large scale tropospheric-stratospheric exchange has been obtained by radiosonde and satellite radiane data. So far mesosphere-stratosphere-troposphere (MST) radars have made mininal contributions, in part due to their recent use as a meteorological tool, intermittent operation at some facilities and sparse geographic distribution. However, as more MST facilities come on-line in more locations, the good time and height resolution data throughtout the troposphere and much of the stratosphere obtainable by MST radars will enhance the detail of stratospheric and tropospheric circulations and interactions. On smaller scales MST radars have already been used to examine convective forcing from the troposphere into the stratosphere and subsequent launching of gravity waves (LARSEN et al., 1982). Observations of persistent turbulent layers in the stratosphere over Arecibo, attributable to inertial oscillations, appear to propagate away from a source region near the tropopause (SATO and WOODMAN, 1982). MST radars offer the availability of high resolution wind data in height and time needed to observe interactions between the troposphere and stratosphere. The lack of geographic coverage (e.g., equatorial regions) and insufficient data bases at many MST facilities presently inhibit studies of large-scale interactions. At present MST radars can be used to examine smaller scale interactions.

Cornish, C. R.↗

Incremental Learning for Passive Microwave Precipitation Retrievals using Advanced Technology Microwave Sounder

Spaceborne passive microwave (PMW) radiometry is central to global precipitation monitoring, yet retrieval uncertainties remain substantial, particularly for cross-track sounders whose variable footprints and channel configurations are optimized for atmospheric temperature and moisture profiling rather than precipitation. Consequently, existing operational products often exhibit angular-dependent biases, limited effective swath utilization, unrealistic rainfall probability distributions, and systematic misclassification of precipitation phase. These limitations are further compounded by the scarcity of globally accurate and representative precipitation observations, as training data from the Dual-frequency Precipitation Radar (DPR) and the Cloud Profiling Radar (CPR) are spatially sparse, lack uniform global coverage, and exhibit heterogeneous error characteristics across precipitation regimes. To address these challenges, this study presents a supervised retrieval algorithm that incrementally trains an ensemble of extreme gradient-boosted decision trees by augmenting base learners with pre-training on reanalysis data and post-training on coincident DPR and CPR observations matched with the Advanced Technology Microwave Sounder (ATMS). By transferring prior information from reanalysis to posterior constraints from radar observations and adopting a sequential detection–estimation strategy for precipitation phase and rate retrieval, the proposed approach yields retrievals across the full ATMS swath that are largely free from persistent deficiencies in current Global Precipitation Measurement (GPM) passive microwave operational products. In particular, the method resolves bimodal artifacts in rainfall retrievals and mitigates systematic high-latitude snowfall biases, including overestimation across the Arctic and underestimation across the Antarctic. Validation against independent Multi-Radar Multi-Sensor (MRMS) data over the Contiguous United States (CONUS) further demonstrates improved performance in precipitation phase detection and rate estimation relative to both reanalysis and current GPM PMW products.

Mahyar Garshasbi↗

Cloud Masking and Surface Temperature Distribution in the Polar Regions Using AVHRR and other Satellite Data

Surface temperature is one of the key variables associated with weather and climate. Accurate measurements of surface air temperatures are routinely made in meteorological stations around the world. Also, satellite data have been used to produce synoptic global temperature distributions. However, not much attention has been paid on temperature distributions in the polar regions. In the polar regions, the number of stations is very sparse. Because of adverse weather conditions and general inaccessibility, surface field measurements are also limited. Furthermore, accurate retrievals from satellite data in the region have been difficult to make because of persistent cloudiness and ambiguities in the discrimination of clouds from snow or ice. Surface temperature observations are required in the polar regions for air-sea-ice interaction studies, especially in the calculation of heat, salinity, and humidity fluxes. They are also useful in identifying areas of melt or meltponding within the sea ice pack and the ice sheets and in the calculation of emissivities of these surfaces. Moreover, the polar regions are unique in that they are the sites of temperature extremes, the location of which is difficult to identify without a global monitoring system. Furthermore, the regions may provide an early signal to a potential climate change because such signal is expected to be amplified in the region due to feedback effects. In cloud free areas, the thermal channels from infrared systems provide surface temperatures at relatively good accuracies. Previous capabilities include the use of the Temperature Humidity Infrared Radiometer (THIR) onboard the Nimbus-7 satellite which was launched in 1978. Current capabilities include the use of the Advance Very High Resolution Radiometer (AVHRR) aboard NOAA satellites. Together, these two systems cover a span of 16 years of thermal infrared data. Techniques for retrieving surface temperatures with these sensors in the polar regions have been developed. Errors have been estimated to range from 1K to 5K mainly due to cloud masking problems. With many additional channels available, it is expected that the EOS-Moderate Resolution Imaging Spectroradiometer (MODIS) will provide an improved characterization of clouds and a good discrimination of clouds from snow or ice surfaces.

Comiso, Joey C.↗

Interhemispheric comparison of atmospheric circulation features as evaluated from NIMBUS satellite data

Findings are presented for IRIS data from NIMBUS 3 in mapping the global ozone distribution. The seasonal and regional variations of ozone, especially in the Southern Hemisphere, reveal features that were not evident from the sparse ground-based ozone observation network in this hemisphere. A regression analysis was undertaken for temperature and height fields on radiance data. Spectrum analyses of upper wind data from the North American section and Australia were completed.

Reiter, E. R.↗

Data Mining and Optimization Tools for Developing Engine Parameters Tools

This project was awarded for understanding the problem and developing a plan for Data Mining tools for use in designing and implementing an Engine Condition Monitoring System. From the total budget of $5,000, Tricia and I studied the problem domain for developing ail Engine Condition Monitoring system using the sparse and non-standardized datasets to be available through a consortium at NASA Lewis Research Center. We visited NASA three times to discuss additional issues related to dataset which was not made available to us. We discussed and developed a general framework of data mining and optimization tools to extract useful information from sparse and non-standard datasets. These discussions lead to the training of Tricia Erhardt to develop Genetic Algorithm based search programs which were written in C++ and used to demonstrate the capability of GA algorithm in searching an optimal solution in noisy datasets. From the study and discussion with NASA LERC personnel, we then prepared a proposal, which is being submitted to NASA for future work for the development of data mining algorithms for engine conditional monitoring. The proposed set of algorithm uses wavelet processing for creating multi-resolution pyramid of the data for GA based multi-resolution optimal search. Wavelet processing is proposed to create a coarse resolution representation of data providing two advantages in GA based search: 1. We will have less data to begin with to make search sub-spaces. 2. It will have robustness against the noise because at every level of wavelet based decomposition, we will be decomposing the signal into low pass and high pass filters.

Dhawan, Atam P.↗

Financial Exposure After a Sealed-Source Release: Insurance Limits, Federal Cost Pathways, and Implications for Gamma Irradiator Substitution

This report assesses whether private insurance and existing federal authorities would likely provide meaningful financial protection to private operators of self-shielded gamma irradiators after a major sealed-source release. It does not answer this question quantitatively because publicly observable data on premiums, limits, uptake, and claims outcomes for this risk class appear sparse. Instead, it uses a qualitative structural analysis based on prior literature, review of relevant federal authorities, limited observable insurance-market evidence, expert outreach, and historical analogs. The analysis finds that available public and private mechanisms do not combine into a clear, dependable, or readily verifiable compensation structure for ordinary private sealed-source operators. Institutions therefore should not assume that either insurance or government response will make them financially whole after a severe incident. Source reduction and replacement remain more dependable than post-event financial mechanisms for reducing institutional exposure and broader radiological risk.

99 GENERAL AND MISCELLANEOUS↗

The growth of the UniTree mass storage system at the NASA Center for Computational Sciences

In October 1992, the NASA Center for Computational Sciences made its Convex-based UniTree system generally available to users. The ensuing months saw the growth of near-online data from nil to nearly three terabytes, a doubling of the number of CPU's on the facility's Cray YMP (the primary data source for UniTree), and the necessity for an aggressive regimen for repacking sparse tapes and hierarchical 'vaulting' of old files to freestanding tape. Connectivity was enhanced as well with the addition of UltraNet HiPPI. This paper describes the increasing demands placed on the storage system's performance and throughput that resulted from the significant augmentation of compute-server processor power and network speed.

Tarshish, Adina↗

Application of Satellite Gravimetry for Water Resource Vulnerability Assessment

The force of Earth's gravity field varies in proportion to the amount of mass near the surface. Spatial and temporal variations in the gravity field can be measured via their effects on the orbits of satellites. The Gravity Recovery and Climate Experiment (GRACE) is the first satellite mission dedicated to monitoring temporal variations in the gravity field. The monthly gravity anomaly maps that have been delivered by GRACE since 2002 are being used to infer changes in terrestrial water storage (the sum of groundwater, soil moisture, surface waters, and snow and ice), which are the primary source of gravity variability on monthly to decadal timescales after atmospheric and oceanic circulation effects have been removed. Other remote sensing techniques are unable to detect water below the first few centimeters of the land surface. Conventional ground based techniques can be used to monitor terrestrial water storage, but groundwater, soil moisture, and snow observation networks are sparse in most of the world, and the countries that do collect such data rarely are willing to share them. Thus GRACE is unique in its ability to provide global data on variations in the availability of fresh water, which is both vital to life on land and vulnerable to climate variability and mismanagement. This chapter describes the unique and challenging aspects of GRACE terrestrial water storage data, examples of how the data have been used for research and applications related to fresh water vulnerability and change, and prospects for continued contributions of satellite gravimetry to water resources science and policy.

Rodell, Matthew↗

Evaluating Drought Indices for Early Warning in East and Southern Africa

Sparsely populated regions of east and southern Africa often have little ground based data to monitor drought, crops, and water resources. The Regional Hydrologic Extremes Assessment System (RHEAS) is a NASA supported data assimilation framework that combines both a hydrologic and crop model, which provide another monitoring approach. RHEAS can also be run in forecast mode, providing outlook on drought and crop yield.

Miller, Sara↗

Learning Physically Interpretable Atmospheric Models From Data With WSINDy

The multiscale and turbulent nature of Earth's atmosphere has historically rendered accurate weather modeling a hard problem. Recently, there has been an explosion of interest surrounding data-driven approaches to weather modeling, which in many cases show improved forecasting accuracy and computational efficiency when compared to traditional methods. However, many of the current data-driven approaches employ highly parameterized neural networks, often resulting in uninterpretable models and limited gains in scientific understanding. In this work, we address the interpretability problem by explicitly discovering partial differential equations governing atmospheric phenomena, identifying symbolic mathematical models with direct physical interpretations. The purpose of this paper is to demonstrate that, in particular, the weak-form sparse identification of nonlinear dynamics (WSINDy) algorithm can learn effective atmospheric models from both simulated and assimilated data. Our approach adapts the standard WSINDy algorithm to work with high-dimensional fluid data of arbitrary spatial dimension.

58 GEOSCIENCES↗

Assessment of Computer-based Geologic Mapping of Rock Units in the LANDSAT-4 Scene of Northern Death Valley, California

Geologists obtain low accuracy levels when maps derived from LANDSAT MSS data are compared with those made by conventional methods. Procedures developed for the IDIMS computer system and used to classify a subset of a TM image of the Death Valley, California - Nevada border are described. Despite the superior resolution, broader spectral coverage, and greater sensitivity inherent to the TM, the actual recorded measured accuracy was in the same narrow range (30 to 60%) recorded for MSS data from earlier LANDSATs. The supervised classification approach appears to be superior to the unsupervised approach when applied to vegetation-sparse surfaces composed of spectrally contrasting rock/soil units distributed in relatively flat to low relief terrain. As spatial resolution improves and optimal spectral bands for identifying rock materials are specified, use of classified multispectral remote sensing data from air and space when coupled with supporting field calibration and checks should become the dominant way in which geologic mapping is carried out in future decades.

Short, N. M.↗

Semi-Analytical Hierarchical Bayesian Inference of Nonlinear Model Structure in Stochastic Dynamics: Applied to Compartmental Models of Infectious Diseases

A Bayesian computational framework for parsimonious inference in stochastic nonlinear dynamical systems is presented. This framework enables the concurrent estimation of system states, time-varying parameters, time-invariant parameters, and the optimal sparsity structure of the model parameters. Because differential equation-based models are often simplified mechanistic or phenomenological representations, robust inference from noisy measurement data requires explicit treatment of model error and uncertainty. Model error and time-varying parameters can be represented as random processes, enabling inference while making minimal assumptions about the underlying sources of discrepancy and variability. Adopting stochastic differential equation representations affords the model significant flexibility, but can also render it susceptible to overfitting during statistical inversion, where the inferred model may track noise rather than the underlying signal. To alleviate the effects of overfitting and to enable the discovery of the optimal sparse representation of the time-invariant parameters, a Bayesian sparse learning algorithm is embedded within the framework. This sparse learning framework adopts an approximate hierarchical Bayesian setting defined by a series of semi-analytical expressions. The model structure inference framework is validated using a stochastic compartmental model for tracking and forecasting active cases of an infectious disease. Compartmental models describe population-level infectious disease dynamics through interactions among population fractions grouped by disease state. Mathematically, such models consist of a system of coupled ordinary differential equations. This example adopts an expressive compartmental model that includes multiple possible interactions between disease states, motivated by early uncertainty surrounding COVID-19 reinfection dynamics and their implications for long-term epidemic forecasting. The sparse learning exercise permits the inference of a priori unknown epidemiological dynamics from simulated public health data, discovering the nested compartmental model that optimizes the trade-off between average data-fit and model complexity. It is shown that inducing sparsity among the model parameters eliminates redundant interactions between compartments, equivalently revealing the optimal coupling structure between differential equations.

97 MATHEMATICS AND COMPUTING↗

Application of Statistical Methods of Rain Rate Estimation to Data From The TRMM Precipitation Radar

The TRMM Precipitation Radar is well suited to statistical methods in that the measurements over any given region are sparsely sampled in time. Moreover, the instantaneous rain rate estimates are often of limited accuracy at high rain rates because of attenuation effects and at light rain rates because of receiver sensitivity. For the estimation of the time-averaged rain characteristics over an area both errors are relevant. By enlarging the space-time region over which the data are collected, the sampling error can be reduced. However. the bias and distortion of the estimated rain distribution generally will remain if estimates at the high and low rain rates are not corrected. In this paper we use the TRMM PR data to investigate the behavior of 2 statistical methods the purpose of which is to estimate the rain rate over large space-time domains. Examination of large-scale rain characteristics provides a useful starting point. The high correlation between the mean and standard deviation of rain rate implies that the conditional distribution of this quantity can be approximated by a one-parameter distribution. This property is used to explore the behavior of the area-time-integral (ATI) methods where fractional area above a threshold is related to the mean rain rate. In the usual application of the ATI method a correlation is established between these quantities. However, if a particular form of the rain rate distribution is assumed and if the ratio of the mean to standard deviation is known, then not only the mean but the full distribution can be extracted from a measurement of fractional area above a threshold. The second method is an extension of this idea where the distribution is estimated from data over a range of rain rates chosen in an intermediate range where the effects of attenuation and poor sensitivity can be neglected. The advantage of estimating the distribution itself rather than the mean value is that it yields the fraction of rain contributed by the light and heavy rain rates. This is useful in estimating the fraction of rainfall contributed by the rain rates that go undetected by the radar. The results at high rain rates provide a cross-check on the usual attenuation correction methods that are applied at the highest resolution of the instrument.

Meneghini, R.↗

Application of Reconfigurable Computing Technology to Multi-KiloHertz Micro-Laser Altimeter (MMLA) Data Processing

The Multi-KiloHertz Micro-Laser Altimeter (MMLA) is an aircraft based instrument developed by NASA Goddard Space Flight Center with several potential spaceflight applications. This presentation describes how reconfigurable computing technology was employed to perform MMLA signal extraction in real-time under realistic operating constraints. The MMLA is a "single-photon-counting" airborne laser altimeter that is used to measure land surface features such as topography and vegetation canopy height. This instrument has to date flown a number of times aboard the NASA P3 aircraft acquiring data at a number of sites in the Mid-Atlantic region. This instrument pulses a relatively low-powered laser at a very high rate (10 kHz) and then measures the time-of-flight of discrete returns from the target surface. The instrument then bins these measurements into a two-dimensional array (vertical height vs. horizontal ground track) and selects the most likely signal path through the array. Return data that does not correspond to the selected signal path are classified as noise returns and are then discarded. The MMLA signal extraction algorithm is very compute intensive in that a score must be computed for every possible path through the two dimensional array in order to select the most likely signal path. Given a typical array size with 50 x 6, up to 33 arrays must be processed per second. And for each of these arrays, roughly 12,000 individual paths must be scored. Furthermore, the number of paths increases exponentially with the horizontal size of the array, and linearly with the vertical size. Yet, increasing the horizontal and vertical sizes of the array offer science advantages such as improved range, resolution, and noise rejection. Due to the volume of return data and the compute intensive signal extraction algorithm, the existing PC-based MMLA data system has been unable to perform signal extraction in real-time unless the array is limited in size to one column, This limits the ability of the MMLA to operate in environments with sparse signal returns and a high number of noise return. However, under an IR&D project, an FPGA-based, reconfigurable computing data system has been developed that has been demonstrated to perform real-time signal extraction under realistic operating constraints. This reconfigurable data system is based on the commercially available Firebird Board from Annapolis Microsystems. This PCI board consists of a Xilinx Virtex 2000E FPGA along with 36 MB of SRAM arranged in five separately addressable banks. This board is housed in a rackmount PC with dual 850MHz Pentium processors running the Windows 2000 operating system. This data system performs all signal extraction in hardware on the Firebird, but also runs the existing "software based" signal extraction in tandem for comparison purposes. Using a relatively small amount of the Virtex XCV2000E resources, the reconfigurable data system has demonstrated to improve performance improvement over the existing software based data system by an order of magnitude. Performance could be further improved by employing parallelism. Ground testing and a preliminary engineering test flight aboard the NASA P3 has been performed, during which the reconfigurable data system has been demonstrated to match the results of the existing data system.

Powell, Wesley↗

Using High Frequency Passive Microwave, A-train, and TRMM Data to Evaluate Hydrometer Structure in the NASA GEOS-5 Data Assimilation System

Validating water vapor and prognostic condensate in global models remains a challenging research task. Model parameterizations are still subject to a large number of tunable parameters; furthermore, accurate and representative in situ observations are very sparse, and satellite observations historically have significant quantitative uncertainties. Progress on improving cloud / hydrometeor fields in models stands to benefit greatly from the growing inventory ofA-Train data sets. ill the present study we are using a variety of complementary satellite retrievals of hydrometeors to examine condensate produced by the emerging NASA Modem Era Retrospective Analysis for Research and Applications, MERRA, and its associated atmospheric general circulation model GEOS5. Cloud and precipitation are generated by both grid-scale prognostic equations and by the Relaxed Arakawa-Schubert (RAS) diagnostic convective parameterization. The high frequency channels (89 to 183.3 GHz) from AMSU-B and MRS on NOAA polar orbiting satellites are being used to evaluate the climatology and variability of precipitating ice from tropical convective anvils. Vertical hydrometeor structure from the Tropical Rainfall Measuring Mission (TRMM) and CloudSat radars are used to develop statistics on vertical hydrometeor structure in order to better interpret the extensive high frequency passive microwave climatology. Cloud liquid and ice water path data retrieved from the Moderate Resolution Imaging Spectroradiometer, MODIS, are used to investigate relationships between upper level cloudiness and tropical deep convective anvils. Together these data are used to evaluate cloud / ice water path, gross aspects of vertical hydrometeor structure, and the relationship between cloud extent and surface precipitation that the MERRA reanalysis must capture.

Robertson, Franklin↗