Search NASA⌕ Search

SEARCH · Search NASA

Results for “sparse data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Arctic Strato-Mesospheric Temperature and Wind Variations

Upper stratosphere and mesosphere rocket measurements are actively used to investigate interaction between the neutral, electrical, and chemical atmospheres and between lower and upper layers of these regions. Satellite temperature measurements from HALOE and from inflatable falling spheres complement each other and allow illustrations of the annual cycle to 85 km altitude. Falling sphere wind and temperature measurements reveal variability that differs as a function of altitude, location, and time. We discuss the state of the Arctic atmosphere during the summer 2002 (Andoya, Norway) and winter 2003 (ESRANGE, Sweden) campaigns of MaCWAVE. Balloon-borne profiles to 30 km altitude and sphere profiles between 50 and 90 km show unique small-scale structure. Nonetheless, there are practical implications that additional measurements are very much needed to complete the full vertical profile picture. Our discussion concentrates on the distribution of temperature and wind and their variability. However, reliable measurements from other high latitude NASA programs over a number of years are available to help properly calculate mean values and the distribution of the individual measurements. Since the available rocket data in the Arctic's upper atmosphere are sparse the results we present are basically a snapshot of atmospheric structure.

Schmidlin, F. J.↗

Method and system for data clustering for very large databases

Multi-dimensional data contained in very large databases is efficiently and accurately clustered to determine patterns therein and extract useful information from such patterns. Conventional computer processors may be used which have limited memory capacity and conventional operating speed, allowing massive data sets to be processed in a reasonable time and with reasonable computer resources. The clustering process is organized using a clustering feature tree structure wherein each clustering feature comprises the number of data points in the cluster, the linear sum of the data points in the cluster, and the square sum of the data points in the cluster. A dense region of data points is treated collectively as a single cluster, and points in sparsely occupied regions can be treated as outliers and removed from the clustering feature tree. The clustering can be carried out continuously with new data points being received and processed, and with the clustering feature tree being restructured as necessary to accommodate the information from the newly received data points.

Zhang, Tian↗

Multiple View Zenith Angle Observations of Reflectance From Ponderosa Pine Stands

Reflectance factors (RF(lambda)) from dense and sparse ponderosa pine (Pinus ponderosa) stands, derived from radiance data collected in the solar principal plane by the Advanced Solid-State Array Spectro-radiometer (ASAS), were examined as a function of view zenith angle (theta(sub v)). RF(lambda) was maximized with theta(sub v) nearest the solar retrodirection, and minimized near the specular direction throughout the ASAS spectral region. The dense stand had much higher RF anisotropy (ma)dmurn RF is minimum RF) in the red region than did the sparse stand (relative differences of 5.3 vs. 2.75, respectively), as a function of theta(sub v), due to the shadow component in the canopy. Anisotropy in the near-infrared (NIR) was more similar between the two stands (2.5 in the dense stand and 2.25 in the sparse stand); the dense stand exhibited a greater hotspot effect than 20 the sparse stand in this spectral region. Two common vegetation transforms, the NIR/red ratio and the normalized difference vegetation index (NDVI), both showed a theta(sub v) dependence for the dense stand. Minimum values occurred near the retrodirection and maximum values occurred near the specular direction. Greater relative differences were noted for the NIR/red ratio (2.1) than for the NDVI (1.3). The sparse stand showed no obvious dependence on theta(sub v) for either transform, except for slightly elevated values toward the specular direction.

Johnson, Lee F.↗

Soil Moisture Active Passive Mission L4_SM Data Product Assessment (Version 2 Validated Release)

During the post-launch SMAP calibration and validation (Cal/Val) phase there are two objectives for each science data product team: 1) calibrate, verify, and improve the performance of the science algorithm, and 2) validate the accuracy of the science data product as specified in the science requirements and according to the Cal/Val schedule. This report provides an assessment of the SMAP Level 4 Surface and Root Zone Soil Moisture Passive (L4_SM) product specifically for the product's public Version 2 validated release scheduled for 29 April 2016. The assessment of the Version 2 L4_SM data product includes comparisons of SMAP L4_SM soil moisture estimates with in situ soil moisture observations from core validation sites and sparse networks. The assessment further includes a global evaluation of the internal diagnostics from the ensemble-based data assimilation system that is used to generate the L4_SM product. This evaluation focuses on the statistics of the observation-minus-forecast (O-F) residuals and the analysis increments. Together, the core validation site comparisons and the statistics of the assimilation diagnostics are considered primary validation methodologies for the L4_SM product. Comparisons against in situ measurements from regional-scale sparse networks are considered a secondary validation methodology because such in situ measurements are subject to up-scaling errors from the point-scale to the grid cell scale of the data product. Based on the limited set of core validation sites, the wide geographic range of the sparse network sites, and the global assessment of the assimilation diagnostics, the assessment presented here meets the criteria established by the Committee on Earth Observing Satellites for Stage 2 validation and supports the validated release of the data. An analysis of the time average surface and root zone soil moisture shows that the global pattern of arid and humid regions are captured by the L4_SM estimates. Results from the core validation site comparisons indicate that "Version 2" of the L4_SM data product meets the self-imposed L4_SM accuracy requirement, which is formulated in terms of the ubRMSE: the RMSE (Root Mean Square Error) after removal of the long-term mean difference. The overall ubRMSE of the 3-hourly L4_SM surface soil moisture at the 9 km scale is 0.035 cubic meters per cubic meter requirement. The corresponding ubRMSE for L4_SM root zone soil moisture is 0.024 cubic meters per cubic meter requirement. Both of these metrics are comfortably below the 0.04 cubic meters per cubic meter requirement. The L4_SM estimates are an improvement over estimates from a model-only SMAP Nature Run version 4 (NRv4), which demonstrates the beneficial impact of the SMAP brightness temperature data. L4_SM surface soil moisture estimates are consistently more skillful than NRv4 estimates, although not by a statistically significant margin. The lack of statistical significance is not surprising given the limited data record available to date. Root zone soil moisture estimates from L4_SM and NRv4 have similar skill. Results from comparisons of the L4_SM product to in situ measurements from nearly 400 sparse network sites corroborate the core validation site results. The instantaneous soil moisture and soil temperature analysis increments are within a reasonable range and result in spatially smooth soil moisture analyses. The O-F residuals exhibit only small biases on the order of 1-3 degrees Kelvin between the (re-scaled) SMAP brightness temperature observations and the L4_SM model forecast, which indicates that the assimilation system is largely unbiased. The spatially averaged time series standard deviation of the O-F residuals is 5.9 degrees Kelvin, which reduces to 4.0 degrees Kelvin for the observation-minus-analysis (O-A) residuals, reflecting the impact of the SMAP observations on the L4_SM system. Averaged globally, the time series standard deviation of the normalized O-F residuals is close to unity, which would suggest that the magnitude of the modeled errors approximately reflects that of the actual errors. The assessment report also notes several limitations of the "Version 2" L4_SM data product and science algorithm calibration that will be addressed in future releases. Regionally, the time series standard deviation of the normalized O-F residuals deviates considerably from unity, which indicates that the L4_SM assimilation algorithm either over- or under-estimates the actual errors that are present in the system. Planned improvements include revised land model parameters, revised error parameters for the land model and the assimilated SMAP observations, and revised surface meteorological forcing data for the operational period and underlying climatological data. Moreover, a refined analysis of the impact of SMAP observations will be facilitated by the construction of additional variants of the model-only reference data. Nevertheless, the “Version 2” validated release of the L4_SM product is sufficiently mature and of adequate quality for distribution to and use by the larger science and application communities.

SMAP L4_SM↗

The SMAP Mission Combined Active-Passive Soil Moisture Product at 9 Km and 3 Km Spatial Resolutions

The NASA Soil Moisture Active Passive (SMAP) mission was launched on January 31st, 2015. The spacecraft was to provide high-resolution (3 km and 9 km) global soil moisture estimates at regular intervals by combining for the first time L-band radiometer and radar observations. On July 7th, 2015, a component of the SMAP radar failed and the radar ceased operation. However, before this occurred the mission was able to collect and process ~2.5 months of the SMAP high-resolution active-passive soil moisture data (L2SMAP) that coincided with the Northern Hemisphere's vegetation green-up and crop growth season. In this study, we evaluate the SMAP high-resolution soil moisture product derived from several alternative algorithms against in situ data from core calibration and validation sites (CVS), and sparse networks. The baseline algorithm had the best comparison statistics against the CVS and sparse networks. The overall unbiased root-mean-square-difference is close to the 0.04 cu. m/cu. m the SMAP mission requirement. A 3 km spatial resolution soil moisture product was also examined. This product had an unbiased root-mean-square-difference of ~0.053 cu. m/cu. m. The SMAP L2SMAP product for ~2.5 months is now validated for use in geophysical applications and research and available to the public through the NASA Distributed Active Archive Center (DAAC) at the National Snow and Ice Data Center (NSIDC). The L2SMAP product is packaged with the geo-coordinates, acquisition times, and all requisite ancillary information. Although limited in duration, SMAP has clearly demonstrated the potential of using a combined L-band radar-radiometer for proving high spatial resolution and accurate global soil moisture.

high resolution soil moisture↗

Application-Controlled Demand Paging for Out-of-Core Visualization

In the area of scientific visualization, input data sets are often very large. In visualization of Computational Fluid Dynamics (CFD) in particular, input data sets today can surpass 100 Gbytes, and are expected to scale with the ability of supercomputers to generate them. Some visualization tools already partition large data sets into segments, and load appropriate segments as they are needed. However, this does not remove the problem for two reasons: 1) there are data sets for which even the individual segments are too large for the largest graphics workstations, 2) many practitioners do not have access to workstations with the memory capacity required to load even a segment, especially since the state-of-the-art visualization tools tend to be developed by researchers with much more powerful machines. When the size of the data that must be accessed is larger than the size of memory, some form of virtual memory is simply required. This may be by segmentation, paging, or by paged segments. In this paper we demonstrate that complete reliance on operating system virtual memory for out-of-core visualization leads to poor performance. We then describe a paged segment system that we have implemented, and explore the principles of memory management that can be employed by the application for out-of-core visualization. We show that application control over some of these can significantly improve performance. We show that sparse traversal can be exploited by loading only those data actually required. We show also that application control over data loading can be exploited by 1) loading data from alternative storage format (in particular 3-dimensional data stored in sub-cubes), 2) controlling the page size. Both of these techniques effectively reduce the total memory required by visualization at run-time. We also describe experiments we have done on remote out-of-core visualization (when pages are read by demand from remote disk) whose results are promising.

Cox, Michael↗

Crater Morphometry and Crater Degradation on Mercury: Mercury Laser Altimeter (MLA) Measurements and Comparison to Stereo-DTM Derived Results

Two types of measurements of Mercury's surface topography were obtained by the MESSENGER (MErcury Surface Space ENvironment, GEochemisty and Ranging) spacecraft: laser ranging data from Mercury Laser Altimeter (MLA) [1], and stereo imagery from the Mercury Dual Imaging System (MDIS) camera [e.g., 2, 3]. MLA data provide precise and accurate elevation meaurements, but with sparse spatial sampling except at the highest northern latitudes. Digital terrain models (DTMs) from MDIS have superior resolution but with less vertical accuracy, limited approximately to the pixel resolution of the original images (in the case of [3], 15-75 m). Last year [4], we reported topographic measurements of craters in the D=2.5 to 5 km diameter range from stereo images and suggested that craters on Mercury degrade more quickly than on the Moon (by a factor of up to approximately 10×). However, we listed several alternative explanations for this finding, including the hypothesis that the lower depth/diameter ratios we observe might be a result of the resolution and accuracy of the stereo DTMs. Thus, additional measurements were undertaken using MLA data to examine the morphometry of craters in this diameter range and assess whether the faster crater degradation rates proposed to occur on Mercury is robust.

morphometry↗

Hyperspectral segmentation of plants in fabricated ecosystems

Hyperspectral imaging provides a powerful tool for analyzing above-ground plant characteristics in fabricated ecosystems, offering rich spectral information across diverse wavelengths. This study presents an efficient workflow for hyperspectral data segmentation and subsequent data analytics, minimizing the need for user annotation through the use of ensembles of sparse mixed scale convolution neural networks. The segmentation process leverages the diversity of ensembles to achieve high accuracy with minimal labeled data, reducing labor-intensive annotation efforts. To further enhance robustness, we incorporate image alignment techniques to address spatial variability in the dataset. Downstream analysis focuses on using the segmented data for processing spectral data, enabling monitoring of plant health. This approach provides a scalable solution for spectral segmentation, and facilitates actionable insights into plant conditions in complex, controlled environments. Our results demonstrate the utility of combining advanced machine learning techniques with hyperspectral analytics for high-throughput plant monitoring.

Zwart, Petrus H.↗

Locating the Isolator Shock-Train Leading Edge with Limited Pressure Information

Real-time detection and control of the isolator shock-train leading edge (STLE) is important to the performance of high-speed air-breathing engines, such as dual-mode scramjets. Typically, the STLE location is determined using wall static-pressure measurements, but there are often restrictions on the placement and overall number of the pressure transducers, reducing the viability and accuracy of such approaches. To address these issues, we introduce the adaptive pressure profile (APP) method for estimating the STLE location. This method does not require extensive prior characterization of the isolator or engine model. Instead, it uses real-time pressure measurements from a small number of transducers to adaptively learn the isolator pressure profile and subsequently uses this deduced profile to estimate the STLE location in a data-driven manner. The APP method works well in situations with sparse transducer placement. It produces accurate estimates when the STLE location is 1) not bounded by two or more transducers or 2) between two transducers that are several isolator duct heights apart. We demonstrate the efficacy of the APP method using simulations and experimental data from direct-connect isolator models. This validation shows that the APP method is accurate and robust for different flow regimes, transducer configurations, and model geometries.

Gregory J. Hunt↗

Computing Gravitational Bumps From Repeating-Orbit Data

Iterative, least-squares algorithm efficiently computes estimates of both position errors indicative of irregularities in gravitational field of Earth and trajectory of satellite in orbit repeating along same ground track. Exploits sparse-matrix techniques. Useful in surveying, navigation, and geophysical research. Particularly useful for processing data on trajectory of satellite in low orbit tracked via Global Positioning System (GPS).

Wu, Jiun-Tsong↗

Linear absorption coefficient of beryllium in the 50-300-A wavelength range

Transmittances of thin-film filters fabricated for an extreme-UV astronomy sounding-rocket experiment yield values for the linear absorption coefficient of beryllium in the 50-300-A wavelength range, in which previous measurements are sparse. The inferred values are consistent with the lowest data previously published and may have important consequences for extreme-UV astronomers.

Barstow, M. A.↗

Pressure dependence of the absolute rate constant for the reaction Cl + C2H2 from 210-361 K

In recent years, considerable attention has been given to the role of chlorine compounds in the catalytic destruction of stratospheric ozone. However, while some reactions have been studied extensively, the kinetic data for the reaction of Cl with C2H2 is sparse with only three known determinations of the rate constant k3. The reactions involved are Cl + C2H2 yields reversibly ClC2H2(asterisk) (3a) and ClC2H2(asterisk) + M yields ClC2H2 + M (3b). In the present study, flash photolysis coupled with chlorine atomic resonance fluorescence have been employed to determine the pressure and temperature dependence of k3 with the third body M = Ar. Room temperature values are also reported for M = N2. The pressure dependence observed in the experiments confirms the expectation that the reaction involves addition of Cl to the unsaturated C2H2 molecule followed by collisional stabilization of the resulting adduct radical.

Brunning, J.↗

Comparison of two matrix data structures for advanced CSM testbed applications

The first section describes data storage schemes presently used by the Computational Structural Mechanics (CSM) testbed sparse matrix facilities and similar skyline (profile) matrix facilities. The second section contains a discussion of certain features required for the implementation of particular advanced CSM algorithms, and how these features might be incorporated into the data storage schemes described previously. The third section presents recommendations, based on the discussions of the prior sections, for directing future CSM testbed development to provide necessary matrix facilities for advanced algorithm implementation and use. The objective is to lend insight into the matrix structures discussed and to help explain the process of evaluating alternative matrix data structures and utilities for subsequent use in the CSM testbed.

Regelbrugge, M. E.↗

Statistical analysis of power-size-redshift distributions of extragalactic jets

This paper investigates whether a hot, sparse, yet cosmologically significant intergalactic medium is consistent with data collected from extragalactic radio sources. This is done by use of Monte Carlo simulations which employ previously run pseudohydrodynamical simulations to cover an observational parameter space. These observational parameters include the scale height, central density, and temperature of a (isothermal) galactic halo, and the power of the central engine which drives the jet. The Monte Carlo simulations generate distribution of sizes in bins of (received) power and redshift, which have been compared with observational data using Kolmogorov-Smirnov tests. Results of this analysis are consistent with the existence of an IGM with temperature and density mentioned above. In addition, this analysis suggests that the active lifetime of powerful extragalactic radio sources decreases with increasing power.

Rosen, Alexander↗

Towards an Automated Classification of Transient Events in Synoptic Sky Surveys

We describe the development of a system for an automated, iterative, real-time classification of transient events discovered in synoptic sky surveys. The system under development incorporates a number of Machine Learning techniques, mostly using Bayesian approaches, due to the sparse nature, heterogeneity, and variable incompleteness of the available data. The classifications are improved iteratively as the new measurements are obtained. One novel featrue is the development of an automated follow-up recommendation engine, that suggest those measruements that would be the most advantageous in terms of resolving classification ambiguities and/or characterization of the astrophysically most interesting objects, given a set of available follow-up assets and their cost funcations. This illustrates the symbiotic relationship of astronomy and applied computer science through the emerging disciplne of AstroInformatics.

classification↗

Central Valley Water Resources: Improving California Groundwater Assessments using GRACE and InSAR Datasets for Water Resource Management

California’s Central Valley is one of the most productive agricultural areas in the world, producing approximately $20 billion in crops annually. The recent California droughts of 2007-2010 and 2012-2019 resulted in increased groundwater pumping in the Central Valley to adequately irrigate farmland. Overdrafting of the Central Valley aquifer results in groundwater depletion, land subsidence, and permanent loss of groundwater storage. In 2014, depletion of groundwater led the state of California to enact the Sustainable Groundwater Management Act (SGMA),requiring critically overdrafted, high, and medium priority sub-basins to reach sustainable levels of groundwater pumping and recharge by 2042. SGMA allows local Groundwater Sustainability Agencies (GSAs) the authority to create Groundwater Sustainability Plans (GSPs) at the sub-basin level. To assist California’s Department of Water Resources (DWR), this project quantified groundwater change and land subsidence in Central Valley sub-basins with sparse or unreliable well and Geographic Positioning Systems (GPS) data. This was done using NASA’s Gravity Recovery and Climate Experiment (GRACE), GRACE FollowOn (GRACE-FO), and interferograms derived from Sentinel-1 C-band Synthetic Aperture Radar (C-SAR) and Advanced Land Observing Satellite 2 (ALOS-2)Phased Array L-band Synthetic Aperture Radar 2 (PALSAR-2). Time series of the GRACE and InSAR data were compared with well and GPS data in data-dense sub-basins to determine the feasibility of these datasets for groundwater storage and subsidence monitoring. We found that GRACE and InSAR data are effective tools for determining groundwater change and land subsidence and can be used on their own to monitor sub-basins in the absence of well and GPS data

Water Resources↗

Daily Ambient Air Pollution Metrics for Five Cities: Evaluation of Data Fusion-Based Estimates and Uncertainties

Spatiotemporal characterization of ambient air pollutant concentrations is increasingly relying on the combination of observations and air quality models to provide well-constrained, spatially and temporally complete pollutant concentration fields. Air quality models, in particular, are attractive, as they characterize the emissions, meteorological, and physiochemical process linkages explicitly while providing continuous spatial structure. However, such modeling is computationally intensive and has biases. The limitations of spatially sparse and temporally incomplete observations can be overcome by blending the data with estimates from a physically and chemically coherent model, driven by emissions and meteorological inputs. We recently developed a data fusion method that blends ambient ground observations and chemical transport-modeled (CTM) data to estimate daily, spatially resolved pollutant concentrations and associatedcorrelations. In this study, we assess the ability of the data fusion method to produce daily metrics (i.e., 1-hr max, 8-hr max, and 24-hr average) of ambient air pollution that capture spatiotemporal air pollution trends for 12 pollutants (CO, NO2, NOx, O3, SO2, PM (sub10), PM (sub 2.5), and five PM (sub 2.5) components) across five metropolitan areas (Atlanta, Birmingham, Dallas, Pittsburgh, and St. Louis), from 2002 to 2008. Three sets of comparisons are performed: (1) the CTM concentrations are evaluated for each pollutant and metropolitan domain, (2) the data fusion concentrations are compared with the monitor data, (3) a comprehensive cross-validation analysis against observed data evaluates the quality of the data fusion model simulations across multiple metropolitan domains. The resulting daily spatial field estimates of air pollutant concentrations and uncertainties are not only consistent with observations, emissions, andmeteorology, but substantially improve CTM-derived results for nearly all pollutants and all cities, with the exception of NO2 for Birmingham. The greatest improvements occur for O3 and PM (sub 2.5). Squared spatiotemporal correlation coefficients range between simulations and observations determined using cross-validation across all cities for air pollutants of secondary and mixed origins are R-squared equal to 0.88-0.93 (O3), 0.81-0.89 (SO4), 0.67-0.83 (PM (sub 2.5)), 0.52-0.72 (NO3), 0.43-0.80 (NH4), 0.32-0.51 (OC), and 0.14-0.71 (PM (sub 10)). Results for relatively homogeneous pollutants of secondary origin, tend to be better than those for more spatially heterogeneous (larger spatial gradients) pollutants of primary origin (NOx, CO, SO2 and EC). Generally, background concentrations and spatial concentration gradients reflect interurban airshed complexity and the effects of regional transport, whereas daily spatial pattern variability shows intra urban consistency in the fused data. With sufficiently high CTM spatial resolution, traffic-related pollutants exhibit gradual concentration gradients that peak toward the urban centers. Ambient pollutant concentration uncertainty estimates for the fused data are both more accurate and smaller than those for either the observations or the model simulations alone.

spatiotemporal fusion↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗