Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Objective sea-level pressure analysis for data-sparse areas.

A computer procedure, described in an earlier study, uses the wind speed field near the ocean surface in combination with a small number of observations of pressure and wind velocity to specify the maritime sea-level pressure field. An improved version was used to analyze the pressure distribution over the North Pacific Ocean for eleven synoptic times in February 1967. Independent knowledge of the central pressures of lows is shown to reduce the analysis errors for very sparse data coverage. The application of planned remote sensing of sea-level wind speeds is shown to make a significant contribution to the quality of the analysis especially in the high gradient midlatitudes and for sparse coverage of conventional observations (such as over Southern Hemisphere oceans). Uniform distribution of the available observations of sea-level pressure and wind velocity yields results far superior to those derived from a random distribution.

Druyan, L. M.↗

Learning control for robotic manipulators with sparse data

Learning control algorithms have been proposed for error compensation in repetitive robotic manipulator tasks. It is shown that the performance of such control algorithms can be seriously degraded when the feedback data they use is relatively sparse in time, such as might be provided by vision systems. It is also shown that learning control algorithms can be modified to compensate for the effects of sparse data and thereby yield performance which approaches that of systems without limitations on the sensory information available for control.

Morita, Atsushi↗

3′ RNA-seq is superior to standard RNA-seq in cases of sparse data but inferior at identifying toxicity pathways in a model organism

The application of RNA-sequencing has led to numerous breakthroughs related to investigating gene expression levels in complex biological systems. Among these are knowledge of how organisms, such as the vertebrate model organism zebrafish (Danio rerio), respond to toxicant exposure. Recently, the development of 3' RNA-seq has allowed for the determination of gene expression levels with a fraction of the required reads compared to standard RNA-seq. While 3' RNA-seq has many advantages, a comparison to standard RNA-seq has not been performed in the context of whole organism toxicity and sparse data. Here, we examined samples from zebrafish exposed to perfluorobutane sulfonamide (FBSA) with either 3' or standard RNA-seq to determine the advantages of each with regards to the identification of functionally enriched pathways. We found that 3' and standard RNA-seq showed specific advantages when focusing on annotated or unannotated regions of the genome. We also found that standard RNA-seq identified more differentially expressed genes (DEGs), but that this advantage disappeared under conditions of sparse data. We also found that standard RNA-seq had a significant advantage in identifying functionally enriched pathways via analysis of DEG lists but that this advantage was minimal when identifying pathways via gene set enrichment analysis of all genes. These results show that each approach has experimental conditions where they may be advantageous. Our observations can help guide others in the choice of 3' RNA-seq vs standard RNA sequencing to query gene expression levels in a range of biological systems.

3’ RNA-seq↗

Source scaling comparison and validation in Central Italy: data intensive direct S waves versus the sparse data coda envelope methodology

SUMMARY Robustness of source parameter estimates is a fundamental issue in understanding the relationships between small and large events; however, it is difficult to assess how much of the variability of the source parameters can be attributed to the physical source characteristics or to the uncertainties of the methods and data used to estimate the values. In this study, we apply the coda method by Mayeda et al. using the coda calibration tool (CCT), a freely available Java-based code (https://github.com/LLNL/coda-calibration-tool) to obtain a regional calibration for Central Italy for estimating stable source parameters. We demonstrate the power of the coda technique in this region and show that it provides the same robustness in source parameter estimation as a data-driven methodology [generalized inversion technique (GIT)], but with much fewer calibration events and stations. The Central Italy region is ideal for both GIT and coda approaches as it is characterized by high-quality data, including recent well-recorded seismic sequences such as L'Aquila (2009) and Amatrice–Norcia–Visso (2016–2017). This allows us to apply data-driven methods such as GIT and coda-based methods that require few, but high-quality data. The data set for GIT analysis includes ∼5000 earthquakes and more than 600 stations, while for coda analysis we used a small subset of 39 events spanning 3.5 < Mw < 6.33 and 14 well-distributed broad-band stations. For the common calibration events, as well as an additional 247 events (∼1.7 < Mw < ∼5.0) not used in either calibration, we find excellent agreement between GIT-derived and CCT-derived source spectra. This confirms the ability of the coda approach to obtain stable source parameters even with few calibration events and stations. Even reducing the coda calibration data set by 75 per cent, we found no appreciable degradation in performance. This validation of the coda calibration approach over a broad range of event size demonstrates that this procedure, once extended to other regions, represents a powerful tool for future routine applications to homogeneously evaluate robust source parameters on a national scale. Furthermore, the coda calibration procedure can homogenize the Mw estimates for small and large events without the necessity of introducing any conversion scale between narrow-band measures such as local magnitude (ML) and Mw, which has been shown to introduce significant bias.

Morasca, Paola (ORCID:0000000265254867)↗

Representation-Independent Iteration of Sparse Data Arrays

An approach is defined that describes a method of iterating over massively large arrays containing sparse data using an approach that is implementation independent of how the contents of the sparse arrays are laid out in memory. What is unique and important here is the decoupling of the iteration over the sparse set of array elements from how they are internally represented in memory. This enables this approach to be backward compatible with existing schemes for representing sparse arrays as well as new approaches. What is novel here is a new approach for efficiently iterating over sparse arrays that is independent of the underlying memory layout representation of the array. A functional interface is defined for implementing sparse arrays in any modern programming language with a particular focus for the Chapel programming language. Examples are provided that show the translation of a loop that computes a matrix vector product into this representation for both the distributed and not-distributed cases. This work is directly applicable to NASA and its High Productivity Computing Systems (HPCS) program that JPL and our current program are engaged in. The goal of this program is to create powerful, scalable, and economically viable high-powered computer systems suitable for use in national security and industry by 2010. This is important to NASA for its computationally intensive requirements for analyzing and understanding the volumes of science data from our returned missions.

James, Mark↗

An Automated Scanning Transmission Electron Microscope Guided by Sparse Data Analytics

Abstract Artificial intelligence (AI) promises to reshape scientific inquiry and enable breakthrough discoveries in areas such as energy storage, quantum computing, and biomedicine. Scanning transmission electron microscopy (STEM), a cornerstone of the study of chemical and materials systems, stands to benefit greatly from AI-driven automation. However, present barriers to low-level instrument control, as well as generalizable and interpretable feature detection, make truly automated microscopy impractical. Here, we discuss the design of a closed-loop instrument control platform guided by emerging sparse data analytics. We hypothesize that a centralized controller, informed by machine learning combining limited a priori knowledge and task-based discrimination, could drive on-the-fly experimental decision-making. This platform may unlock practical, automated analysis of a variety of material features, enabling new high-throughput and statistical studies.

47 OTHER INSTRUMENTATION↗

Efficient mapping between void shapes and stress fields using Deep Convolutional Neural Networks with sparse data

Establishing fast and accurate structure-to-property relationships is an important component in the design and discovery of advanced materials. Physics-based simulation models like the finite element method (FEM) are often used to predict deformation, stress, and strain fields as a function of material microstructure in material and structural systems. Such models may be computationally expensive and time intensive if the underlying physics of the system is complex. This limits their application to solve inverse design problems and identify structures that maximize performance. In such scenarios, surrogate models are employed to make the forward mapping computationally efficient to evaluate. However, the high dimensionality of the input microstructure and the output field of interest often renders such surrogate models inefficient, especially when dealing with sparse data. Deep convolutional neural network (CNN) based surrogate models have shown great promise in handling such high-dimensional problems. In this paper, a single ellipsoidal void structure under a uniaxial tensile load represented by a linear elastic, high-dimensional and expensive-to-query, FEM model. We consider two deep CNN architectures, a modified convolutional autoencoder framework with a fully connected bottleneck and a UNet CNN, and compare their accuracy in predicting the von Mises stress field for any given input void shape in the FEM model. Additionally, a sensitivity analysis study is performed using the two approaches, where the variation in the prediction accuracy on unseen test data is studied through numerical experiments by varying the number of training samples from 20 to 100.

surrogate modeling; convolutional neural networks;↗

First experimental study of multiple orientation muon tomography, with image optimization in sparse data environments

Due to the high penetrating power of cosmic ray muons, they can be used to probe very thick and dense objects. As charged particles, they can be tracked by ionization detectors, determining the position and direction of the muons. With detectors on either side of an object, particle direction changes can be used to extract scattering information within an object. This can be used to produce a scattering intensity image within the object related to density and atomic number. Such imaging is typically performed with a single detector-object orientation, taking advantage of the more intense downward flux of muons, producing planar imaging with some depth-of-field information in the third dimension. Several simulation studies have been published with multi-orientation tomography, which can form a three-dimensional representation faster than a single orientation view. In this work we present the first experimental multiple orientation muon tomography study. Experimental muon-scatter based tomography was performed using a concrete filled steel drum with several different metal wedges inside, between detector planes. Data was collected from different detector-object orientations by rotating the steel drum. The data collected from each orientation were then combined using two different tomographic methods. Results showed that using a combination of multiple depth-of-field reconstructions, rather than a traditional inverse Radon transform approach used for CT, resulted in more useful images for sparser data. As cosmic ray muon flux imaging is rate limited, the imaging techniques were compared for sparse data. Using the combined depth-of-field reconstruction technique, fewer detector-object orientations were needed to reconstruct images that could be used to differentiate the metal wedge compositions.

Applied Physics (physics.app-ph)↗

Experimental study of multiple-orientation muon tomography with image optimization in sparse data environments

Due to the high penetrating power of cosmic-ray muons, they can be used to probe very thick and dense objects. As muons are charged particles, they can be tracked by ionization detectors, determining the position and direction of the muons. With detectors on either side of an object to measure particle direction change, scattering information within the object can be found. This can be used to produce a scattering-intensity image within the object related to density and atomic number. Such imaging is typically performed with a single detector-object orientation, taking advantage of the more intense downward flux of muons, producing planar imaging with some depth-of-field information in the third dimension. Several simulation studies were published with multiorientation tomography, which can form a three-dimensional representation faster than a single-orientation view. In this study, experimental muon-scatter-based tomography was performed using a concrete filled steel drum with several different metal wedges inside, with the drum between detector planes. Data were collected from different detector-object orientations by rotating the steel drum. The data collected from each orientation were combined using two different tomographic methods. A traditional inverse Radon transform approach used for computed tomography and a combination of multiple depth-of-field reconstructions were applied to the data. As cosmic-ray muon flux imaging is rate limited, the imaging techniques were compared for sparse data. Using the combined depth-of-field reconstruction technique, fewer detector-object orientations were needed to reconstruct images that could be used to differentiate the metal wedges.

47 OTHER INSTRUMENTATION↗

An Efficient, Scalable IO Framework for Sparse Data: larcv3

Neutrino physics is one of the fundamental areas of research into the origins and properties of the Universe. Many experimental neutrino projects use sophisticated detectors to observe properties of these particles, and have turned to deep learning and artificial intelligence techniques to analyze their data. From this, we have developed \texttt{larcv}, a \texttt{C++} and \texttt{Python} based framework for efficient IO of sparse data with particle physics applications in mind. We describe in this paper the \texttt{larcv} framework and some benchmark IO performance tests. \texttt{larcv} is designed to enable fast and efficient IO of ragged and irregular data, at scale on modern HPC systems, and is compatible with the most popular open source data analysis tools in the Python ecosystem.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Objective sea level pressure analysis for sparse data areas

A computer procedure was used to analyze the pressure distribution over the North Pacific Ocean for eleven synoptic times in February, 1967. Independent knowledge of the central pressures of lows is shown to reduce the analysis errors for very sparse data coverage. The application of planned remote sensing of sea-level wind speeds is shown to make a significant contribution to the quality of the analysis especially in the high gradient mid-latitudes and for sparse coverage of conventional observations (such as over Southern Hemisphere oceans). Uniform distribution of the available observations of sea-level pressure and wind velocity yields results far superior to those derived from a random distribution. A generalization of the results indicates that the average lower limit for analysis errors is between 2 and 2.5 mb based on the perfect specification of the magnitude of the sea-level pressure gradient from a known verification analysis. A less than perfect specification will derive from wind-pressure relationships applied to satellite observed wind speeds.

Druyan, L. M.↗

Sparse-Data Deep Learning Strategies for Radiographic Non-Destructive Testing

Radiography is an imaging technique used in a variety of applications, such as medical diagnosis, airport security, and nondestructive testing. We present a deep learning system for extracting information from radiographic images. We perform various prediction tasks using our system, including material classification and regression on the dimensions of a given object that is being radiographed. Our system is designed to address the sparse-data issue for radiographic nondestructive testing applications. It uses a radiographic simulation tool for synthetic data augmentation, and it uses transfer learning with a pre-trained convolutional neural network model. Using this system, our preliminary results indicate that the object geometry regression task saw an improvement of 70% in the R-squared value when using a multi-regime model. In addition, we increase the performance of the object material classification tasks by utilizing data from different imaging systems. In particular, using neutron imaging improved the material classification accuracy by 20% when compared to x-ray imaging.

convolutional neural networks↗

Automatic Calibration of a Geomechanical Model from Sparse Data for Estimating Stress in Deep Geological Formations

Summary In this study, we demonstrate geomechanical modeling with fully automatic parameter calibration to estimate the full geomechanical stress fields of a prospective US carbon dioxide (CO2) storage site, based on sparse measurement data. The goal is to compute full stress tensor field estimates (principal stresses and orientations) that are maximally compatible with observations within the constraints of the model assumptions, thereby extending pointwise, incomplete partial stress measurement to a simulated full formation stress field, as well as a rough assessment of the associated error. We use the Perch site, located in Otsego County, Michigan, USA, as our case study. The input data consist of partial stress tensor information inferred from in-situ borehole tests, geophysical well logs, and processing of seismic data. A static earth model (SEM) of the site was developed, and geomechanical simulation functionality of the open-source MATLAB Reservoir Simulation Toolbox (MRST) was used to model the stress field. Adjoint-based nonlinear optimization was used to adjust boundary conditions and material properties to calibrate simulated results of observations. Results were interpreted through a Bayesian framework. The focus of this paper is to demonstrate how the fully automatic calibration procedure works and discuss the results obtained; it does not attempt a detailed analysis of the stress field in the context of the proposed CO2 storage initiatives. Our work is part of a larger effort to noninvasively determine in-situ stresses in deep formations considered for CO2 storage. Guided by previously published research on geomechanical model calibration, our work presents a novel calibration approach supporting a potentially large number of linear or nonlinear calibration parameters to produce results optimally agreeing with available measurements and thus extend partial pointwise estimates to full tensor fields compatible with the physics of the site.

Engineering↗

Providing Data Access and Analysis Capabilities to SERVIR’s Data-Sparse Regions

In developing regions of the world, the communications infrastructure pose enormous challenges for using Earth observation data. Limited internet bandwidth along with the high costs make it almost impossible to process and extract zonal statistics over large periods of time for even small geographic areas. In such cases, downloading daily rainfall data or dekadal series of NDVI data would take days and consume all the bandwidth allocated to an organization (for reference, internet connections in Niger would cost thousands of dollars per month at a maximum - and unreliable - bandwidth of just 10 Mbps). Running crop models or hydrological models typically require several years of historic data over the area of interest (AOI). In some cases, these AOIs are relatively small compared to the footprint of individual earth observation granules. Hence, systems that let the stakeholders subset the data to download to a user specified area, or even submit processing requests that let them download small result files for the AOI become critical. The SERVIR program has developed a tool to provide this type of access to help decision makers in developing regions use long time series of adjusted rainfall data (CHIRPS), NDVI values, seasonal weather forecasts, evaporative stress indices and others in a very efficient manner. This system, named ClimateSERV (https://climateserv.servirglobal.net) ingests the datasets in an automated fashion and allows interactive access (through a web application), or automated access through a simple API that developers can quickly incorporate in independent applications. This way, the extraction of daily averages of rainfall over a 50 square Km area through 30 years of archived data takes only a few seconds to process, and the results can be presented on an online chart or downloaded in a comma separated file that's only a few Kb.

Ashmall, William↗

Satellite-derived synoptic climatology in data-sparse regions

Synoptic-scale 'moisture bursts' are defined, based on infrared GOES imagery, and their synoptic climatology is developed. Quantitative analysis of satellite-derived individual channel radiance data and vertical eigenfunctions of complete channel data yield rich structural detail; these details do not appear in FGGE analyses in regions void of conventional meteorological data.

Mcguirk, J. P.↗

Building Alternative Indices of Socioeconomic Status for Population Modeling in Data-Sparse Contexts (Short Paper)

Population modeling requires clear definitions of socioeconomic status (SES) to ensure overall estimate accuracy and locate potentially underserved subpopulations. This presents a challenge as SES can be measured in myriad ways and for divergent purposes, and the data required to calculate these metrics may be lacking, particularly in low and middle income countries (LMICs). To support more refined SES measurement, we explore improvements upon the Demographic and Health Survey’s (DHS) Wealth Index (DHS-WI) using alternative characterizations of SES based on multiple correspondence analysis (MCA) and hierarchical clustering. We produce the MCA-derived metrics first on a full suite of household economic, demographic, and dwelling variables, then on a reduced set of occupant-only SES characteristics. We explore the utility of these metrics relative to DHS-WI based on their ability to 1) differentiate DHS household types and 2) identify mixtures of SES levels within DHS samples and mapped at high spatial resolution. We find that our full suite MCA yields more clearly defined SES segments and that our reduced MCA delineates occupant SES most clearly, suggesting potential pathways to improve upon the DHS-WI in future population modeling efforts for LMICs.

Cunningham, Angela↗