Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

3′ RNA-seq is superior to standard RNA-seq in cases of sparse data but inferior at identifying toxicity pathways in a model organism

The application of RNA-sequencing has led to numerous breakthroughs related to investigating gene expression levels in complex biological systems. Among these are knowledge of how organisms, such as the vertebrate model organism zebrafish (Danio rerio), respond to toxicant exposure. Recently, the development of 3' RNA-seq has allowed for the determination of gene expression levels with a fraction of the required reads compared to standard RNA-seq. While 3' RNA-seq has many advantages, a comparison to standard RNA-seq has not been performed in the context of whole organism toxicity and sparse data. Here, we examined samples from zebrafish exposed to perfluorobutane sulfonamide (FBSA) with either 3' or standard RNA-seq to determine the advantages of each with regards to the identification of functionally enriched pathways. We found that 3' and standard RNA-seq showed specific advantages when focusing on annotated or unannotated regions of the genome. We also found that standard RNA-seq identified more differentially expressed genes (DEGs), but that this advantage disappeared under conditions of sparse data. We also found that standard RNA-seq had a significant advantage in identifying functionally enriched pathways via analysis of DEG lists but that this advantage was minimal when identifying pathways via gene set enrichment analysis of all genes. These results show that each approach has experimental conditions where they may be advantageous. Our observations can help guide others in the choice of 3' RNA-seq vs standard RNA sequencing to query gene expression levels in a range of biological systems.

3’ RNA-seq↗

Source scaling comparison and validation in Central Italy: data intensive direct S waves versus the sparse data coda envelope methodology

SUMMARY Robustness of source parameter estimates is a fundamental issue in understanding the relationships between small and large events; however, it is difficult to assess how much of the variability of the source parameters can be attributed to the physical source characteristics or to the uncertainties of the methods and data used to estimate the values. In this study, we apply the coda method by Mayeda et al. using the coda calibration tool (CCT), a freely available Java-based code (https://github.com/LLNL/coda-calibration-tool) to obtain a regional calibration for Central Italy for estimating stable source parameters. We demonstrate the power of the coda technique in this region and show that it provides the same robustness in source parameter estimation as a data-driven methodology [generalized inversion technique (GIT)], but with much fewer calibration events and stations. The Central Italy region is ideal for both GIT and coda approaches as it is characterized by high-quality data, including recent well-recorded seismic sequences such as L'Aquila (2009) and Amatrice–Norcia–Visso (2016–2017). This allows us to apply data-driven methods such as GIT and coda-based methods that require few, but high-quality data. The data set for GIT analysis includes ∼5000 earthquakes and more than 600 stations, while for coda analysis we used a small subset of 39 events spanning 3.5 < Mw < 6.33 and 14 well-distributed broad-band stations. For the common calibration events, as well as an additional 247 events (∼1.7 < Mw < ∼5.0) not used in either calibration, we find excellent agreement between GIT-derived and CCT-derived source spectra. This confirms the ability of the coda approach to obtain stable source parameters even with few calibration events and stations. Even reducing the coda calibration data set by 75 per cent, we found no appreciable degradation in performance. This validation of the coda calibration approach over a broad range of event size demonstrates that this procedure, once extended to other regions, represents a powerful tool for future routine applications to homogeneously evaluate robust source parameters on a national scale. Furthermore, the coda calibration procedure can homogenize the Mw estimates for small and large events without the necessity of introducing any conversion scale between narrow-band measures such as local magnitude (ML) and Mw, which has been shown to introduce significant bias.

Morasca, Paola (ORCID:0000000265254867)↗

An Automated Scanning Transmission Electron Microscope Guided by Sparse Data Analytics

Abstract Artificial intelligence (AI) promises to reshape scientific inquiry and enable breakthrough discoveries in areas such as energy storage, quantum computing, and biomedicine. Scanning transmission electron microscopy (STEM), a cornerstone of the study of chemical and materials systems, stands to benefit greatly from AI-driven automation. However, present barriers to low-level instrument control, as well as generalizable and interpretable feature detection, make truly automated microscopy impractical. Here, we discuss the design of a closed-loop instrument control platform guided by emerging sparse data analytics. We hypothesize that a centralized controller, informed by machine learning combining limited a priori knowledge and task-based discrimination, could drive on-the-fly experimental decision-making. This platform may unlock practical, automated analysis of a variety of material features, enabling new high-throughput and statistical studies.

47 OTHER INSTRUMENTATION↗

Efficient mapping between void shapes and stress fields using Deep Convolutional Neural Networks with sparse data

Establishing fast and accurate structure-to-property relationships is an important component in the design and discovery of advanced materials. Physics-based simulation models like the finite element method (FEM) are often used to predict deformation, stress, and strain fields as a function of material microstructure in material and structural systems. Such models may be computationally expensive and time intensive if the underlying physics of the system is complex. This limits their application to solve inverse design problems and identify structures that maximize performance. In such scenarios, surrogate models are employed to make the forward mapping computationally efficient to evaluate. However, the high dimensionality of the input microstructure and the output field of interest often renders such surrogate models inefficient, especially when dealing with sparse data. Deep convolutional neural network (CNN) based surrogate models have shown great promise in handling such high-dimensional problems. In this paper, a single ellipsoidal void structure under a uniaxial tensile load represented by a linear elastic, high-dimensional and expensive-to-query, FEM model. We consider two deep CNN architectures, a modified convolutional autoencoder framework with a fully connected bottleneck and a UNet CNN, and compare their accuracy in predicting the von Mises stress field for any given input void shape in the FEM model. Additionally, a sensitivity analysis study is performed using the two approaches, where the variation in the prediction accuracy on unseen test data is studied through numerical experiments by varying the number of training samples from 20 to 100.

surrogate modeling; convolutional neural networks;↗

First experimental study of multiple orientation muon tomography, with image optimization in sparse data environments

Due to the high penetrating power of cosmic ray muons, they can be used to probe very thick and dense objects. As charged particles, they can be tracked by ionization detectors, determining the position and direction of the muons. With detectors on either side of an object, particle direction changes can be used to extract scattering information within an object. This can be used to produce a scattering intensity image within the object related to density and atomic number. Such imaging is typically performed with a single detector-object orientation, taking advantage of the more intense downward flux of muons, producing planar imaging with some depth-of-field information in the third dimension. Several simulation studies have been published with multi-orientation tomography, which can form a three-dimensional representation faster than a single orientation view. In this work we present the first experimental multiple orientation muon tomography study. Experimental muon-scatter based tomography was performed using a concrete filled steel drum with several different metal wedges inside, between detector planes. Data was collected from different detector-object orientations by rotating the steel drum. The data collected from each orientation were then combined using two different tomographic methods. Results showed that using a combination of multiple depth-of-field reconstructions, rather than a traditional inverse Radon transform approach used for CT, resulted in more useful images for sparser data. As cosmic ray muon flux imaging is rate limited, the imaging techniques were compared for sparse data. Using the combined depth-of-field reconstruction technique, fewer detector-object orientations were needed to reconstruct images that could be used to differentiate the metal wedge compositions.

Applied Physics (physics.app-ph)↗

Experimental study of multiple-orientation muon tomography with image optimization in sparse data environments

Due to the high penetrating power of cosmic-ray muons, they can be used to probe very thick and dense objects. As muons are charged particles, they can be tracked by ionization detectors, determining the position and direction of the muons. With detectors on either side of an object to measure particle direction change, scattering information within the object can be found. This can be used to produce a scattering-intensity image within the object related to density and atomic number. Such imaging is typically performed with a single detector-object orientation, taking advantage of the more intense downward flux of muons, producing planar imaging with some depth-of-field information in the third dimension. Several simulation studies were published with multiorientation tomography, which can form a three-dimensional representation faster than a single-orientation view. In this study, experimental muon-scatter-based tomography was performed using a concrete filled steel drum with several different metal wedges inside, with the drum between detector planes. Data were collected from different detector-object orientations by rotating the steel drum. The data collected from each orientation were combined using two different tomographic methods. A traditional inverse Radon transform approach used for computed tomography and a combination of multiple depth-of-field reconstructions were applied to the data. As cosmic-ray muon flux imaging is rate limited, the imaging techniques were compared for sparse data. Using the combined depth-of-field reconstruction technique, fewer detector-object orientations were needed to reconstruct images that could be used to differentiate the metal wedges.

47 OTHER INSTRUMENTATION↗

An Efficient, Scalable IO Framework for Sparse Data: larcv3

Neutrino physics is one of the fundamental areas of research into the origins and properties of the Universe. Many experimental neutrino projects use sophisticated detectors to observe properties of these particles, and have turned to deep learning and artificial intelligence techniques to analyze their data. From this, we have developed \texttt{larcv}, a \texttt{C++} and \texttt{Python} based framework for efficient IO of sparse data with particle physics applications in mind. We describe in this paper the \texttt{larcv} framework and some benchmark IO performance tests. \texttt{larcv} is designed to enable fast and efficient IO of ragged and irregular data, at scale on modern HPC systems, and is compatible with the most popular open source data analysis tools in the Python ecosystem.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Sparse-Data Deep Learning Strategies for Radiographic Non-Destructive Testing

Radiography is an imaging technique used in a variety of applications, such as medical diagnosis, airport security, and nondestructive testing. We present a deep learning system for extracting information from radiographic images. We perform various prediction tasks using our system, including material classification and regression on the dimensions of a given object that is being radiographed. Our system is designed to address the sparse-data issue for radiographic nondestructive testing applications. It uses a radiographic simulation tool for synthetic data augmentation, and it uses transfer learning with a pre-trained convolutional neural network model. Using this system, our preliminary results indicate that the object geometry regression task saw an improvement of 70% in the R-squared value when using a multi-regime model. In addition, we increase the performance of the object material classification tasks by utilizing data from different imaging systems. In particular, using neutron imaging improved the material classification accuracy by 20% when compared to x-ray imaging.

convolutional neural networks↗

Automatic Calibration of a Geomechanical Model from Sparse Data for Estimating Stress in Deep Geological Formations

Summary In this study, we demonstrate geomechanical modeling with fully automatic parameter calibration to estimate the full geomechanical stress fields of a prospective US carbon dioxide (CO2) storage site, based on sparse measurement data. The goal is to compute full stress tensor field estimates (principal stresses and orientations) that are maximally compatible with observations within the constraints of the model assumptions, thereby extending pointwise, incomplete partial stress measurement to a simulated full formation stress field, as well as a rough assessment of the associated error. We use the Perch site, located in Otsego County, Michigan, USA, as our case study. The input data consist of partial stress tensor information inferred from in-situ borehole tests, geophysical well logs, and processing of seismic data. A static earth model (SEM) of the site was developed, and geomechanical simulation functionality of the open-source MATLAB Reservoir Simulation Toolbox (MRST) was used to model the stress field. Adjoint-based nonlinear optimization was used to adjust boundary conditions and material properties to calibrate simulated results of observations. Results were interpreted through a Bayesian framework. The focus of this paper is to demonstrate how the fully automatic calibration procedure works and discuss the results obtained; it does not attempt a detailed analysis of the stress field in the context of the proposed CO2 storage initiatives. Our work is part of a larger effort to noninvasively determine in-situ stresses in deep formations considered for CO2 storage. Guided by previously published research on geomechanical model calibration, our work presents a novel calibration approach supporting a potentially large number of linear or nonlinear calibration parameters to produce results optimally agreeing with available measurements and thus extend partial pointwise estimates to full tensor fields compatible with the physics of the site.

Engineering↗

Building Alternative Indices of Socioeconomic Status for Population Modeling in Data-Sparse Contexts (Short Paper)

Population modeling requires clear definitions of socioeconomic status (SES) to ensure overall estimate accuracy and locate potentially underserved subpopulations. This presents a challenge as SES can be measured in myriad ways and for divergent purposes, and the data required to calculate these metrics may be lacking, particularly in low and middle income countries (LMICs). To support more refined SES measurement, we explore improvements upon the Demographic and Health Survey’s (DHS) Wealth Index (DHS-WI) using alternative characterizations of SES based on multiple correspondence analysis (MCA) and hierarchical clustering. We produce the MCA-derived metrics first on a full suite of household economic, demographic, and dwelling variables, then on a reduced set of occupant-only SES characteristics. We explore the utility of these metrics relative to DHS-WI based on their ability to 1) differentiate DHS household types and 2) identify mixtures of SES levels within DHS samples and mapped at high spatial resolution. We find that our full suite MCA yields more clearly defined SES segments and that our reduced MCA delineates occupant SES most clearly, suggesting potential pathways to improve upon the DHS-WI in future population modeling efforts for LMICs.

Cunningham, Angela↗

Sparse Data Machine Learning Integration with Theory, Experiment and Uncertainty Quantification: Process-Structure-Property-Performance of Friction Deformation Processing

Computer vision and deep learning tools that advance the ability to establish processing-structure-property-performance (PSPP) relations are presented. The Bayesian binning method for image segmentation enables quantitative analysis of microstructural features in an automated way, while the analysis of shapes and relative orientation of these features reveals local deformation maps indicative of both, material flow and residual stresses due to materials processing. The deep learning method leads to the previous knowledge agnostic mapping of empirically observed microstructural zones in friction stir welding (FSW) process and synthetic microstructure generation capability that is statistically equivalent to experimentally collected data.

97 MATHEMATICS AND COMPUTING↗

Variable rate neural compression for sparse detector data

Particle colliders produce data at extraordinary rates, posing major challenges for transmission and storage. High-throughput compression algorithms are therefore essential. In the sPHENIX experiment taking data at the Relativistic Heavy Ion Collider, a time projection chamber records three-dimensional (3D) particle trajectories that are highly sparse, making conventional learning-free lossy compression ineffective. Convolutional neural networks have surpassed traditional methods in compression ratio and accuracy. However, they fail to exploit sparsity for efficiency. To address these gaps, we present BCAE-VS, a bicephalous convolutional autoencoder with variable compression ratio for sparse data, which adapts compression to input complexity through key-point identification and sparse convolution. BCAE-VS achieves higher accuracy and compression ratios than prior neural approaches while being orders of magnitude smaller. Moreover, its throughput increases with sparsity—a property not observed in other methods. Although it was developed for collider experiments, BCAE-VS readily extends to other sparse data domains, such as light detection and ranging (LiDAR) sensing and 3D microscopy.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗