Search NASASearch

SEARCH · Search NASA

Results for “sparse data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Pulse: An Outlier Sensitive Downsampling Algorithm For Timeseries Data

Pulse is a downsampling algorithm for timeseries data. Frequently datasets become so large that visualization tools and web browsers cannot effectively render graphics due to memory constraints. Downsampling algorithms are commonly applied to minimize the quantity of data required to visualize important features or trends in the data, but some datasets are composed by distinct enough features and trends that most existing downsampling algorithms fail to preserve them. Pule was developed to downsample timeseries data for galvanostatic stack test data at the Idaho National Laboratory. These datasets were composed by approximately 4 million records, most of them being extremely uniform. However, during relatively brief time periods when the stack test changes state, for example when the test article is powered on, or a load is added, the data produce sparse asymptotes. No existing downsampling algorithm was capable of preserving the sparse asymptotes in electrolysis stack test data. Instead, we develop a downsampling algorithm that preserves important outliers in data, and otherwise aggressively downsamples uniform data. The algorithm has applications in other domains like seismology, in the measurement of earthquakes, or astronomy, in the measurement of quasars or transit photometry.

Woodruff, Nathan [Idaho National Laboratory (INL),

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database

Identification of a precambrian rift through Missouri by digital image processing of geophysical and geological data

A newly discovered feature in the midcontinent - a gravity low that begins at a break in the midcontinent gravity high in SE Nebraska, extends across Missouri in a NW-SE direction, and intersects the Mississippi Valley graben to form the Pascola arch - is discussed. The anomaly varies from 120 to 160 km in width, extends approximately 700 km, and is best expressed in southern Missouri, where it has a Bouguer amplitude of about -34 mGal. It is noted that the magnitude of the anomaly cannot be explained on the basis of a thickened section of Paleozoic sedimentary rock. The gravity data and the sparse seismic refraction data for the region are found to be consistent with an increased crustal thickness beneath the gravity low. It is thought that the gravity anomaly is probably the present expression of a failed arm of a rifting event, perhaps one associated with the spreading that led to or preceded formation of the granite and rhyolite terrain of southern Missouri.

Guinness, E. A.

Inferring demographic and selective histories from population genomic data using a 2-step approach in species with coding-sparse genomes: an application to human data

Abstract The demographic history of a population, and the distribution of fitness effects (DFE) of newly arising mutations in functional genomic regions, are fundamental factors dictating both genetic variation and evolutionary trajectories. Although both demographic and DFE inference has been performed extensively in humans, these approaches have generally either been limited to simple demographic models involving a single population, or, where a complex population history has been inferred, without accounting for the potentially confounding effects of selection at linked sites. Taking advantage of the coding-sparse nature of the genome, we propose a 2-step approach in which coalescent simulations are first used to infer a complex multi-population demographic model, utilizing large non-functional regions that are likely free from the effects of background selection. We then use forward-in-time simulations to perform DFE inference in functional regions, conditional on the complex demography inferred and utilizing expected background selection effects in the estimation procedure. Throughout, recombination and mutation rate maps were used to account for the underlying empirical rate heterogeneity across the human genome. Importantly, within this framework it is possible to utilize and fit multiple aspects of the data, and this inference scheme represents a generalized approach for such large-scale inference in species with coding-sparse genomes.

Soni, Vivak (ORCID:0000000294969562)

Assessment of and standardization for quantitative nondestructive test

Present capabilities and limitations of nondestructive testing (NDT) as applied to aerospace structures during design, development, production, and operational phases are assessed. It will help determine what useful structural quantitative and qualitative data may be provided from raw materials to vehicle refurbishment. This assessment considers metal alloys systems and bonded composites presently applied in active NASA programs or strong contenders for future use. Quantitative and qualitative data has been summarized from recent literature, and in-house information, and presented along with a description of those structures or standards where the information was obtained. Examples, in tabular form, of NDT technique capabilities and limitations have been provided. NDT techniques discussed and assessed were radiography, ultrasonics, penetrants, thermal, acoustic, and electromagnetic. Quantitative data is sparse; therefore, obtaining statistically reliable flaw detection data must be strongly emphasized. The new requirements for reusable space vehicles have resulted in highly efficient design concepts operating in severe environments. This increases the need for quantitative NDT evaluation of selected structural components, the end item structure, and during refurbishment operations.

Neuschaefer, R. W.

Poisson vs. Gaussian statistics for sparse X-ray data: Application to the soft X-ray spectrometer

Reliable results when fitting X-ray data require proper consideration of the statistics involved. We probe the impact of Gaussian versus Poisson statistics at low count levels using both the standard χ^(2) method and maximum likelihood based on Poisson (C) statistics. The difference is studied and quantified through simulated spectra with known properties. We then test the results through analysis of Mn Kα calibration data taken with the flight spare microcalorimeter for the Hitomi soft X-ray spectrometer. Through comparison with simulations, our results show that the χ^(2) method tends to give overly optimistic estimates of the detector energy resolution, in particular when there are few counts. Given an energy resolution of ∼5 eV and a line with about 100 photons, the line width becomes ∼10% lower in the χ^(2) method than in Poisson statistics. This is a consequence of the uncertainties being dominated by counting statistics, and therefore highlights the need to choose the appropriate fit statistic.

Shinya Yamada

NASA Tech Briefs, March 2014

Topics include: Data Fusion for Global Estimation of Forest Characteristics From Sparse Lidar Data; Debris and Ice Mapping Analysis Tool - Database; Data Acquisition and Processing Software - DAPS; Metal-Assisted Fabrication of Biodegradable Porous Silicon Nanostructures; Post-Growth, In Situ Adhesion of Carbon Nanotubes to a Substrate for Robust CNT Cathodes; Integrated PEMFC Flow Field Design for Gravity-Independent Passive Water Removal; Thermal Mechanical Preparation of Glass Spheres; Mechanistic-Based Multiaxial-Stochastic-Strength Model for Transversely-Isotropic Brittle Materials; Methods for Mitigating Space Radiation Effects, Fault Detection and Correction, and Processing Sensor Data; Compact Ka-Band Antenna Feed with Double Circularly Polarized Capability; Dual-Leadframe Transient Liquid Phase Bonded Power Semiconductor Module Assembly and Bonding Process; Quad First Stage Processor: A Four-Channel Digitizer and Digital Beam-Forming Processor; Protective Sleeve for a Pyrotechnic Reefing Line Cutter; Metabolic Heat Regenerated Temperature Swing Adsorption; CubeSat Deployable Log Periodic Dipole Array; Re-entry Vehicle Shape for Enhanced Performance; NanoRacks-Scale MEMS Gas Chromatograph System; Variable Camber Aerodynamic Control Surfaces and Active Wing Shaping Control; Spacecraft Line-of-Sight Stabilization Using LWIR Earth Signature; Technique for Finding Retro-Reflectors in Flash LIDAR Imagery; Novel Hemispherical Dynamic Camera for EVAs; 360 deg Visual Detection and Object Tracking on an Autonomous Surface Vehicle; Simulation of Charge Carrier Mobility in Conducting Polymers; Observational Data Formatter Using CMOR for CMIP5; Propellant Loading Physics Model for Fault Detection Isolation and Recovery; Probabilistic Guidance for Swarms of Autonomous Agents; Reducing Drift in Stereo Visual Odometry; Future Air-Traffic Management Concepts Evaluation Tool; Examination and A Priori Analysis of a Direct Numerical Simulation Database for High-Pressure Turbulent Flows; and Resource-Constrained Application of Support Vector Machines to Imagery.

Source record

TPSAS-NF1676L-35322-DND

This talk discusses emerging methods that seek to fuse and integrate physics-based modeling with machine learning. With the recent rise of machine learning and artificial intelligence, there has been a huge surge in data-driven approaches to solve computational science and engineering problems. However, neglecting a priori knowledge of established physical laws and relying solely on data-driven methods can yield unreliable, less interpretable, and/or non-physical results, especially when data is sparse or predictions are required outside of the training data domain. This two part talk presents two distinct approaches for accelerating predictions with machine learning that are grounded and constrained by relevant physics and their application to problems at NASA.

Julian Cuevas Paniagua

Data-Driven Closures and Assimilation for Stiff Multiscale Random Dynamics

Here, we introduce a data-driven and physics-informed framework for propagating uncertainty in stiff, multiscale random ordinary differential equations (RODEs) driven by correlated (colored) noise. Unlike systems subjected to Gaussian white noise, a deterministic equation for the joint probability density function (PDF) of RODE state variables does not exist in closed form. Moreover, such an equation would require as many phase-space variables as there are states in the RODE system. To alleviate this curse of dimensionality, we instead derive exact, albeit unclosed, reduced-order PDF (RoPDF) equations for low-dimensional observables/quantities of interest. The unclosed terms take the form of state-dependent conditional expectations, which are directly estimated from data at sparse observation times. However, for systems exhibiting stiff, multiscale dynamics, data sparsity introduces regression discrepancies that compound during RoPDF evolution. This is overcome by introducing a kinetic-like defect term to the RoPDF equation, which is learned by assimilating in sparse, low-fidelity RoPDF estimates. Two assimilation methods are considered, namely nudging and deep neural networks, which are successfully tested against Monte Carlo simulations.

97 MATHEMATICS AND COMPUTING

Developing Data-Driven Synthetic Infrastructure Models for Resilience Analysis

Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.

24 POWER TRANSMISSION AND DISTRIBUTION

Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale. This has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. In this work, we implement ML-based geothermal exploration and an enhanced geothermal systems (EGS) design tool to achieve the above goals. Our exploration tool is GeoThermalCloud (GTC) EGS design tool is GeoDT-ML. GTC (github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. It enables the identification of critical measurements needed to identify geothermal resource signatures. GeoDT-ML (github.com/SmartTensors/GeoThermalCloud.jl/tree/master/) adds coupling to GeoDT (https://github.com/GeoDesignTool/GeoDT.git) for stochastic EGS design optimization and performance prediction. GeoDT-ML leverages recent advances in deep learning and high-performance computing. Contributors to this effort include LANL, PNNL, Google, Stanford, and Julia Computing.

15 GEOTHERMAL ENERGY

Gravity field improvement using global positioning system data from TOPEX/Poseidon - A covariance analysis

The TOPEX/Poseidon satellite data can be used to improve the knowledge of the earth's gravitational field. The GPS data are especially useful for improving the gravity field over the world's oceans, where the current tracking data are sparse. Using realistic scenario for processing 10 days of GPS data, a covariance analysis is performed to obtain the expected improvement to the GEM-T2 gravity field. The large amount of GPS data and the large number of parameters (1979 parameters for the gravity field, plus carrier-phase biases, etc.) required special filtering techniques for efficient solution. The gravity-bin technique is used to compute the covariance matrix associated with the spherical harmonic gravity field. The covariance analysis shows that the GPS data from one 10-day arc of TOPEX/Poseidon with no a priori constraints can resolve medium degree and order (3-26) parameters with sigmas (standard deviations) that are an order of magnitude smaller than the corresponding sigmas of GEM-T2. When the information from GEM-T2 is combined with the TOPEX/Poseidon GPS measurements, an order-of-magnitude improvement is observed in low- and medium-degree terms with significant improvements spread over a wide range of degree and order.

Bertiger, Willy I.

Solar Occultation Satellite Data and Derived Meteorological Products: Sampling Issues and Comparisons with Aura MLS

Derived Meteorological Products (DMPs, including potential temperature (theta), potential vorticity, equivalent latitude (EqL), horizontal winds and tropopause locations) have been produced for the locations and times of measurements by several solar occultation (SO) instruments and the Aura Microwave Limb Sounder (MLS). DMPs are calculated from several meteorological analyses for the Atmospheric Chemistry Experiment-Fourier Transform Spectrometer, Stratospheric Aerosol and Gas Experiment II and III, Halogen Occultation Experiment, and Polar Ozone and Aerosol Measurement II and III SO instruments and MLS. Time-series comparisons of MLS version 1.5 and SO data using DMPs show good qualitative agreement in time evolution of O3, N2O, H20, CO, HNO3, HCl and temperature; quantitative agreement is good in most cases. EqL-coordinate comparisons of MLS version 2.2 and SO data show good quantitative agreement throughout the stratosphere for most of these species, with significant biases for a few species in localized regions. Comparisons in EqL coordinates of MLS and SO data, and of SO data with geographically coincident MLS data provide insight into where and how sampling effects are important in interpretation of the sparse SO data, thus assisting in fully utilizing the SO data in scientific studies and comparisons with other sparse datasets. The DMPs are valuable for scientific studies and to facilitate validation of non-coincident measurements.

Manney, Gloria

Reevaluation of Stratospheric Ozone Trends From SAGE II Data Using a Simultaneous Temporal and Spatial Analysis

This paper details a new method of regression for sparsely sampled data sets for use with time-series analysis, in particular the Stratospheric Aerosol and Gas Experiment (SAGE) II ozone data set. Non-uniform spatial, temporal, and diurnal sampling present in the data set result in biased values for the long-term trend if not accounted for. This new method is performed close to the native resolution of measurements and is a simultaneous temporal and spatial analysis that accounts for potential diurnal ozone variation. Results show biases, introduced by the way data is prepared for use with traditional methods, can be as high as 10%. Derived long-term changes show declines in ozone similar to other studies but very different trends in the presumed recovery period, with differences up to 2% per decade. The regression model allows for a variable turnaround time and reveals a hemispheric asymmetry in derived trends in the middle to upper stratosphere. Similar methodology is also applied to SAGE II aerosol optical depth data to create a new volcanic proxy that covers the SAGE II mission period. Ultimately this technique may be extensible towards the inclusion of multiple data sets without the need for homogenization.

Damadeo, R. P.

Global detailed gravimetric geoid

A global detailed gravimetric geoid has been computed by combining the Goddard Space Flight Center GEM-4 gravity model derived from satellite and surface gravity data and surface 1 deg-by-1 deg mean free air gravity anomaly data. The accuracy of the geoid is + or - 2 meters on continents, 5 to 7 meters in areas where surface gravity data are sparse, and 10 to 15 meters in areas where no surface gravity data are available. Comparisons have been made with the astrogeodetic data provided by Rice (United States), Bomford (Europe), and Mather (Australia). Comparisons have also been carried out with geoid heights derived from satellite solutions for geocentric station coordinates in North America, the Caribbean, Europe, and Australia.

Vincent, S.

Global detailed gravimetric geoid

A global detailed gravimetric geoid has been computed by combining the Goddard Space Flight Center GEM-4 gravity model derived from satellite and surface gravity data and surface 1 x 1-deg mean free-air gravity anomaly data. The accuracy of the geoid is plus or minus 2 meters on continents, 5 to 7 meters in areas where surface gravity data are sparse, and 10 to 15 meters in areas where no surface gravity data are available. Comparisons have been made with the astrogeodetic data provided by Rice (United States), Bomford (Europe), and Mather (Australia). Comparisons have also been carried out with geoid heights derived from satellite solutions for geocentric station coordinates in North America, the Caribbean, Europe and Australia.

Vincent, S.

Requirements for facilities and measurement techniques to support CFD development for hypersonic aircraft

The design of a hypersonic aircraft poses unique challenges to the engineering community. Problems with duplicating flight conditions in ground based facilities have made performance predictions risky. Computational fluid dynamics (CFD) has been proposed as an additional means of providing design data. At the present time, CFD codes are being validated based on sparse experimental data and then used to predict performance at flight conditions with generally unknown levels of uncertainty. This paper will discuss the facility and measurement techniques that are required to support CFD development for the design of hypersonic aircraft. Illustrations are given of recent success in combining experimental and direct numerical simulation in CFD model development and validation for hypersonic perfect gas flows.

Sellers, William L., III

A Global Map of Mars' Crustal Magnetic Field Based on Electron Reflectometry

One of the great surprises of the Mars Global Surveyor mission was the discovery of intensely magnetized crust. Magnetic sources on Mars are at least ten times stronger than their terrestrial counterparts, probably requiring large volumes of coherently magnetized material, very strong remanence, or both. Although much of the attention so far has been placed on the strong crustal fields in the southern highlands, magnetic sources do exist in the younger low-lying plains. The strength and morphology of these sources could yield clues to the thermal and magnetic history of the northern plains. Low altitude (approx. 100 km) Magnetometer (MAG) data obtained during aerobraking have the greatest spatial resolution and sensitivity for identifying crustal magnetic sources from orbit, but those data are sparse and therefore limit the ability to discern morphology. Fully sampled MAG data obtained in the 400-km altitude mapping orbit have been differenced with respect to latitude (Br/Lat) to minimize the influence of induced fields from the solar wind interaction and thus enhance the sensitivity to weak crustal sources. Here we describe independent results from the Electron Reflectometer (ER), which remotely measures the magnetic field intensity at approx. 170 km altitude, and is roughly seven times more sensitive to crustal magnetic sources than measurements of Br from the mapping orbit.

Mitchell, D. L.