Search NASA⌕ Search

SEARCH · Search NASA

Results for “identify data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS↗

Compressional ULF waves in the outer magnetosphere. 2: A case study of Pc 5 type wave activity

In previously published work (Zhu and Kivelson, 1991) the spatial distribution of compressional magnetic pulsations of period 2 - 20 min in the outer magnetosphere was described. In this companion paper, we study some specific compressional events within our data set, seeking to determine the structure of the waves and identifying the wave generation mechanism. We use both the magnetic field and three-dimensional plasma data observed by the International Sun-Earth Explorer (ISEE) 1 and/or 2 spacecraft to characterize eight compressional ultra low frequency (ULF) wave events with frequencies below 8 mHz in the outer magnetosphere. High time resolution plasma data for the event of July 24, 1978, made possible a detailed analysis of the waves. Wave properties specific to the event of July 24, 1978, can be summarized as follows: (1) Partial plasma pressures in the different energy ranges responded to the magnetic field pressure differently. In the low-energy range they oscillated in phase with the magnetic pressure, while oscillations in higher-energy ranges were out-of-phase; (2) Perpendicular wavelengths for the event were determined to be 60,000 and 30,000 km in the radial and azimuthal directions, respectively. Wave properties common to all events can be summarized as follows: (1) Compressional Pc 5 wave activity is correlated with Beta, the ratio of the plasma pressure to the magnetic pressure; the absolute magnitude of the plasma pressure plays a minor role for the wave activity; (2) The magnetic equator is a node of the compressional perturbation of the magnetic field; (3) The criterion for the mirror mode instability is often satisfied near the equator in the outer magnetosphere when the compressional waves are present. We believe these waves are generated by internal magnetohydrodynamic (MHD) instabilities.

Zhu, Xiaoming↗

Towards a unified nonlocal, peridynamics framework for the coarse-graining of molecular dynamics data with fractures

Molecular dynamics (MD) has served as a powerful tool for designing materials with reduced reliance on laboratory testing. However, the use of MD directly to treat the deformation and failure of materials at the mesoscale is still largely beyond reach. In this work, we propose a learning framework to extract a peridynamics model as a mesoscale continuum surrogate from MD simulated material fracture data sets. Firstly, we develop a novel coarse-graining method, to automatically handle the material fracture and its corresponding discontinuities in the MD displacement data sets. Inspired by the weighted essentially non-oscillatory (WENO) scheme, the key idea lies at an adaptive procedure to automatically choose the locally smoothest stencil, then reconstruct the coarse-grained material displacement field as the piecewise smooth solutions containing discontinuities. Then, based on the coarse-grained MD data, a two-phase optimization-based learning approach is proposed to infer the optimal peridynamics model with damage criterion. In the first phase, we identify the optimal nonlocal kernel function from the data sets without material damage to capture the material stiffness properties. Then, in the second phase, the material damage criterion is learnt as a smoothed step function from the data with fractures. As a result, a peridynamics surrogate is obtained. As a continuum model, our peridynamics surrogate model can be employed in further prediction tasks with different grid resolutions from training, and hence allows for substantial reductions in computational cost compared with MD. We illustrate the efficacy of the proposed approach with several numerical tests for the dynamic crack propagation problem in a single-layer graphene. Our tests show that the proposed data-driven model is robust and generalizable, in the sense that it is capable of modeling the initialization and growth of fractures under discretization and loading settings that are different from the ones used during training.

97 MATHEMATICS AND COMPUTING↗

Land Surfaces at the Tipping-Point for Water and Energy Balance Coupling

The surface water and energy balances can be coupled or uncoupled depending on whether the evaporation regime is water-limited or energy-limited. As the landscape loses soil moisture during drydowns, a transition between the regimes may occur, which signifies a nonlinear change in water-energy-carbon coupling. Regions that switch often between these two regimes, that is, are dominated by neither regime, are particularly vulnerable to climate variability and change. To robustly identify these tipping points, we identify drydown events based on global soil moisture data sets from remote sensing. The event identification does not rely on precipitation information and is robust with respect to measurement noise. Then, the soil moisture thresholds delineating the evaporation regime transitions are determined by Sequential Monte Carlo Sampling and a two-stage parametrization strategy. Based on the estimated soil moisture thresholds across the globe, we estimate observation-based water availability indices which quantify the nonlinear controls of soil moisture on evaporation. This framework is tested and applied globally using Soil Moisture Active Passive soil moisture retrievals. Combined with a new tippling-point metric that describes the frequency of evaporation regime transitions, we identify regions that switch often between different evaporation regimes at the global scale. Given unit shifts in soil moisture, these regions will experience the most change in how their surface water and energy are coupled.

Jianzhi Dong↗

Venus small volcano classification and description

The high resolution and global coverage of the Magellan radar image data set allows detailed study of the smallest volcanoes on the planet. A modified classification scheme for volcanoes less than 20 km in diameter is shown and described. It is based on observations of all members of the 556 significant clusters or fields of small volcanoes located and described by this author during data collection for the Magellan Volcanic and Magmatic Feature Catalog. This global study of approximately 10 exp 4 volcanoes provides new information for refining small volcano classification based on individual characteristics. Total number of these volcanoes was estimated to be 10 exp 5 to 10 exp 6 planetwide based on pre-Magellan analysis of Venera 15/16, and during preparation of the global catalog, small volcanoes were identified individually or in clusters in every C1-MIDR mosaic of the Magellan data set. Basal diameter (based on 1000 measured edifices) generally ranges from 2 to 12 km with a mode of 34 km, and follows an exponential distribution similar to the size frequency distribution of seamounts as measured from GLORIA sonar images. This is a typical distribution for most size-limited natural phenomena unlike impact craters which follow a power law distribution and continue to infinitely increase in number with decreasing size. Using an exponential distribution calculated from measured small volcanoes selected globally at random, we can calculate total number possible given a minimum size. The paucity of edifice diameters less than 2 km may be due to inability to identify very small volcanic edifices in this data set; however, summit pits are recognizable at smaller diameters, and 2 km may represent a significant minimum diameter related to style of volcanic eruption. Guest, et al, discussed four general types of small volcanic edifices on Venus: (1) small lava shields; (2) small volcanic cones; (3) small volcanic domes; and (4) scalloped margin domes ('ticks'). Steep-sided domes or 'pancake domes', larger than 20 km in diameter, were included with the small volcanic domes. For the purposes of this study, only volcanic edifices less than 20 km in diameter are discussed. This forms a convenient cutoff since most of the steep-sided domes ('pancake domes') and scalloped margin domes ('ticks') are 20 to 100 km in diameter, are much less numerous globally than are the smaller diameter volcanic edifices (2 to 3 orders of magnitude lower in total global number), and do not commonly occur in large clusters or fields of large numbers of edifices.

Aubele, J. C.↗

Co-registration of Laser Altimeter Tracks with Digital Terrain Models and Applications in Planetary Science

We have derived algorithms and techniques to precisely co-register laser altimeter profiles with gridded Digital Terrain Models (DTMs), typically derived from stereo images. The algorithm consists of an initial grid search followed by a least-squares matching and yields the translation parameters at sub-pixel level needed to align the DTM and the laser profiles in 3D space. This software tool was primarily developed and tested for co-registration of laser profiles from the Lunar Orbiter Laser Altimeter (LOLA) with DTMs derived from the Lunar Reconnaissance Orbiter (LRO) Narrow Angle Camera (NAC) stereo images. Data sets can be co-registered with positional accuracy between 0.13 m and several meters depending on the pixel resolution and amount of laser shots, where rough surfaces typically result in more accurate co-registrations. Residual heights of the data sets are as small as 0.18 m. The software can be used to identify instrument misalignment, orbit errors, pointing jitter, or problems associated with reference frames being used. Also, assessments of DTM effective resolutions can be obtained. From the correct position between the two data sets, comparisons of surface morphology and roughness can be made at laser footprint- or DTM pixel-level. The precise co-registration allows us to carry out joint analysis of the data sets and ultimately to derive merged high-quality data products. Examples of matching other planetary data sets, like LOLA with LRO Wide Angle Camera (WAC) DTMs or Mars Orbiter Laser Altimeter (MOLA) with stereo models from the High Resolution Stereo Camera (HRSC) as well as Mercury Laser Altimeter (MLA) with Mercury Dual Imaging System (MDIS) are shown to demonstrate the broad science applications of the software tool.

Laser↗

HEPOM: Using Graph Neural Networks for the Accelerated Predictions of Hydrolysis Free Energies in Different pH Conditions

Hydrolysis is a fundamental family of chemical reactions where water facilitates the cleavage of bonds. The process is ubiquitous in biological and chemical systems, owing to water’s remarkable versatility as a solvent. However, accurately predicting the feasibility of hydrolysis through computational techniques is a difficult task, as subtle changes in reactant structure like heteroatom substitutions or neighboring functional groups can influence the reaction outcome. Furthermore, hydrolysis is sensitive to the pH of the aqueous medium, and the same reaction can have different reaction properties at different pH conditions. In this work, we have combined reaction templates and high-throughput ab initio calculations to construct a diverse data set of hydrolysis free energies. The developed framework automatically identifies reaction centers, generates hydrolysis products, and utilizes a trained graph neural network (GNN) model to predict ΔG values for all potential hydrolysis reactions in a given molecule. The long-term goal of the work is to develop a data-driven, computational tool for high-throughput screening of pH-specific hydrolytic stability and the rapid prediction of reaction products, which can then be applied in a wide array of applications including chemical recycling of polymers and ion-conducting membranes for clean energy generation and storage.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Vision for Coupling Operation of US Fusion Facilities with HPC Systems and the Implications for Workflows and Data Management

The operation of large US Department of Energy (DOE) research facilities, like the DIII-D National Fusion Facility, results in the collection of complex multi-dimensional scientific datasets, both experimental and model-generated. In the future, it is envisioned that integrated data analysis coupled with large-scale high performance computing (HPC) simulations will be used to improve experimental planning and operation. Practically, massive data sets from these simulations provide the physics basis for generation of both reduced semi-analytic and machine-learning-based models. Storage of both HPC simulation datasets (generated from US DOE leadership computing facilities) and experimental datasets presents significant challenges. In this paper, we present a vision for a DOE-wide data management workflow that integrates US DOE fusion facilities with leadership computing facilities. Data persistence and long-term availability beyond the length of allocated projects is essential, particularly for verification and recalibration of artificial intelligence and machine learning (AI/ML) models. Because these data sets are often generated and shared among hundreds of users across multiple leadership computing facility centers, they would benefit from cross-platform accessibility, persistent identifiers (e.g. DOI, or digital object identifier), and provenance tracking. Here, the ability to handle different data access patterns suggests that a combination of low cost, high latency (e.g. for storing ML training sets) and high cost, low latency systems (e.g. for real-time, integrated machine control feedback) may be needed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Seeing the Invisible: Embedding Tests in Code That Cannot be Modified

The difficulty of characterizing and observing valid software behavior during testing can be very difficult in flight systems. To address this issue, we evaluated several approaches to increasing test observability on the Shuttle Abort Flight Management (SAFM) system. To increase test observability, we added probes into the running system to evaluate the internal state and analyze test data. To minimize the impact of the instrumentation and reduce manual effort, we used Aspect-Oriented Programming (AOP) tools to instrument the source code. We developed and elicited a spectrum of properties, from generic to application specific properties, to be monitored via the instrumentation. To evaluate additional approaches, SAFM was ported to Linux, enabling the use of gcov for measuring test coverage, Valgrind for looking for memory usage errors, and libraries for finding non-normal floating point values. An in-house C++ source code scanning tool was also used to identify violations of SAFM coding standards, and other potentially problematic C++ constructs. Using these approaches with the existing test data sets, we were able to verify several important properties, confirm several problems and identify some previously unidentified issues.

O'Malley, Owen↗

Wavelet and Deep-Learning-Based Approach for Generation System Problematic Parameters Identification and Calibration

Accurate models of generation systems are critical for maintaining reliable and secure grid operations. In this paper, a novel and systematic approach is proposed to identify and calibrate the generation system problematic parameters using continuous wavelet transform (CWT) and advanced deep-learning technology. The phasor measurement unit (PMU) data are used through “event playback” to check whether the parameter calibration is required, and if yes, a group of suspicious parameters will be identified as the primary problematic parameter candidates (PPCs). These primary PPCs are randomly perturbed to generate the event playback simulation data, which are used by the CWT and convolutional neural networks (CNNs) to further narrow down the primary PPCs into a smaller set of candidates. Then, the identified candidates are perturbed again to generate massive event playback simulation data for training a parameter calibration neural network. Here, we designed a multi-output neural network structure to find the mappings between the perturbed parameters and the simulation data using both CNN and long short-term memory (LSTM) models. Finally, the well-trained and tested CNN-LSTM model is used to estimate the accurate value of the suspicious parameters with actual PMU measurements. The proposed CNN-LSTM network can accurately and reliably estimate the generation-system problematic parameters, and has better performance when compared to other machine-learning methods, such as the multilayer perceptron network and the conditional variational autoencoder method. The accuracy and effectiveness of the proposed approach have been validated through simulation and real-world data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Leveraging generative adversarial networks to create realistic scanning transmission electron microscopy images

Abstract The rise of automation and machine learning (ML) in electron microscopy has the potential to revolutionize materials research through autonomous data collection and processing. A significant challenge lies in developing ML models that rapidly generalize to large data sets under varying experimental conditions. We address this by employing a cycle generative adversarial network (CycleGAN) with a reciprocal space discriminator, which augments simulated data with realistic spatial frequency information. This allows the CycleGAN to generate images nearly indistinguishable from real data and provide labels for ML applications. We showcase our approach by training a fully convolutional network (FCN) to identify single atom defects in a 4.5 million atom data set, collected using automated acquisition in an aberration-corrected scanning transmission electron microscope (STEM). Our method produces adaptable FCNs that can adjust to dynamically changing experimental variables with minimal intervention, marking a crucial step towards fully autonomous harnessing of microscopy big data.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Machine-Learning Assisted Identification of Accurate Battery Lifetime Models with Uncertainty

Reduced-order battery lifetime models, which consist of algebraic expressions for various aging modes, are widely utilized for extrapolating degradation trends from accelerated aging tests to real-world aging scenarios. Identifying models with high accuracy and low uncertainty is crucial for ensuring that model extrapolations are believable, however, it is difficult to compose expressions that accurately predict multivariate data trends; a review of cycling degradation models from literature reveals a wide variety of functional relationships. Here, a machine-learning assisted model identification method is utilized to fit degradation in a stand-out LFP-Gr aging data set, with uncertainty quantified by bootstrap resampling. The model identified in this work results in approximately half the mean absolute error of a human expert model. Models are validated by converting to a state-equation form and comparing predictions against cells aging under varying loads. Parameter uncertainty is carried forward into an energy storage system simulation to estimate the impact of aging model uncertainty on system lifetime. The new model identification method used here reduces life-prediction uncertainty by more than a factor of three (86% ± 5% relative capacity at 10 years for human-expert model, 88.5% ± 1.5% for machine-learning assisted model), empowering more confident estimates of energy storage system lifetime.

25 ENERGY STORAGE↗

State Identification for Planetary Rovers: Learning and Recognition

A planetary rover must be able to identify states where it should stop or change its plan. With limited and infrequent communication from ground, the rover must recognize states accurately. However, the sensor data is inherently noisy, so identifying the temporal patterns of data that correspond to interesting or important states becomes a complex problem. In this paper, we present an approach to state identification using second-order Hidden Markov Models. Models are trained automatically on a set of labeled training data; the rover uses those models to identify its state from the observed data. The approach is demonstrated on data from a planetary rover platform.

Aycard, Olivier↗

Search and identification of transient and variable radio sources using MeerKAT observations: a case study on the MAXI J1820+070 field

ABSTRACT Many transient and variable sources detected at multiple wavelengths are also observed to vary at radio frequencies. However, these samples are typically biased towards sources that are initially detected in wide-field optical, X-ray, or gamma-ray surveys. Many sources that are insufficiently bright at higher frequencies are therefore missed, leading to potential gaps in our knowledge of these sources and missing populations that are not detectable in optical, X-rays, or gamma-rays. Taking advantage of new state-of-the-art radio facilities that provide high-quality wide-field images with fast survey speeds, we can now conduct unbiased surveys for transient and variable sources at radio frequencies. In this paper, we present an unbiased survey using observations obtained by MeerKAT, a mid-frequency (∼GHz) radio array in South Africa’s Karoo Desert. The observations used were obtained as part of a weekly monitoring campaign for X-ray binaries (XRBs) and we focus on the field of MAXI J1820+070. We develop methods to efficiently filter transient and variable candidates that can be directly applied to other data sets. In addition to MAXI J1820+070, we identify four likely active galactic nuclei, one source that could be a Galactic source (pulsar or quiescent XRB) or an AGN, and one variable pulsar. No transient sources, defined as being undetected in deep images, were identified leading to a transient surface density of <3.7 × 10−2 deg−2 at a sensitivity of 1 mJy on time-scales of 1 week at 1.4 GHz.

Rowlinson, A. (ORCID:0000000211957022)↗

Multimodality in the Search for New Physics in Pulsar Timing Data and the Case of Kination-amplified Gravitational-wave Background from Inflation

We investigate the kination-amplified inflationary gravitational-wave background (GWB) interpretation of the signal recently reported by various pulsar timing array (PTA) experiments. Kination is a post-inflationary phase in the expansion history dominated by the kinetic energy of some scalar field, characterized by a stiff equation of state w = 1. Within the inflationary GWB model, we identify two modes that can fit the current data sets (NANOGrav and EPTA) with equal likelihood: the kination-amplification (KA) mode and the ordinary, no-kination-amplification (no-KA) mode. The multimodality of the likelihood motivates a Bayesian analysis with nested sampling. We analyze the free spectra of current PTA data and mock free spectra constructed with higher signal-to-noise ratios using nested sampling. The analysis of the mock spectrum designed to be consistent with the best fit to the NANOGrav 15 yr (NG15) data successfully reveals the expected bimodal posterior for the first time while excluding the reheating mode that appears in the fit to the current NG15 data, making a case for our correct and comprehensive treatment of potential multimodal posteriors arising from future PTA data sets. The resultant Bayes factor is $\mathcal{B}$ $\equiv$ Z no–KA /Z KA = 2.9 ± 1.9, indicating comparable statistical significance between the two modes. Given the theoretical model-building challenges of producing highly blue-tilted primordial tensor spectra, the KA mode has the advantage of requiring less blue primordial spectra, compared with the no-KA mode. The synergy between future cosmic microwave background polarization, pulsar timing, and laser interferometer measurements of gravitational waves will help resolve the ambiguity implied by the multimodal posterior in PTA-only searches.

Cosmology↗

Parameter identifiability of linear dynamical systems

It is assumed that the system matrices of a stationary linear dynamical system were parametrized by a set of unknown parameters. The question considered here is, when can such a set of unknown parameters be identified from the observed data? Conditions for the local identifiability of a parametrization are derived in three situations: (1) when input/output observations are made, (2) when there exists an unknown feedback matrix in the system and (3) when the system is assumed to be driven by white noise and only output observations are made. Also a sufficient condition for global identifiability is derived.

Glover, K.↗

Detecting agricultural to urban land use change from multi-temporal MSS digital data

Conversion of agricultural land to a variety of urban uses is a major problem along the Wasatch Front, Utah. Although LANDSAT MSS data is a relatively coarse tool for discriminating categories of change in urban-size plots, its availability prompts a thorough test of its power to detect change. The procedures being applied to a test area in Salt Lake County, Utah, where the land conversion problem is acute are presented. The identity of land uses before and after conversion was determined and digital procedures for doing so were compared. Several algorithms were compared, utilizing both raw data and preprocessed data. Verification of results involved high quality color infrared photography and field observation. Two data sets were digitally registered, specific change categories internally identified in the software, results tabulated by computer, and change maps printed at 1:24,000 scale.

Ridd, M. K.↗

Activities of the Pilot Land Data System project

The University of Maryland's Remote Sensing Systems Laboratory submitted to NASA/Goddard an interim progress report on the work being conducted within its Pilot Land Data System IPLDS project. The Remote Sensing Systems Laboratory addressed the following tasks: (1) identify data types and data sources needed to describe the selected test sites in collaboration with Goddard's Hydrological Sciences Branch; (2) define the procedures necessary to access/acquire this data; (3) conduct meetings with the PLDS Systems Engineering Group to identify functional specification priorities for PLDS development; (4) assemble documentation on historical remotely sensed imagery and transfer of such information to the PLDS Data Management Group; (5) collect data identified by Goodard's Hydrological Sciences Branch for data set inventory in PLD; (6) develop a Workstation-PLDS system interface over high speed lines, (7) develop and test through a Phase 1 demonstration of a micro workstation to access PLDS; and (8) establish interdepartmental agreement of development of computer link for electronic access of water resources data from USGS.

Sircar, J. K.↗