Search NASASearch

SEARCH · Search NASA

Results for “training data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Feature Acquisition with Imbalanced Training Data

This work considers cost-sensitive feature acquisition that attempts to classify a candidate datapoint from incomplete information. In this task, an agent acquires features of the datapoint using one or more costly diagnostic tests, and eventually ascribes a classification label. A cost function describes both the penalties for feature acquisition, as well as misclassification errors. A common solution is a Cost Sensitive Decision Tree (CSDT), a branching sequence of tests with features acquired at interior decision points and class assignment at the leaves. CSDT's can incorporate a wide range of diagnostic tests and can reflect arbitrary cost structures. They are particularly useful for online applications due to their low computational overhead. In this innovation, CSDT's are applied to cost-sensitive feature acquisition where the goal is to recognize very rare or unique phenomena in real time. Example applications from this domain include four areas. In stream processing, one seeks unique events in a real time data stream that is too large to store. In fault protection, a system must adapt quickly to react to anticipated errors by triggering repair activities or follow- up diagnostics. With real-time sensor networks, one seeks to classify unique, new events as they occur. With observational sciences, a new generation of instrumentation seeks unique events through online analysis of large observational datasets. This work presents a solution based on transfer learning principles that permits principled CSDT learning while exploiting any prior knowledge of the designer to correct both between-class and withinclass imbalance. Training examples are adaptively reweighted based on a decomposition of the data attributes. The result is a new, nonparametric representation that matches the anticipated attribute distribution for the target events.

Thompson, David R.

An Investigation of Interaction of Saharan Dust and Atlantic ITCZ Using Cloudsat-Calipso and A-Train Data

In this study, we investigate the radiative forcing of Saharan dust, its interactions with the Atlantic Intertropical Convergence Zone (ITCZ), through African easterly waves (AEW), African easterly jets (AEJ), and its impacts in short term numerical forecasts of tropical cyclogenesis using the GOCART-GEOS5 forecast system. Our approach is to develop and use an A-Train satellite simulator (ATSS) to constrain the observed aerosol index of refraction and particle size distribution by finding the values that simultaneously minimize the difference between observed CALIOP, CloudSat, OMI, and MODIS radiances and simulated radiances inverted from atmospheric model output using procedures and physical principles consistent with those used in corresponding retrieval algorithms. We use observations from the A-train and TRMM to determine relationships among the Saharan dust layer, transport by the AEW, and possible responses to dust radiative forcing in developing tropical cyclones in the A-ITCZ. Preliminary model results showing physical processes associated with the generation and transport of the Saharan dust layer, their interactions with the incipient moisture, clouds and rainfall in developing tropical cyclones will be presented. Also presented will be results of a case study of possible radiative impacts on AEW and AEJ during the NAMMA field campaign.

Lau, W.

Optimal A-Train Data Utilization: A Use Case of Aura OMI L2G and MERRA-2 Aerosol Products

Ozone Monitoring Instrument (OMI) aboard NASA's Aura mission measures ozone column and profile, aerosols, clouds, surface UV irradiance, and the trace gases including NO2, SO2, HCHO, BrO, and OClO using UltraViolet electromagnetic spectrum (280 - 400 nm) with a daily global coverage and a pixel spatial resolution of 13 km × 24 km at nadir, and it's been one of the key instruments to study the Earth's atmospheric composition and chemistry. The second Modern-Era Retrospective analysis for Research and Applications (MERRA-2) is NASA's atmospheric reanalysis using an upgraded version of Goddard Earth Observing System Model, version 5 (GEOS-5) data assimilation system. Compared to its predecessor MERRA, MERRA-2 is enhanced with more aspects of the Earth system among which is aerosol assimilation. When comparing between satellite pixel measurements and modeled grid data, how to properly handle counterpart pairing is critical considering their spatial and temporal variations. The comparison between satellite and model data by simply using Level 3 (L3) products may result biases due to lack of detailed temporal information. It has been preferred to inter-compare or implement satellite derived physical quantity (i.e., Level 2 (L2) Swath type) directly with/to model measurements with higher temporal and spatial resolution as possible. However, this has posed a challenge in the community to handle. Rather than directly handling the L2 or L3 data, there is a Level 2G (L2G) product conserving L2 pixel scientific data quality but in Grid type with the global coverage. In this presentation, we would like to demonstrate the optimal utilization of OMI L2G daily aerosol products by comparing with MERRA-2 hourly aerosol simulations matched well in both space and time.

MERRA-2 reanalysis

Similarity Metric for Data Optimization and Efficient Training of Reactive Machine Learning Force Fields for Hydrocarbon Radiolysis

Radiolysis is a common approach to sterilize polymers, chemically modify them for upcycling, and accelerate their decomposition for recycling purposes. Reactive molecular dynamics (MD) simulations provide a powerful tool to generate atomic-level trajectories of the reactive processes and quantify radiolytic chemical degradation pathways. For this, machine learning (ML) surrogate models for reactive force fields with quantum mechanical accuracy are now widely used, which require ML training data sets that can provide information on atomic environments for target chemical systems. However, radiolysis chemistry can be highly complex and diverse, which poses significant challenges for generating training data to parametrize ML models. In this regard, we developed a method for optimizing the training data set using a cosine similarity metric to help guide training set selection for radiolysis of polyethylene, a model hydrocarbon polymer, as well as to enhance the transferability of our reactive ML force field (MLFF) to a variety of molecular and polymeric systems. Our approach performs atom-by-atom comparisons between local atomic environments to pinpoint important data points associated with rare and localized events, such as radiolysis damage within structures. We apply this approach to train the Chebyshev Interaction Model for Efficient Simulation (ChIMES) MLFF model, which expresses the atomic interaction potentials in terms of linear combinations of many-body Chebyshev polynomials. We first show that our method can reduce our training set size by ∼70% while improving overall accuracy compared to more standard MD model fitting approaches. We then validate our optimum model against diverse hydrocarbon simulation data, including simple alkanes and systems with unsaturated carbon bonds, over a wide range of thermodynamic conditions. Finally, we use our ChIMES model to perform MD simulations of radiolytic damage with large-scale systems that help avoid system size effects. Overall, our approach yields an MD force field that retains most of the accuracy of the underlying quantum method while yielding many orders of improvement in computational efficiency. In conclusion, our efforts will have impact on future hydrocarbon polymer radiolysis studies, where the chemical details of the polymer–radiation interactions can have a strong effect on the resulting products observed in experiments.

Hydrocarbons

Neutral Networks for Subpixel Classification of Multispectral Images

In this work, the implementation and use of AVHRR (Advanced Very High Resolution Radiometer) images for subpixel analysis is studied. The work consists of two parts, the first is making training data, the second is using the training data to design a neural net subpixel analyzer. Most work on subpixel analysis has been done with images with more spectral bands. AVHRR images were chosen because of their easy acquisition, and because the five spectral bands allow investigation into the development of training data. The first step in subpixel analysis is the development of training data. This consists of image to be classified, and the classification of each pixel. In order to do the classification, a high spatial resolution image is typically needed in order to manually create a classified image. It is difficult to have both the image of interest, and a high spatial resolution image of the same area taken at the same time. Thus it was studied whether a subsampled image taken from the image of interest could serve as the training data. Statistical work has been done showing the unusefulness of this approach. In the second part of the analysis, a feedforward neural net was trained and used to classify the AVHRR images. Results of these tests, comparing the neural net with typical statistical based schemes are shown.

Figueroa, Ricardo R.

When more data hurts: Optimizing data coverage while mitigating diversity-induced underfitting in an ultrafast machine-learned potential

Machine-learned interatomic potentials (MLIPs) are becoming an essential tool in materials modeling. However, optimizing the generation of training data used to parametrize the MLIPs remains a significant challenge. This is because MLIPs can fail when encountering local environments too different from those present in the training data. The difficulty of determining a priori the environments that will be encountered during molecular dynamics simulation necessitates diverse, high-quality training data. Here, this study investigates how training data diversity affects the performance of MLIPs using the Ultra-Fast force field (UF 3 ) to model amorphous silicon nitride. We employ expert and autonomously generated data to create the training data and fit four force field variants to subsets of the data. Our findings reveal a critical balance in training data diversity: insufficient diversity hinders generalization, while excessive diversity can exceed the MLIP's learning capacity, reducing simulation accuracy. Specifically, we found that the UF 3 variant trained on a subset of the training data, in which nitrogen-rich structures were removed, offered vastly better prediction and simulation accuracy than any other variant. By comparing these UF 3 variants, we highlight the nuanced requirements for creating accurate MLIPs, emphasizing the importance of application-specific training data to achieve optimal performance in modeling complex material behaviors.

ab initio molecular dynamics

Converting sWeights to probabilities with density ratios

The use of machine learning approaches continues to have many benefits in experimental nuclear and particle physics. One common issue is generating training data which is sufficiently realistic to give reliable results. Here we advocate using real experimental data as the source of training data and demonstrate how one might subtract background contributions through the use of probabilistic weights which can be readily applied to training data. The sPlot formalism is a common tool used to isolate distributions from different sources. However, the negative sWeights produced by the sPlot technique can cause training problems and poor predictive power. This article demonstrates how density ratio estimation can be applied to convert sWeights to event probabilities, which we call drWeights. The drWeights can then be applied to produce the distributions of interest and are consistent with direct use of the sWeights. This article will also show how decision trees are particularly well suited to convert sWeights, with the benefit of fast prediction rates and adaptability to aspects of experimental data such as the data sample size and proportions of different event sources. We also show that a density ratio product approach in which the initial drWeights are reweighted by an additional converter gives substantially better results.

Glazier, Derek I. [Univ. of Glasgow, Scotland (Uni

Interpretable machine learning models classify minerals via spectroscopy

Developing methods to identify mineral species confidently and rapidly from Raman spectral analysis is critical to numerous fields. Traditionally, analysis relies on pattern matching the Raman spectrum of an unknown dataset with a supporting library of well-characterized spectral data, which may prove difficult for environmental samples that are poorly crystalline or phase mixtures. Here, we developed interpretable machine learning models that can classify uranium minerals by secondary oxyanion chemistry and other physicochemical properties based solely on Raman spectra. This new ML method produces a mineral profile of physical and chemical properties for an unknown sample and can rapidly classify or identify unknown minerals from Raman data, without the need for an exact pattern match in a spectral library. Training models are validated by 1. Strong correlation of high confidence model regions with published spectroscopic assignments and 2. Correct classification of a mineral not present in training data. Training data are from the Compendium of Uranium Raman and Infrared Experimental Spectra and available crystallographic information files within the open-source Smart Spectral Matching scientific framework. Physically meaningful classifier models can rapidly identify key structural and chemical information about unknown uranium minerals and the overall methodology is broadly applicable for mineral phases.

Machine learning

Reliability Modeling of Microelectromechanical Systems Using Neural Networks

Microelectromechanical systems (MEMS) are a broad and rapidly expanding field that is currently receiving a great deal of attention because of the potential to significantly improve the ability to sense, analyze, and control a variety of processes, such as heating and ventilation systems, automobiles, medicine, aeronautical flight, military surveillance, weather forecasting, and space exploration. MEMS are very small and are a blend of electrical and mechanical components, with electrical and mechanical systems on one chip. This research establishes reliability estimation and prediction for MEMS devices at the conceptual design phase using neural networks. At the conceptual design phase, before devices are built and tested, traditional methods of quantifying reliability are inadequate because the device is not in existence and cannot be tested to establish the reliability distributions. A novel approach using neural networks is created to predict the overall reliability of a MEMS device based on its components and each component's attributes. The methodology begins with collecting attribute data (fabrication process, physical specifications, operating environment, property characteristics, packaging, etc.) and reliability data for many types of microengines. The data are partitioned into training data (the majority) and validation data (the remainder). A neural network is applied to the training data (both attribute and reliability); the attributes become the system inputs and reliability data (cycles to failure), the system output. After the neural network is trained with sufficient data. the validation data are used to verify the neural networks provided accurate reliability estimates. Now, the reliability of a new proposed MEMS device can be estimated by using the appropriate trained neural networks developed in this work.

Perera. J. Sebastian

A Universal Augmentation Framework for Long-Range Electrostatics in Machine Learning Interatomic Potentials

Most current machine learning interatomic potentials (MLIPs) rely on short-range approximations, without explicit treatment of long-range electrostatics. To address this, we recently developed the Latent Ewald Summation (LES) method, which infers electrostatic interactions, polarization, and Born effective charges (BECs), just by learning from energy and force training data. Here, in this study, we present LES as a standalone library, compatible with any short-range MLIP, and demonstrate its integration with methods such as MACE, NequIP, Allegro, CACE, CHGNet, and UMA. We benchmark LES-enhanced models on distinct systems, including bulk water, polar dipeptides, and gold dimer adsorption on defective substrates, and show that LES not only captures correct electrostatics but also improves accuracy. Additionally, we scale LES to large and chemically diverse data by training MACELES-OFF on the SPICE set containing molecules and clusters, making a universal MLIP with electrostatics for organic systems, including biomolecules. MACELES-OFF is more accurate than its short-range counterpart (MACE-OFF) trained on the same data set, predicts dipoles and BECs reliably, and has better descriptions of bulk liquids. By enabling efficient long-range electrostatics without directly training on electrical properties, LES paves the way for electrostatic foundation MLIPs.

Kim, Dongjin [University of California, Berkeley,

Exploring and Visualizing A-Train Instrument Data

The succession of US and international satellites that follow each other in close succession, known as the A-Train, affords an opportunity to atmospheric researchers that no single platform could provide: Increasing the number of observations at any given geographic location.. . a more complete "virtual science platform". However, vertically and horizontally, co-registering and regridding datasets from independently developed missions, Aqua, Calipso, Cloudsat, Parasol, and Aura, so that they can be inter-compared can be daunting to some, and may be repeated by many. Scientists will individually spend much of their time and resources acquiring A-Train datasets of interest residing at various locations, developing algorithms to match up and graph datasets along the A-Train track, and search through large amounts of data for areas and/or phenomena of interest. The aggregate amount of effort that can be expended on repeating pre-science tasks could climb into the tens of millions of dollars. The goal of the A-Train Data Depot (ATDD) is to enable free movement of remotely located A-Train data so that they are combined to create a consolidated vertical view of the Earth's Atmosphere along the A-Train tracks. The innovative approach of analyzing and visualizing atmospheric profiles along the platforms track (i.e., time) is accomplished by through the ATDDs Giovanni data analysis and visualization tool. Giovanni brings together data from Aqua (MODIS, AIRS, AMSR-E), Cloudsat (cloud profiling radar) and Calipso (CALIOP, IIR), as well as the Aura (OMI, MLS, HIRDLS, TES) to create a consolidated vertical view of the Earth's Atmosphere along the A-Train tracks. This easy to learn and use exploration tool will allow users to create vertical profiles of any desired A-Train dataset, for any given time of choice. This presentation shows the power of Giovanni by describing and illustrating how this tool facilitates and aids A-Train science and research. A web based display system Giovanni provides users with the capability of creating co-located profile images of temperature and humidity data from the MODIS, MLS and AIRS instruments for a user specified time and spatial area. In addition, Cloud and Aerosol profiles may also be displayed for the Cloudsat and Caliop instruments. The ability to modify horizontal and vertical axis range, data range and dynamic color range is also provided. Two dimensional strip plots of MODIS, AIRS, OM1 and POLDER parameters, co-located along the Cloudsat reference track, can also be plotted along with the Cloudsat cloud profiling data. Center swath pixels for the same parameters can also be shown as line plots overlaying the Cloudsat or Calipso profile images. Images and subsetted data produced in each analysis run may be downloaded. Users truly can explore and discover data specific to their needs prior to ever transferring data to their analysis tools.

Kempler, S.

Surveillance system and method having parameter estimation and operating mode partitioning

A system and method for monitoring an apparatus or process asset including partitioning an unpartitioned training data set into a plurality of training data subsets each having an operating mode associated thereto; creating a process model comprised of a plurality of process submodels each trained as a function of at least one of the training data subsets; acquiring a current set of observed signal data values from the asset; determining an operating mode of the asset for the current set of observed signal data values; selecting a process submodel from the process model as a function of the determined operating mode of the asset; calculating a current set of estimated signal data values from the selected process submodel for the determined operating mode; and outputting the calculated current set of estimated signal data values for providing asset surveillance and/or control.

Bickford, Randall L.

Particle Filter Based Inference Testing

The primary intent of PAR-FIT (Particle Filter based Inference Testing) is to provide hard inductive evidence that a machine learning model is capable and proven for an individual test input. By examining training data used to form the underlying model functional correlation, an estimate of the reliability that a model will make the correct prediction can be made. The Sequential Probability Ratio Test is used to derive a qualitative evaluation for reliability based on hypothesis testing. The PAR-FIT framework achieves this by implementing a particle filter and the sequential probability ratio test algorithms on the machine learning model training data to determine relevancy of new individual test samples to the training dataset. The kernel function evaluates the local proximity and density of training data used to derive a prediction outcome. Particles are used to probabilistically determine which training data to evaluate for proximity. For test samples that are within a close proximity to and surrounded by multiple training data points, the evaluated reliability of the prediction is high. For test samples that are anomalies not represented by the training dataset, in low density data clusters, or are far from existing data points, the evaluated reliability is low as insufficient training evidence exists to suggest the model is capable of making the correct prediction. Sequential Probability Ratio Test is further used to determine when a hypothesis on whether a signal can be rejected or accepted for use. The ratio test collects sequence information from the particle filter to test whether the signal is anomalous or normal via hypothesis testing of the underlying distributions.

Chen, Edward [Idaho National Laboratory (INL), Ida

FiberFlex: Real-time FPGA-based Intelligent and Distributed Fiber Sensor System for Pedestrian Recognition

In recent years, security monitoring of public places and critical infrastructure has heavily relied on the widespread use of cameras, raising concerns about personal privacy violations. To balance the need for effective security monitoring with the protection of personal privacy, we explore the potential of optical fiber sensors for this application. This article proposes FiberFlex, an intelligent and distributed fiber sensor system. Ultizing Field Programmable Gate Arrays (FPGA) high-level synthesis (HLS) acceleration, FiberFlex offers real-time pedestrian detection by co-designing the entire pipeline of optical signal acquisition, processing, and recognition networks based on the principles of optical fiber sensing. As a promising alternative to traditional camera-based monitoring systems, FiberFlex achieves pedestrian detection by analyzing the vibration patterns caused by pedestrian footsteps, enabling security monitoring while preserving individual privacy. FiberFlex comprises three modules: First , fiber-optic sensing system: A fiber-optic distributed acoustic sensing (DAS) system is built and used to measure the ground vibration waves generated by people walking. Second , algorithms: We first collect the training data by measuring the ground vibration waves, label the data, and use the data to train the neural network models to perform pedestrian recognition. Third , hardware accelerators: We use HLS tools to design hardware modules on FPGA for data collection and pre-processing and integrate them with the downstream neural network accelerators to perform in-line real-time pedestrian detection. The final detection results are sent back from FPGA to the host CPU. We implement our system FiberFlex with the in-house built DAS system and AMD/Xilinx Kintex7 FPGA KC705 board and verify the whole system using the real-world collected data. We conduct recognition tests on five test subjects of varying ages, heights, and weights in a fixed sensing area. Each subject experienced 20 real-time recognition tests using their daily walking habits, and the subjects were given adequate rest between tests. After 100 tests on five test subjects, the overall real-time recognition accuracy exceeded \(88.0\%\) . The whole system uses 55 W of power, 33 W in the optical DAS system and 22 W in the FPGA. Relying on its end-to-end interdisciplinary design, FiberFlex seamlessly combines fiber-optic sensors with FPGA accelerators to enable low-power real-time security monitoring without compromising privacy, making it a valuable addition to the existing security monitoring network. According to FiberFlex, more valuable research can be conducted in the future, such as fall monitoring for the elderly, migration of identification networks between different application scenarios, and improvement of anti-interference performance in more complex environments. In future perception networks, where the “eyes” are not feasible, let’s use fiber optic touch instead.

Distributed

Accelerating multilevel Markov Chain Monte Carlo using machine learning models

Here, this work presents an efficient approach for accelerating multilevel Markov Chain Monte Carlo (MCMC) sampling for large-scale problems using low-fidelity machine learning models. While conventional techniques for large-scale Bayesian inference often substitute computationally expensive high-fidelity models with machine learning models, thereby introducing approximation errors, our approach offers a computationally efficient alternative by augmenting high-fidelity models with low-fidelity ones within a hierarchical framework. The multilevel approach utilizes the low-fidelity machine learning model (MLM) for inexpensive evaluation of proposed samples thereby improving the acceptance of samples by the high-fidelity model. The hierarchy in our multilevel algorithm is derived from geometric multigrid hierarchy. We utilize an MLM to accelerate the coarse level sampling. Training machine learning model for the coarsest level significantly reduces the computational cost associated with generating training data and training the model. We present an MCMC algorithm to accelerate the coarsest level sampling using MLM and account for the approximation error introduced. We provide theoretical proofs of detailed balance and demonstrate that our multilevel approach constitutes a consistent MCMC algorithm. Additionally, we derive the expression for cost reduction due to machine learning model to facilitate cost analysis of the hierarchical sampling algorithm. Our technique is demonstrated on a standard benchmark inference problem in groundwater flow, where we estimate the probability density of a quantity of interest using a four-level MCMC algorithm. Our proposed algorithm accelerates multilevel sampling by a factor of two while achieving similar accuracy compared to sampling using the standard multilevel algorithm.

97 MATHEMATICS AND COMPUTING