Search NASASearch

SEARCH · Search NASA

Results for “Machine Learning Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Machine Learning the COSMO Model for Predicting Thermodynamics of Electrolyte Mixtures

Bottom-up design of electrolyte mixtures for battery systems requires predicting macro thermodynamic properties from molecular constituents. For instance, molten salt electrolyte batteries require conditions far above room temperature to operate. Therefore, discovering mixtures with increasingly lower eutectic melting points is desirable. A model that can approximate chemical activity is a valuable tool to search through the vast compositional design space. Machine learning can predict properties of materials such as vibrational free energies, electronic energy gaps, and thermal conductivities. Moreover, they can learn physical models such as interatomic potentials. The COSMO-SAC model uses theory and empirical parameterization to predict liquid-vapor and liquid-solid properties using first-principles calculations. However, obtaining activity coefficients required for parameterizing the COSMO-SAC model is costly and limited to a select chemical space. In this work, we explored if machine learning methods could improve the COSMO-SAC model and bridge density functional theory calculations to liquid phase thermodynamic properties. Our data-driven approach uses existing databases for sigma-profiles of organic solvents and reconciles their methodological differences via ensemble averaging. First, an optimal machine learning model is constructed for each dataset. Our machine learning algorithms use the sigma-profile as an input feature to predict binary mixtures' activity coefficients using multi-output regression. Each dataset uses different choices of functionals, methods, and basis sets. Therefore, our ensemble model attempts to predict corrected activity coefficients given the combination of all the model outputs. The activity coefficients used for training are generated using the COSMO-SAC model. This approach enables the extraction of meaningful information from the existing datasets to improve the COSMO-SAC model for obtaining thermodynamic properties of electrolyte mixtures. With the liquid phase activities, we can identify electrolyte mixtures that meet desired phase equilibria conditions.

Thermodynamics

NASA GPM Status and Future Activities

The joint U.S.-Japan Global Precipitation Measurement (GPM) mission is approaching a decade of operations, and continues to pursue research, dataset production, and outreach related to precipitation. One key activity over the last year was the release of an improved “Version 07” of all GPM precipitation and latent heating products. This talk summarizes key improvements to the GPM products for which NASA has lead responsibility and provides some examples of the changes between Versions 06 and 07 in algorithm performance. One important operational change that affected Version 07 is that the scanning strategy for the Ka-band radar channel changed in May 2018; all products that depend on Ka were revised to accommodate this change. For example, in Version 07 the Goddard Profiling (GPROF) algorithm has implemented improvements in regions where orographic enhancement and suppression take place and where the surface is snowy/icy, and again covers radiometers reaching back to 1987. The Combined Radar Radiometer Algorithm (CORRA) now incorporates modified drop-size distribution constraints that substantially reduce bias. Revisions to the Convective-Stratiform Heating (CSH) algorithm employ new radiative transfer retrievals as well as accounting for terrain in the vertical coordinates. Each algorithm was adjusted to ensure continuity for each product across the boundary in 2014 between the predecessor Tropical Rainfall Measuring Mission (TRMM) and the GPM Core Observatory. The U.S. Science Team’s Integrated Multi-satellitE Retrievals for GPM (IMERG) was upgraded to account for distortions in the probability density function of regional precipitation rates due to weighted averaging in the Kalman filter used for “morphing” the passive microwave data. The talk will conclude by considering major issues that require continued attention, including the use of machine learning algorithms, the operational challenge of swarms of “small”, perhaps short-lived satellites, and estimates of the remaining lifespan of the Core Observatory.

Global Precipitation Measuremen

Educational and Scientific Applications of Climate Model Diagnostic Analyzer

Climate Model Diagnostic Analyzer (CMDA) is a web-based information system designed for the climate modeling and model analysis community to analyze climate data from models and observations. CMDA provides tools to diagnostically analyze climate data for model validation and improvement, and to systematically manage analysis provenance for sharing results with other investigators. CMDA utilizes cloud computing resources, multi-threading computing, machine-learning algorithms, web service technologies, and provenance-supporting technologies to address technical challenges that the Earth science modeling and model analysis community faces in evaluating and diagnosing climate models. As CMDA technology and infrastructure have matured, we have developed the educational and scientific applications of CMDA. Educationally, CMDA supported the summer school of the JPL Center for Climate Sciences in 2014, 2015, and 2016. In the summer school, the students work on group research projects where CMDA provide datasets, analysis tools, and provenance support utility tools. Each student is assigned to a virtual machine with CMDA installed in Amazon Web Services. Scientifically, we have developed several science use cases of CMDA covering various topics, datasets, and analysis types. Each of the science use cases is described in terms of a scientific goal, datasets used, the analysis tools used, scientific results discovered, an analysis result such as output plots and data files, and a link to the corresponding analysis service call with all the input arguments filled.

Bao, Qihao

Extracting Lessons of Resilience Using Machine Mining of the ASRS Database

NASA’s Aviation Safety Reporting System (ASRS) database is the world's largest repository of voluntary, confidential safety information provided by aviation's frontline personnel, including pilots, air traffic controllers, mechanics, flight attendants, dispatchers, and other members of the aviation community and the public. The database contains close to 2 million narratives, many of which describe everyday situations in which people saved the day. In these situations, people’s resilient behavior solved a problem, dealt with a malfunction, and maintained a safe operation despite a serious perturbation. To be able to extract lessons of such resilience from this large database, the use of machine learning algorithms is being explored. In this report, we describe a comparison between two such algorithms: Perilog and Word2Vec. An identical search using both programs was done on a database containing approximately 470,000 ASRS reports submitted between 1988 and 2022. The comparison reveals some of the strength and weaknesses of each algorithm as well as the challenges inherent in using such algorithms to extract lessons of resilience from the ASRS database.

resilience

NASA GPM Status and Future Activities

The joint U.S.-Japan Global Precipitation Measurement (GPM) mission is approaching a decade of operations, and continues to pursue research, dataset production, and outreach related to precipitation. Key activities over the last year were the release of an improved “Version 07” of all GPM precipitation and latent heating products, boosting the orbit of the GPM Core Observatory (GPM CO) to 435 km, and improving quality control on precipitation retrievals from the GPM constellation of passive microwave satellites. This presentation summarizes key improvements to the GPM products and provides some examples of the changes between Versions 06 and 07 in algorithm performance. One important operational change that affected Version 07 is that the scanning strategy for the Ka-band radar channel changed in May 2018; all products that depend on Ka were revised to accommodate this change. For example, in Version 07 the Goddard Profiling (GPROF) algorithm has implemented improvements in regions where orographic enhancement and suppression take place and where the surface is snowy/icy, and again covers radiometers reaching back to 1987. The Combined Radar Radiometer Algorithm (CORRA) now incorporates modified drop-size distribution constraints that substantially reduce bias. Revisions to the Convective-Stratiform Heating (CSH) algorithm employ new radiative transfer retrievals as well as accounting for terrain in the vertical coordinates. Each algorithm was adjusted to ensure continuity for each product across the boundary in 2014 between the predecessor Tropical Rainfall Measuring Mission (TRMM) and the GPM CO. The U.S. Science Team’s Integrated Multi-satellitE Retrievals for GPM (IMERG) was upgraded to account for distortions in the probability density function of regional precipitation rates due to weighted averaging in the Kalman filter used for “morphing” the passive microwave data. Maintaining the GPM CO orbital altitude in the the current very active solar cycle has been forcing the use of more fuel than planned and consequently shortening the forecasted life of the mission from the early 2030's to the late 2020's. It was considered vital to regain some of this lifetime to ensure overlap with the upcoming Atmosphere Observing System mission to provide crosscalibration of instruments. To accomplish this, the orbital altitude was raised from 400 to 435 km on 7-8 November 2023. Thereafter, the primary GPM CO algorithms had to be revised to account for the change in observing parameters. By meeting time this action should be complete. Recently, a screening algorithm based on auto-encoding was developed that uncovered 162 orbits (out of the many thousands of orbits across all years and all satellites) of passive microwave retrievals that had highly anomalous values. Removing these defective retrievals has improved the integrity of both the GPROF and IMERG records. However, the nature of the IMERG processing interacted sufficiently badly with the now-discovered anomalous orbits that it was necessary to completely reprocess the IMERG Final Run record, now labeled Version 07B. The presentation also considers major issues that require continued attention, including the use of machine learning algorithms and the operational challenge of swarms of “small”, perhaps short-lived satellites.

GPM

Understanding the Scalability of Bayesian Network Inference Using Clique Tree Growth Curves

One of the main approaches to performing computation in Bayesian networks (BNs) is clique tree clustering and propagation. The clique tree approach consists of propagation in a clique tree compiled from a Bayesian network, and while it was introduced in the 1980s, there is still a lack of understanding of how clique tree computation time depends on variations in BN size and structure. In this article, we improve this understanding by developing an approach to characterizing clique tree growth as a function of parameters that can be computed in polynomial time from BNs, specifically: (i) the ratio of the number of a BN s non-root nodes to the number of root nodes, and (ii) the expected number of moral edges in their moral graphs. Analytically, we partition the set of cliques in a clique tree into different sets, and introduce a growth curve for the total size of each set. For the special case of bipartite BNs, there are two sets and two growth curves, a mixed clique growth curve and a root clique growth curve. In experiments, where random bipartite BNs generated using the BPART algorithm are studied, we systematically increase the out-degree of the root nodes in bipartite Bayesian networks, by increasing the number of leaf nodes. Surprisingly, root clique growth is well-approximated by Gompertz growth curves, an S-shaped family of curves that has previously been used to describe growth processes in biology, medicine, and neuroscience. We believe that this research improves the understanding of the scaling behavior of clique tree clustering for a certain class of Bayesian networks; presents an aid for trade-off studies of clique tree clustering using growth curves; and ultimately provides a foundation for benchmarking and developing improved BN inference and machine learning algorithms.

Mengshoel, Ole J.

A Machine-Learning Approach to Assess Aircraft Engine System Performance

Artificial intelligence (AI)/machine learning, and big data are transforming the global business environment. They have become the most disruptive technologies for organizations to improve workplace efficiency and productivity. This work explored the application of machine learning-based predictive analytics that would enable aircraft engine designers to estimate engine system performance quickly during the conceptual design stage. Supervised machine-learning algorithm was employed to study patterns in an existing database of production and research turbofan engines, and built predictive analytics for use in predicting system performance of new turbofan designs. Specifically, the author developed deep-learning analytics to predict turbofan system weight, using turbofan design parameters as the input. The predictive analytics were trained and deployed in Keras, an open-source neural networks API (application program interface) written in Python, with TensorFlow (an open-source artificial AI library developed by Google) serving as the backend engine. The current engine-weight prediction results, together with those for the TSFC (thrust specific fuel consumption) and core-size predictions that were studied previously by the author, show that machine learning-based predictive analytics can be an effective, time-saving tool for aircraft engine design-space exploration during the conceptual design stage. It would enable expeditious identification of the best engine design amongst several candidates.

Michael T Tong

Using Federated Learning to Overcome Data Gravity in Space

Humans intend to take longer missions to outer space. Understanding the impact that space has on human health is paramount to the success of these missions. Controlled experiments with model organisms are run to infer the impact of space conditions on human health, but the data these experiments generate are too large to transfer to Earth for building models. The same is true for space-relevant data generated on Earth. Ideally, these datasets should be combined to improve statistical power and model accuracy without having to transfer data. Federated learning is such a method which trains an algorithm across decentralized computing systems, each of which has their own local copy of training and testing data. In this research, made possible by NASA@Work, the AI for Life in Space group at NASA demonstrates the use of federated learning to train an ensemble of causality inference models on a combination of data residing on the International Space Station (ISS) and in the cloud. Our work leverages CRISP, a causal inference platform developed during the 2020 Frontier Development Lab’s “Astronaut Health Challenge.” We also leverage the OpenFL federated learning library which was collaboratively developed at Intel and UPenn. We used publicly available data from the NASA Ames Life Sciences Data Archive to identify features in ionizing radiation experiments as causal of changes in cardiac blood velocity. This research demonstrates, for the first time, the possibility of running machine learning algorithms on datasets separated by astronomical distances. In this experiment, all the data were generated in terra, half of which were transferred to the ISS and analyzed on the Spaceborne Computer. In the future, our research will leverage federated learning on data generated in situ on the ISS with data generated terrestrially to predict the impact of spaceflight on mammalian female reproductive capacity.

James Casaletto

Snow Depth from AMSR-2 Using Multispectral Satellite Data in an Artificial Neural Network

By using diffusion theory and Monte Carlo lidar radiative transfer simulations, Hu et al. (2022b) has derived snow depth from the first-, second- and third-order moments of the lidar backscattering pathlength distribution. Lu et al. (2022) calculated the snow depth by applying the methods to the satellite ICESat-2 lidar measurements over the Arctic sea ice, as well as land surfaces of Northern Hemisphere. In this paper, an artificial neural network (ANN) algorithm, employing several channels from Advanced Microwave Scanning Radiometer 2 (AMSR-2) and the humidity vertical profiles from Global Modeling and Assimilation Office (GMAO) Goddard Earth Observing System for Instrument Teams (GEOS-IT) product, is trained to determine snow depth identified by time and geolocation matched 2019 ICESat-2 snow-depth data during winter months over the Arctic sea ice. The trained ANN snow-depth was applied to 2018 AMSR-2 clear pixel data, although the algorithms perform reasonably well in thinner clouds. The validation data (different from the training set) of ANN snow depth from AMSR-2 showed a good agreement with time matched and co-located snow-depth values from ICESat-2. The bias was near zero, with mean absolute error (MAE) 0.05 cm and a root-mean-square-error (RMSE) 0.08 cm. Prior applying the trained ANN snow depth to AMSR-2 data, a cloud screening algorithm was developed with a similar approach. A separate ANN cloud mask was trained to determine an AMSR-2 pixel is clear or cloudy with time and geolocation matched 2015 CALIOP Vertical Feature Mask (VFM) over Arctic sea ice. The ANN cloud mask from AMSR-2 under-estimated cloud fraction by 3-6% compared to CALIOP . The additional research is needed to conclusively evaluate the ANN cloud mask accuracy. Finally, this paper will lay the foundation for a sustained long-term snowfall and snow-storm monitoring system. The future Cloud Aerosol LIdar for Global scale Observations of the ocean-Land Atmosphere system (CALIGOLA) mission will provide a means to calculate snow depth from the lidar backscattering pathlength distribution, benefiting from the UV, visible and infrared pulses. With the calculated snow depth as the truth one could develop a machine learning algorithm, as it was done in this paper, using a passive microwave instrument available at that time to generate a wide range of snow depth data, covering extensive spatial areas in the cross-orbit direction.

Neural Network

U.S. GPM Status

The joint U.S.-Japan Global Precipitation Measurement (GPM) mission, now in its tenth year of operations, continues to pursue research on scientific and operational shortcomings in algorithms for retrieving global precipitation from satellite observations. A summary of operations, including the orbit boost, processing, science, and applications is provided. This includes key improvements to the GPM products for which NASA has lead responsibility. This includes refinements to ensure continuity for each product across the boundary in 2014 between the predecessor Tropical Rainfall Measuring Mission (TRMM) and the GPM Core Observatory. Early evaluations of the U.S. Science Team’s Integrated Multi-satellitE Retrievals for GPM (IMERG) in V07 show interesting behavior across the TRMM orbit boost that is highly relevant to the planned orbit boost for the GPM Core observatory. We will present summary statistics that illustrate the tremendous demand that exists in the user community for GPM precipitation products, which on the U.S. side is primarily focused on IMERG, and end by considering major issues that might be addressed in up-coming versions. These include the use of machine learning algorithms, the operational challenge of swarms of “small”, perhaps short-lived satellites, and continuation of precessing calibration satellites after GPM.

GPM

Acquisition and production of skilled behavior in dynamic decision-making tasks

Detailed summaries of two NASA-funded research projects are provided. The first project was an ecological task analysis of the Star Cruiser model. Star Cruiser is a psychological model designed to test a subject's level of cognitive activity. Ecological task analysis is used as a framework to predict the types of cognitive activity required to achieve productive behavior and to suggest how interfaces can be manipulated to alleviate certain types of cognitive demands. The second project is presented in the form of a thesis for the Masters Degree. The thesis discusses the modeling of decision-making through the use of neural network and genetic-algorithm machine learning technologies.

Kirlik, Alex

Population-based learning of load balancing policies for a distributed computer system

Effective load-balancing policies use dynamic resource information to schedule tasks in a distributed computer system. We present a novel method for automatically learning such policies. At each site in our system, we use a comparator neural network to predict the relative speedup of an incoming task using only the resource-utilization patterns obtained prior to the task's arrival. Outputs of these comparator networks are broadcast periodically over the distributed system, and the resource schedulers at each site use these values to determine the best site for executing an incoming task. The delays incurred in propagating workload information and tasks from one site to another, as well as the dynamic and unpredictable nature of workloads in multiprogrammed multiprocessors, may cause the workload pattern at the time of execution to differ from patterns prevailing at the times of load-index computation and decision making. Our load-balancing policy accommodates this uncertainty by using certain tunable parameters. We present a population-based machine-learning algorithm that adjusts these parameters in order to achieve high average speedups with respect to local execution. Our results show that our load-balancing policy, when combined with the comparator neural network for workload characterization, is effective in exploiting idle resources in a distributed computer system.

Mehra, Pankaj

Application of Domain Knowledge to Software Quality Assurance

This work focused on capturing, using, and evolving a qualitative decision support structure across the life cycle of a project. The particular application of this study was towards business process reengineering and the representation of the business process in a set of Business Rules (BR). In this work, we defined a decision model which captured the qualitative decision deliberation process. It represented arguments both for and against proposed alternatives to a problem. It was felt that the subjective nature of many critical business policy decisions required a qualitative modeling approach similar to that of Lee and Mylopoulos. While previous work was limited almost exclusively to the decision capture phase, which occurs early in the project life cycle, we investigated the use of such a model during the later stages as well. One of our significant developments was the use of the decision model during the operational phase of a project. By operational phase, we mean the phase in which the system or set of policies which were earlier decided are deployed and put into practice. By making the decision model available to operational decision makers, they would have access to the arguments pro and con for a variety of actions and can thus make a more informed decision which balances the often conflicting criteria by which the value of action is measured. We also developed the concept of a 'monitored decision' in which metrics of performance were identified during the decision making process and used to evaluate the quality of that decision. It is important to monitor those decision which seem at highest risk of not meeting their stated objectives. Operational decisions are also potentially high risk decisions. Finally, we investigated the use of performance metrics for monitored decisions and audit logs of operational decisions in order to feed an evolutionary phase of the the life cycle. During evolution, decisions are revisisted, assumptions verified or refuted, and possible reassessments resulting in new policy are made. In this regard we implemented a machine learning algorithm which automatically defined business rules based on expert assessment of the quality of operational decisions as recorded during deployment.

Wild, Christian W.

Visual Inference Programming

The goal of visual inference programming is to develop a software framework data analysis and to provide machine learning algorithms for inter-active data exploration and visualization. The topics include: 1) Intelligent Data Understanding (IDU) framework; 2) Challenge problems; 3) What's new here; 4) Framework features; 5) Wiring diagram; 6) Generated script; 7) Results of script; 8) Initial algorithms; 9) Independent Component Analysis for instrument diagnosis; 10) Output sensory mapping virtual joystick; 11) Output sensory mapping typing; 12) Closed-loop feedback mu-rhythm control; 13) Closed-loop training; 14) Data sources; and 15) Algorithms. This paper is in viewgraph form.

Wheeler, Kevin

Software for Partly Automated Recognition of Targets

The Feature Analyst is a computer program for assisted (partially automated) recognition of targets in images. This program was developed to accelerate the processing of high-resolution satellite image data for incorporation into geographic information systems (GIS). This program creates an advanced user interface that embeds proprietary machine-learning algorithms in commercial image-processing and GIS software. A human analyst provides samples of target features from multiple sets of data, then the software develops a data-fusion model that automatically extracts the remaining features from selected sets of data. The program thus leverages the natural ability of humans to recognize objects in complex scenes, without requiring the user to explain the human visual recognition process by means of lengthy software. Two major subprograms are the reactive agent and the thinking agent. The reactive agent strives to quickly learn the user's tendencies while the user is selecting targets and to increase the user's productivity by immediately suggesting the next set of pixels that the user may wish to select. The thinking agent utilizes all available resources, taking as much time as needed, to produce the most accurate autonomous feature-extraction model possible.

Opitz, David

Understanding the Scalability of Bayesian Network Inference using Clique Tree Growth Curves

Bayesian networks (BNs) are used to represent and efficiently compute with multi-variate probability distributions in a wide range of disciplines. One of the main approaches to perform computation in BNs is clique tree clustering and propagation. In this approach, BN computation consists of propagation in a clique tree compiled from a Bayesian network. There is a lack of understanding of how clique tree computation time, and BN computation time in more general, depends on variations in BN size and structure. On the one hand, complexity results tell us that many interesting BN queries are NP-hard or worse to answer, and it is not hard to find application BNs where the clique tree approach in practice cannot be used. On the other hand, it is well-known that tree-structured BNs can be used to answer probabilistic queries in polynomial time. In this article, we develop an approach to characterizing clique tree growth as a function of parameters that can be computed in polynomial time from BNs, specifically: (i) the ratio of the number of a BN's non-root nodes to the number of root nodes, or (ii) the expected number of moral edges in their moral graphs. Our approach is based on combining analytical and experimental results. Analytically, we partition the set of cliques in a clique tree into different sets, and introduce a growth curve for each set. For the special case of bipartite BNs, we consequently have two growth curves, a mixed clique growth curve and a root clique growth curve. In experiments, we systematically increase the degree of the root nodes in bipartite Bayesian networks, and find that root clique growth is well-approximated by Gompertz growth curves. It is believed that this research improves the understanding of the scaling behavior of clique tree clustering, provides a foundation for benchmarking and developing improved BN inference and machine learning algorithms, and presents an aid for analytical trade-off studies of clique tree clustering using growth curves.

Mengshoel, Ole Jakob

Finite-difference simulation and visualization of elastodynamics in time-evolving generalized curvilinear coordinates

Modeling and simulation of free and forced structural vibrations is essential to an overall structural health monitoring capability. In the various embodiments, a first principles finite-difference approach is adopted in modeling a structural subsystem such as a mechanical gear by solving elastodynamic equations in generalized curvilinear coordinates. Such a capability to generate a dynamic structural response is widely applicable in a variety of structural health monitoring systems. This capability (1) will lead to an understanding of the dynamic behavior of a structural system and hence its improved design, (2) will generate a sufficiently large space of normal and damage solutions that can be used by machine learning algorithms to detect anomalous system behavior and achieve a system design optimization and (3) will lead to an optimal sensor placement strategy, based on the identification of local stress maxima all over the domain.

Kaul, Upender K.

Collaborative Supervised Learning for Sensor Networks

Collaboration methods for distributed machine-learning algorithms involve the specification of communication protocols for the learners, which can query other learners and/or broadcast their findings preemptively. Each learner incorporates information from its neighbors into its own training set, and they are thereby able to bootstrap each other to higher performance. Each learner resides at a different node in the sensor network and makes observations (collects data) independently of the other learners. After being seeded with an initial labeled training set, each learner proceeds to learn in an iterative fashion. New data is collected and classified. The learner can then either broadcast its most confident classifications for use by other learners, or can query neighbors for their classifications of its least confident items. As such, collaborative learning combines elements of both passive (broadcast) and active (query) learning. It also uses ideas from ensemble learning to combine the multiple responses to a given query into a single useful label. This approach has been evaluated against current non-collaborative alternatives, including training a single classifier and deploying it at all nodes with no further learning possible, and permitting learners to learn from their own most confident judgments, absent interaction with their neighbors. On several data sets, it has been consistently found that active collaboration is the best strategy for a distributed learner network. The main advantages include the ability for learning to take place autonomously by collaboration rather than by requiring intervention from an oracle (usually human), and also the ability to learn in a distributed environment, permitting decisions to be made in situ and to yield faster response time.

Wagstaff, Kiri L.