Search NASASearch

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

An Analysis of Barriers Preventing the Widespread Adoption of Predictive and Prescriptive Maintenance in Aviation

The aviation industry has long recognized the potential benefits of predictive maintenance, a maintenance strategy that leverages sensor and operational data to predict the future degradation of components. Prescriptive maintenance takes this a step further and considers the entire aviation ecosystem to schedule maintenance actions optimally. With the ability to reduce maintenance costs by up to 30%, as reported by the Department of Energy, these maintenance strategies have been identified to be an important investment to reduce a airline costs. However, despite great interest and technological advances in areas such as diagnostics, prognostics, sensing, computation, and machine learning, the adoption of predictive and prescriptive maintenance has not been widely applied in aviation. To shed light on this issue, we conducted an analysis of the barriers preventing or limiting the adoption of predictive and prescriptive maintenance in aviation. Through discussions with subject matter experts across industry, academia, standards bodies, and government, we identified five key challenges: complexity of prediction; validation, safety assurance, and regulatory challenges; cost of adoption; difficulty in quantifying impact and informing decisions; and data availability, quality, and ownership challenges. This study provides a detailed overview of these barriers and areas where stakeholders could invest to overcome them, aiming to support the scaled adoption of predictive and prescriptive maintenance in aviation.

Christopher Teubert

Prediction of Pushback Times and Ramp Taxi Times for Departures at Charlotte Airport

When optimizing the takeoff sequence and schedule for departures at busy airports, it is important to accurately predict the taxi times from gate to runway because those are used to calculate the earliest possible takeoff times. Several airports like Charlotte Douglas International Airport show relatively long taxi times inside the ramp area with large variations, with respect to the travel times in the airport movement area. Also, the pushback process times have not been accurately modeled so far mainly due to the lack of accurate data. The recent deployment of the integrated arrival, departure, and surface traffic management system at Charlotte airport by NASA enables more accurate flight data in the airport surface operations to be obtained. Taking advantage of this system, actual pushback times and ramp taxi times from historical flight data at this airport are analyzed. Based on the analysis, a simple, data-driven prediction model is introduced for estimating pushback times and ramp transit times of individual departure flights. To evaluate the performance of this prediction model, several machine learning techniques are also applied to the same dataset. The prediction results show that the data-driven prediction model is as good as the machine learning algorithms when comparing various prediction performance metrics.

Lee, Hanbong

Prediction of Pushback Times and Ramp Taxi Times for Departures at Charlotte Airport

When optimizing the takeoff sequence and schedule for departures at busy airports, it is important to accurately predict the taxi times from gate to runway because those are used to calculate the earliest possible takeoff times. Several airports like Charlotte Douglas International Airport show relatively long taxi times inside the ramp area with large variations, with respect to the travel times in the airport movement area. Also, the pushback process times have not been accurately modeled so far mainly due to the lack of accurate data. The recent deployment of the integrated arrival, departure, and surface traffic management system at Charlotte airport by NASA enables more accurate flight data in the airport surface operations to be obtained. Taking advantage of this system, actual pushback times and ramp taxi times from historical flight data at this airport are analyzed. Based on the analysis, a simple, data-driven prediction model is introduced for estimating pushback times and ramp transit times of individual departure flights. To evaluate the performance of this prediction model, several machine learning techniques are also applied to the same dataset. The prediction results show that the data-driven prediction model is as good as the machine learning algorithms when comparing various prediction performance metrics.

airport surface operations

Towards a program of record of inland water quality: Exploiting present and heritage multispectral sensors for maximum information extraction

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This research exploits recent advancements in bio-optical modeling, cloud computing, and machine learning to enhance our capacity to leverage present and heritage satellite data. Recent research suggests that sensors with low spectral resolution, such as Sentinel 2 and Landsat missions, contain enough hidden spectral variation which can be exploited using data-driven approaches. The availability of three decades of archival imagery will open doors to discover global trends of eutrophication and increased cyanobacteria dominance and provide valuable insight to the development of predictive methodologies. Preliminary efforts in synthetic emulation of global natural inland waters will be discussed and contextualized against satellite radiometric measurement uncertainty, satellite data product uncertainties and causal signal ambiguity over the visible wavelength range, supported by high quality field and image data for selected inland aquatic sites. Insights on water quality estimation via data-driven machine learning models versus matrix inversions will be discussed, and how we can exploit spectral-spatial relationships in high spatial resolution data. A cross-sensor synergistic approach with detailed uncertainty analysis based on optical water types, will allow for unprecedented global snapshots of fine scale ecological dynamics of inland waters.

Inland

Integrated Modeling for Payload Test of the Roman Space Telescope

The Nancy Grace Roman Telescope (RST) is a NASA observatory designed to unravel the secrets of dark energy and dark matter, search for and image exoplanets, and explore many topics in infrared optics. Scheduled to launch in no earlier than October 2026, this 2.4 meter aperture telescope has a field of view 100 times greater than the Hubble Space Telescope. The mission is currently in its construction phase, where integrated modeling between thermal, structural, and optical models of the observatory is necessary to demonstrate science quality images over the range of operational parameters. This presentation discusses the most recent integrated modeling analysis cycle for Roman, including model correlation with our instrument level testing. We include a discussion on improved processes of the handling of the various flows of data between the modeling disciplines and discipline specific monte-carlo analysis predictions. We will finish with the predicted uncertainties and expected performance for our upcoming observatory alignment verification test using machine learning algorithms.

telescope

Investigating the Use of Machine Learning (ML) to Assess Tropospheric Doppler Radar Wind Profiler (TDRWP) Data Quality

Manual Quality Control (MQC) of Tropospheric Doppler Radar Wind Profiler (TDRWP) data is essential for defining an accurate climatology for downstream aerospace vehicle assessments. MQC traditionally takes around 30.5 hours per year of radar data. The Marshall Space Flight Center Natural Environments Branch (MSFC NE) used machine learning (ML) to test the feasibility of automating the MQC process, showing a potential to reduce labor by 300%. However, analysis of the model showed some false positives. We compared a neural network to the model to validate it and develop a process for assessing comparable solutions in the future.

Corey Walker

Introduction to Analysis Methods for Big Earth Data

Big Earth Data are too big to be tractable to simple data inspection. Thus, they typically require models to make sense of all the data. Useful models for Big Earth Data may be physical, statistical, or machine learning based. While physical models are ideal for understanding the data, they are not always feasible, particularly when our ability to observe at finer scales exceeds our ability to incorporate the physics. Statistical models are more generalized, but computationally intensive for many Earth Observation datasets. Machine Learning models generally scale well but are sometimes limited in the physical understanding they can offer. Hybrid models combine attributes—and advantages—of two or more of these types.

Christopher Lynnes

A Quantitative Analysis on the Use of Supervised Machine Learning in Earth Science

Recent review papers (Ball et al., 2017; Reichstein et al., 2019) have investigated the opportunities and challenges in applying supervised machine learning (ML) techniques to Earth science problems. A common challenge is the lack of training (or labeled) data. Supervised ML, and especially deep learning (DL), require large training datasets. While there are large, open access Earth science archives, the data typically require preprocessing in preparation for supervised ML, frequently including manual labeling. Our objective is to understand the landscape of supervised ML in the Earth sciences, including which research communities have most rapidly adopted supervised ML, which algorithms are applied, and what data are used to train these algorithms. We conducted a literature survey of Earth science papers published during the last 10 years in journals from the American Geophysical Union (AGU), American Meteorological Society (AMS), the Institute of Electrical and Electronics Engineers(IEEE), and the Society of Photo-Optical Instrumentation Engineers (SPIE). We identified papers containing the terms ML, DL, or the names of individual supervised ML algorithms. "Earth science" is an additional required search term for IEEE and SPIE. We investigate trends in supervised ML usage during the 10-year study period, and manually analyzed AGU papers from 2018-2019 to enable deep-dive statistics.

Katrina S Virts

Machine Learning Applications to Metal-Silicate Equilibria and their Insights into Core Formation

An extensive number of studies have experimentally investigated how elements distribute between metal and silicate phases, to better constrain core-mantle chemical equilibrium. Here, we present a new database compiling all (to our knowledge) experimental data on liquid metal-silicate partitioning from 118 peer-reviewed publications. We applied various machine learning techniques to gain further insights into these partitioning equilibria and their dependencies. We performed a network analysis to investigate the relationship between experiments and partition coefficients, which enables visualizing gaps in the experimental dataset and biases related to varying experimental conditions and analytical setup. In addition, semi-empirical thermodynamic models are commonly used to extrapolate these chemical reactions to the wide range of pressure, temperature and compositional conditions of planetary differentiation. These models are based on linear regressions that assume continuous relationship between partition coefficients and experimental variables. Here, we considered random forest regressions, which are algorithms based on ensembles of decision trees and does not consider continuous effects of each variable. The application of this regression significantly improves the prediction of metal-silicate partitioning for several elements including Ni, Si and Cr. We will show how this new approach improves our understanding of elemental exchange between metal and silicate and their implications for the Earth’s core formation.

siderophile element

Earth Independent Medical Operations (EIMO) Datascope: Challenges and Potential Solutions

Data flows and storage/retrieval capacity are severely constrained during missions in space and challenges will become even greater during exploration class missions. There is a need for an artificial intelligence (AI)-based clinical decision support system (CDSS) to monitor and analyze data to provide real-time consultative support for crew medical officer (CMO) decision-making. EIMO is defined as the gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. While a hallmark of this paradigm shift from low-earth orbit is that on-board care will increasingly become the responsibility of the astronauts for primary management and decision making, terrestrial assets will continue to be paramount in pre-mission screening and planning, as well as prevention, health maintenance and long-term care contingencies. New capabilities and systems that enable progressively more robust and resilient systems and crews will be necessary to reduce risk and increase probability of deep space exploration mission success. An aspiration for EIMO is to develop AI-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A “system of systems” approach is envisioned whereby EIMO will deploy AI-supported natural language processing and machine learning (ML) techniques to utilize embedded reference databases and real-time data streams [input vectors] from multiple data sources. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will feature mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats. Large amounts and variable sources of data can be leveraged to diagnose, inform treatment strategies, and potentially predict medical events and performance decrements. Inclusion of advanced training tools using extended reality will enable increasingly autonomous medical care to aid a CMO when ground support is unavailable or time-delayed beyond required action window, e.g., emergent medical situations. EIMO CDSS would require very large datasets to train pre-flight and significant amounts of data are needed to support ML via in-flight CDSS operations. An additional challenge will be to find sufficient data to train a model relevant to astronaut demographics. The rapid, accelerating evolution of this field creates a propitious solution space to leverage multi-modal AI through public-private partnership(s). The status of multi-modal AI systems today would preclude their use for long duration missions as they remain unreliable and are subject to “digital hallucinations” and other errors that could pose operational risk. A federated labs structure is being considered to test and optimize data flow from the multiple input vectors leading to field testing in suitable ground/flight analogs. Critical to the success of an EIMO CDSS will be integration and interoperability and success will be defined by a system that can serve as an in-flight medical consult for the CMO providing critical support during medical contingencies. Benefits to terrestrial medicine may be significant as an outflow of the EIMO medical system, particularly for remote areas and communities lacking significant infrastructure, personnel and resources.

J Lemery

Earth Independent Medical Operations (EIMO) Datascope: Challenges and Potential Solutions

Data flows and storage/retrieval capacity are severely constrained during missions in space and challenges will become even greater during exploration class missions. There is a need for an artificial intelligence (AI)-based clinical decision support system (CDSS) to monitor and analyze data to provide real-time consultative support for crew medical officer (CMO) decision-making. EIMO is defined as the gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. While a hallmark of this paradigm shift from low-earth orbit is that on-board care will increasingly become the responsibility of the astronauts for primary management and decision making, terrestrial assets will continue to be paramount in pre-mission screening and planning, as well as prevention, health maintenance and long-term care contingencies. New capabilities and systems that enable progressively more robust and resilient systems and crews will be necessary to reduce risk and increase probability of deep space exploration mission success. An aspiration for EIMO is to develop AI-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A “system of systems” approach is envisioned whereby EIMO will deploy AI-supported natural language processing and machine learning (ML) techniques to utilize embedded reference databases and real-time data streams [input vectors] from multiple data sources. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will feature mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats. Large amounts and variable sources of data can be leveraged to diagnose, inform treatment strategies, and potentially predict medical events and performance decrements. Inclusion of advanced training tools using extended reality will enable increasingly autonomous medical care to aid a CMO when ground support is unavailable or time-delayed beyond required action window, e.g., emergent medical situations. EIMO CDSS would require very large datasets to train pre-flight and significant amounts of data are needed to support ML via in-flight CDSS operations. An additional challenge will be to find sufficient data to train a model relevant to astronaut demographics. The rapid, accelerating evolution of this field creates a propitious solution space to leverage multi-modal AI through public-private partnership(s). The status of multi-modal AI systems today would preclude their use for long duration missions as they remain unreliable and are subject to “digital hallucinations” and other errors that could pose operational risk. A federated labs structure is being considered to test and optimize data flow from the multiple input vectors leading to field testing in suitable ground/flight analogs. Critical to the success of an EIMO CDSS will be integration and interoperability and success will be defined by a system that can serve as an in-flight medical consult for the CMO providing critical support during medical contingencies. Benefits to terrestrial medicine may be significant as an outflow of the EIMO medical system, particularly for remote areas and communities lacking significant infrastructure, personnel and resources.

Medical Operations

Machine learning based noise reduction for satellite products: application to solar-induced fluorescence retrievals using simulated and real data

In the past two decades, global satellite measurements of terrestrial chlorophyll solar-induced fluorescence (SIF) have been used widely for a number of different applications related to physiology, phenology, and productivity of plants. However, SIF retrievals are inherently noisy due to the relatively small SIF spectral signature in comparison with observational noise. In this work, we examine how a spectral-based approach that employs principal component analysis along with a relatively shallow artificial neural network can be used to reduce noise and other artifacts in satellite level 2 (L2) products. We first apply the approach in a controlled environment in which radiance spectra are simulated with a full atmospheric and surface radiative transfer model for different scenarios including various SIF values that are known. Various levels of noise can be added to the simulated spectra. Resulting noisy and noise-reduced SIF retrievals are compared with the true values to assess performance. We then apply the noise reduction approach to real SIF derived from instruments flying on meteorological satellites. The results are evaluated by comparing SIF retrievals from different platforms with each other and with other independent data sets, showing enhanced capability to capture seasonal and interannual variability in SIF.

Chlorophyll fluorescence

Machine Learning Enabled Quantitative Risk Assessment of Aerial Wildfire Response

Aerial wildfire operations are high risk and account for a large number of firefighter deaths. Increasing intensity of wildfires is driving a surge in aerial operations, while simultaneously there is growing interest in improving system safety and performance. In this work, wildfire aviation mishaps documented using the SAFECOM system are analyzed using a previously developed framework for hazard extraction and analysis of trends (HEAT). Hazards and specific failure modes are extracted from the narrative data in SAFECOM forms using natural language processing techniques. Metrics for each hazard are calculated, including frequency, rate, and severity. We examine whether these metrics change over time, and whether they are related to metadata, such as region and aircraft type. The results of the hazard analysis are presented in a risk matrix, identifying the highest and lowest risk hazards based on rate of occurrence and average severity. Results identify jumper operations hazards as high-risk, in addition to bucket drop failures, cargo let down failures, and severe weather as medium risk.

machine learning

Machine Learning Enabled Quantitative Risk Assessment of Aerial Wildfire Response

Aerial wildfire operations are high risk and account for a large number of firefighter deaths. Increasing intensity of wildfires is driving a surge in aerial operations, while simultaneously there is growing interest in improving system safety and performance. In this work, wildfire aviation mishaps documented using the SAFECOM system are analyzed using a previously developed framework for hazard extraction and analysis of trends (HEAT). Hazards and specific failure modes are extracted from the narrative data in SAFECOM forms using natural language processing techniques. Metrics for each hazard are calculated, including frequency, rate, and severity. We examine whether these metrics change over time, and whether they are related to metadata, such as region and aircraft type. The results of the hazard analysis are presented in a risk matrix, identifying the highest and lowest risk hazards based on rate of occurrence and average severity. Results identify jumper operations hazards as high-risk, in addition to bucket drop failures, cargo let down failures, and severe weather as medium risk.

machine learning

PALMO: An OVERFLOW Machine Learning Airfoil Performance Database

The OVERFLOW Machine Learning Airfoil Performance (PALMO) database has been created to enable robust modeling of airfoil performance in a variety of applications. The database uses OVERFLOW simulation data second-order accurate in time and fourth-order accurate in space with Spalart-Allmaras turbulence closure. The foundation of the in-development PALMO database is the airfoil base cube. Each base cube includes simulation data parametrized over a range of Mach numbers, Reynolds numbers, and angles-of-attack. This first release of the database includes the NACA 4-series airfoils, with parametrization in airfoil thickness and camber from an NACA 0006 to an NACA 4424. In total, 52,480 NACA 4-series calculations were run on the NASA High-End Compute Capability (HECC) supercomputer and the corresponding airfoil performance coefficients are embedded in the Appendix of this document for public distribution. This provides high-order-accurate simulation data covering a wide range of aerospace design applications, which enables users to develop OVERFLOW-quality airfoil performance look-up tables without additional high-performance computing. In addition to engineering design and analysis of aerospace vehicles, PALMO is well suited to be a benchmark dataset for the development and testing of machine learning methods in aerospace engineering. Downstream surrogate models enable OVERFLOW- quality airfoil performance predictions for any arbitrary combination of camber, thickness, Mach number, Reynolds number, and angle-of-attack within the bounds of the database.

Database

Distributed Lunar Data Platform with Advanced Machine Learning Capabilities in Support of Lunar Science and Exploration

The United States 2020 Space Policy directive declares that NASA, in cooperation with private industry, will “extend human economic activity into deep space by establishing a permanent human presence on the Moon”. This goal will require advanced data management, as well as analysis, modeling and representation of lunar information in order to prepare for Artemis human missions, lunar science investigations and exploration. To meet this requirement, we conceptualize and present an implementation strategy for a distributed platform for lunar data retrieval, inferencing and analysis, which will be based on federated learning and the NASA Celestial Mapping System (CMS). In addition to demonstrating the imperative of enabling lunar-borne data to remain in-situ but still accessible, this presentation will also include examples of how third parties could contribute both datasets and new functionality into this platform using an AI-based data import pipeline and a plug-in architecture respectively.

Artificial Intelligence

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential in mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail walking and US-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986 to 2022. Using data from Landsat 5 and 8, Sentinel-2, the National Agriculture Imagery Program, and PlanetScope, the team computed vegetation indices including the Normalized Difference Vegetation Index, Normalized Difference Moisture Index, Modified Soil Adjusted Vegetation Index 2, Enhanced Vegetation Index, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on the vegetation indices and spectral bands before running k-means++ clustering and random forest classification algorithms. Between all datasets, we found the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use the end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants. The NPS’s collection of ground data for 2022–2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive spread over time.

Carson Schuetze

Automating sky object classification in astronomical survey images

We describe the application of machine classification techniques to the development of an automated tool for the reduction of a large scientific data set. The 2nd Palomer Observatory Sky Survey is nearly completed. This survey provides comprehensive coverage of the northern celestial hemisphere in the form of photographic plates. The plates are being transformed into digitized images whose quality will probably not be surpassed in the next ten to twenty years. The images are expected to contain on the order of 10(exp 7) galaxies and 10(exp 8) stars. Astronomers wish to determine which of these sky objects belong to various classes of galaxies and stars. The size of this data set precludes manual analysis. Our approach is to develop a software system which integrates the functions of independently developed techniques for image processing and data classification. Digitized sky images are passed through image processing routines to identify sky objects and to extract a set of features for each object. These routines are used to help select a useful set of attributes for classifying sky objects. Then GID3* and O-BTree, two inductive learning techniques, learn classification decision trees from examples. These classifiers will be used to process the rest of the data. This paper gives an overview of the machine learning techniques used, describes the details of our specific application, and reports the initial encouraging results. The results indicate that our approach is well-suited to the problem. The primary benefits of the approach are increased data reduction throughput and consistency of classification. The classification rules which are the product of the inductive learning techniques will form an object, examinable basis for classifying sky objects. A final, not to be underestimated benefit is that astronomers will be freed from the tedium of an intensely visual task to pursue more challenging analysis and interpretation problems based on automatically cataloged data.

Fayyad, Usama M.