Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Gaussian Process for Flight Delay Prediction: Learning a Stochastic Process

This paper presents a machine-learning approach to predict flight delays. Whereas neural networks are extensively studied for predictive capabilities, they involve non-intuitive design and extensive analysis, particularly in training and optimization processes. Instead, the proposed framework employs Gaussian Processes as a supervised learning technique for flight delay prediction. This data-driven approach trains the model using prior information, specifically the mean and covariance tied to existing data. The proposed Gaussian Process Regression (GPR) model employs the day of flight as a pivotal feature for delay forecasting. We analyze flights from various routes and gauge the accuracy of the presented learning technique by comparing the predicted delays with the actual ones. Given the inherent challenges in precisely forecasting delays, we predict the delays with a 95 % confidence interval. Also, an error propagation analysis in the prediction horizon is carried out to determine the optimal time frame for prediction. The proposed method for flight delay prediction is important as airlines can strategize flight operations and issue timely advisories.

stochastic↗

Celestial Mapping System and Digital Lunar Library Initiative

We are preparing to create an interactive, global 3D lunar environment with integrated dataset and AI/ML tools to provide unique value to mission planners, scientists and the entire lunar community. This lunar environment will be based on NASA Ames Celestial Mapping System (CMS) [1] and Digital Lunar Library (DLL) Initiative. CMS provides a 3D virtual Lunar Globe with extensive user friendly tool sets, that include high resolution terrain visualization, elevation profiles, measurement kits, slope analysis, path optimization, line of sight analysis, equipment planning and placement tools and many other functionalities [1]. It has a thick client with less overhead to access hardware resources. This allows features such as terrain profiling and distance calculations to be performed on the client and on the fly. The application is developed to provide situational and domain awareness on the Lunar surface, planning capabilities for equipment placements and traverse path optimization. As data becomes available, CMS has the capabilities to integrate data sets that change dynamically in real-time, which will be useful for monitoring satellites and remotely-sensed data on the Lunar surface. CMS supports importing synthetic features in a variety of 3D, 2D, vector and raster formats. In the future, these capabilities will be enhanced by incorporating AI/ML tools and a plug-in architecture to enable customization by the user groups. With the help of DLL we will be able to : 1) Amplify the value of lunar information with AI-powered data enhancements 2) Acquire and integrate lunar data with AI-assisted georectification and homogenization 3) Analyze lunar data with advanced 3D visualization, intelligent search-by-example 4) Apply lunar data insights to specific use cases with an open plug-in architecture. The CMS-DLL initiative will have several potential use cases for NASA and the lunar community in general, including subsurface lava tube visualization and analysis, soil analysis, in-situ lunar resource visualization and representation on 3D globe, and data analytics for utilization. REFERENCES: [1] https://celestial.arc.nasa.gov/

3D Globe↗

Building a landslide hazard indicator with machine learning and land surface models

The U.S. Pacific Northwest has a history of frequent and occasionally deadly landslides caused by various factors. Using a multivariate, machine-learning approach, we combined a Pacific Northwest Landslide Inventory with a 36-year gridded hydrologic dataset from the National Climate Assessment – Land Data Assimilation System to produce a landslide hazard indicator (LHI) on a daily 0.125-degree grid. The LHI identified where and when landslides were most probable over the years 1979–2016, addressing issues of bias and completeness that muddy the analysis of multi-decadal landslide inventories. The seasonal cycle was strong along the west coast, with a peak in the winter, but weaker east of the Cascade Range. This lagging indicator can fill gaps in the observational record to identify the seasonality of landslides over a large spatiotemporal domain and show how landslide hazard has responded to a changing climate.

XGBoost↗

Machine Learning for NASA Advanced Information Systems

NASA's Advanced Information Systems Technology (AIST) Program is one of several Technology programs managed by the Earth Science Technology Office (ESTO) in the Earth Science Division (ESD). The AIST Program focuses on advanced information systems and novel computer science technologies that will be needed by NASA Earth Science in the next 5 to 10 years. The three main thrusts of the AIST Program deal with Novel Observing Strategies (NOS), Analytic Collaborative Frameworks (ACF) and Earth System Digital Twins (ESDT). For all these thrusts, Machine Learning (ML) is increasingly being used in multiple aspects of Earth science systems, e.g., for onboard autonomy and decision making, for the analysis of massive and diverse datasets as well as more recently for developing surrogate models that will represent one of the main components of future Digital Twins of the Earth. Particularly, ESDT technologies developed by the AIST Program will allow to develop integrated Earth Science frameworks that will mirror the Earth with state-of-the-art models (Earth system models and others), timely and relevant observations, and analytic tools. These information systems will be used for supporting near- and long-term science and policy decisions. ESDT frameworks will build on previously developed AIST capabilities and technologies to integrate interconnected models with continuous streams of observations, data analytics, data assimilation, simulations, advanced visualizations and the ability to conduct "what-if" scenarios. This talk will describe the three thrusts of the AIST Program with a special focus on Machine Learning and how it is being used at all steps of the Earth Science data lifecycle.

Mathematical and Computer Sciences (General)↗

Software Development Cost Estimation Executive Summary

Identify simple fully validated cost models that provide estimation uncertainty with cost estimate. Based on COCOMO variable set. Use machine learning techniques to determine: a) Minimum number of cost drivers required for NASA domain based cost models; b) Minimum number of data records required and c) Estimation Uncertainty. Build a repository of software cost estimation information. Coordinating tool development and data collection with: a) Tasks funded by PA&E Cost Analysis; b) IV&V Effort Estimation Task and c) NASA SEPG activities.

data mining↗

Big-data Efficient and Automated Science Transfer (BEAST): An Open-Source Software Architecture for Arc Jet Data Management, Modeling, and Automation

Big-data Efficient and Automated Science Transfer (BEAST) was conceived to address the existing ground testing data management of the NASA Ames arc jet facilities (e.g., manually entered Excel files and USB drive data transfers). These data management practices were seen as a choke point for future thermal protection system (TPS) development as they limit statistical tracking, resolution of diagnostics, coordination between video/time series, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management↗

Big-data Efficient and Automated Science Transfer (BEAST): An Open-Source Software Architecture for Arc Jet Data Management, Modeling, and Automation

Big-data Efficient and Automated Science Transfer (BEAST) is a facility data management application developed for the NASA Ames arc jet facilities. The current decentralized data management practices limit statistical tracking, synchronization between video/time series, search capability, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management↗

Big-data Efficient Automated Science Transfer (BEAST): an open-source software architecture for arc jet data management, modeling, and automation

Big-data Efficient and Automated Science Transfer (BEAST) was conceived to address the existing ground testing data management of the NASA Ames arc jet facilities (e.g., manually entered Excel files and USB drive data transfers). These data management practices were seen as a choke point for future thermal protection system (TPS) development as they limit statistical tracking, resolution of diagnostics, coordination between video/time series, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management↗

Time series comparisons in Deep Space Network

The Deep Space Network (DSN) is NASA’s international array of antennas that support interplanetary spacecraft missions. DSN provides radar and radio astronomy observations that enhance our understanding of the solar system and the larger universe. A track is a block of continuous multi-dimensional time series from the beginning to end of DSN communication with the target spacecraft, containing 129 monitor data items lasting several hours at a frequency of 0.2-1Hz. Monitor data on each track reports on the performance of specific spacecraft operations and the DSN itself. DSN is receiving signals from 32 spacecraft across the solar system. DSN has pressure to reduce costs while maintaining the quality of support for DSN mission users. DSN operators need to simultaneously monitor multiple tracks and identify anomalies in real time. DSN has seen that as the number of missions increases, the data that needs to be processed increases over time. In this project, we look at the last 8 years of data for analysis. Any anomaly in the track indicates a problem with either the spacecraft, DSN equipment, or weather conditions. DSN operators typically write “discrepancy reports” for further analysis. It is recognized that it would be quite helpful to identify 10 similar historical tracks out of the huge database to quickly find/match anomalies. This tool has three functions: (1) identification of the top 10 similar historical tracks, (2) detection of anomalies compared to the reference normal track, and (3) comparison of statistical differences between two given tracks. The requirements for these features were confirmed by survey responses from 21 DSN operators and engineers. The preliminary machine learning model has shown promising performance (AUC=0.92). We plan to increase the number of data sets and perform additional testing to improve performance further before its planned integration into the Track Visualizer to assist DSN field operators and engineers.

Rebbapragada, Umaa↗

A pattern recognition system for locating small volvanoes in Magellan SAR images of Venus

The Magellan data set constitutes an example of the large volumes of data that today's instruments can collect, providing more detail of Venus than was previously available from Pioneer Venus, Venera 15/16, or ground-based radar observations put together. However, data analysis technology has not kept pace with data collection and storage technology. Due to the sheer size of the data, complete and comprehensive scientific analysis of such large volumes of image data is no longer feasible without the use of computational aids. Our progress towards developing a pattern recognition system for aiding in the detection and cataloging of small-scale natural features in large collections of images is reported. Combining classical image processing, machine learning, and a graphical user interface, the detection of the 'small-shield' volcanoes (less than 15km in diameter) that constitute the most abundant visible geologic feature in the more that 30,000 synthetic aperture radar (SAR) images of the surface of Venus are initially targeted. Our eventual goal is to provide a general, trainable tool for locating small-scale features where scientists specify what to look for simply by providing examples and attributes of interest to measure. This contrasts with the traditional approach of developing problem specific programs for detecting Specific patterns. The approach and initial results in the specific context of locating small volcanoes is reported. It is estimated, based on extrapolating from previous studies and knowledge of the underlying geologic processes, that there should be on the order of 10(exp 5) to 10(exp 6) of these volcanoes visible in the Magellan data. Identifying and studying these volcanoes is fundamental to a proper understanding of the geologic evolution of Venus. However, locating and parameterizing them in a manual manner is forbiddingly time-consuming. Hence, the development of techniques to partially automate this task were undertaken. The primary constraints for this particular problem are that the method must be reasonably robust and fast. Unlike most geological features, the small volcanoes of Venus can be ascribed to a basic process that produces features with a short list of readily defined characteristics differing significantly from other surface features on Venus. For pattern recognition purposes the relevant criteria include (1) a circular planimetric outline, (2) known diameter frequency distribution from preliminary studies, (3) a limited number of basic morphological shapes, and (4) the common occurrence of a single, circular summit pit at the center of the edifice.

Burl, M. C.↗

Remote Sensing of Lineage Functional Types for Modeling and Monitoring Biodiversity

Hyperspectral remote sensing has the potential to continuously scale plant function and plant diversity information from landscape to global extents. Numerous studies have indicated that VSWIR (400-2500 nm) reflectance properties of vegetation capture evolutionarily conserved biochemical, structural, and other functional attributes of plant species. Spectral properties conserved in plants provide the opportunity to both 1) aggregate species into lineages with improved classification accuracy and 2) link those lineages directly to plant traits. Full realization of this goal will enable parameterization of Land Surface Models (LSMs) with remotely sensed information, e.g., canopy nitrogen, and better representations of biodiversity and functional diversity in biogeographic studies. In this study, we use hyperspectral AVIRIS data from the 2013 HyspIRI campaign over the Southern Sierra Nevada, California flight box to investigate the potential for incorporating evolutionary thinking into landcover classification. We link the airborne hyperspectral data with vegetation plot data from roughly 1372 surveys and a phylogeny representing 1361 species. We aggregate species into lineages ranging from species level groups down to similar number of Plant Functional Types as often used in LSMs. We assessed the ability of Random Forest and Partial Least Squares Discriminant Analysis to discriminate across these different phylogenetic scales and determine the optimal number of lineages to classify. Although there are some temporal and spatial differences in our training data, our best approaches achieved moderate classification accuracy (Kappa > 0.65). Given an optimal number of lineages, we explored approaches to improve classifications including machine learning and unmixing approaches. This work suggests that lineage-based methods may be a promising way to leverage the huge amounts of data that will come from high resolution and high return interval hyperspectral data planned for the Surface Biology and Geology mission with sparsely sampled existing ground-based ecological data.

Hyperspectral↗

Toward Intelligent Software Defect Detection

Source code level software defect detection has gone from state of the art to a software engineering best practice. Automated code analysis tools streamline many of the aspects of formal code inspections but have the drawback of being difficult to construct and either prone to false positives or severely limited in the set of defects that can be detected. Machine learning technology provides the promise of learning software defects by example, easing construction of detectors and broadening the range of defects that can be found. Pinpointing software defects with the same level of granularity as prominent source code analysis tools distinguishes this research from past efforts, which focused on analyzing software engineering metrics data with granularity limited to that of a particular function rather than a line of code.

Benson, Markland J.↗

Future of Big Earth Data Analytics

The state of the art of Big Earth Data Analytics can be expected to evolve rapidly in the coming years. The forces driving evolution come from both growth in the data and advancement in the field of data analytics. In the data area, advances in sensor instrumentation and platform miniaturization are increasing both data resolution and coverage, resulting in enormous growth in data Volume. Increases in temporal resolution in particular also generate demands for higher data Velocity. At the same time, the proliferation of instruments and the platforms on which they reside is increasing the Variety of datasets. The Variety increase in turn leads to questions about the Veracity of the data. In the algorithm area, powerful machine learning methods are coming to the fore, particularly Deep Neural Networks. These are powerful at detecting interesting features in the data, integrating many different measurements (i.e., data fusion), and classification problems. However, they are still challenging when seeking explanations of how natural or socio-economic phenomena work using Earth Observations. Thus, classical analysis techniques will remain relevant when the emphasis is on forming or testing explanations, as well as to support interactive data exploration.

Lynnes, Christopher↗

SWIPE: Spectral Water Inversion Processor and Emulator

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This presentation will be discussing the progress made developing SWIPE: Spectral Water Inversion Processor and Emulator. SWIPE is a platform for advanced modeling of coastal and inland aquatic habitats. The goal is create a comprehensive and cohesive system to leverage recent advancements in computation and machine learning to develop a synthetic training ground for sensitivity studies and algorithm development. The four principal facets of SWIPE include: 1. Advanced two-layer coated sphere bio-optical modeling and GPU radiative transfer modeling, 2. Big Data involving massive synthetic spectral libraries of optical properties of various global aquatic particles, surface reflectance, and top-of-atmosphere reflectance, all at hyperspectral resolution leveraging high-end computing systems at NASA Ames Research Center, 3. Deep Learning for algorithm development for water quality inversion of concentrations of common biogeophysical variables as well as optics, full uncertainty characterization by water type, and forward emulation, and lastly, 4. Image Processing for application of developed retrieval algorithms for both hyperspectral and multispectral sensors with experimental corrections for global adjacency, noise, sunglint, and benthic reflectance. This presentation will demonstrate the Equivalent Algal Populations (EAP) two-layer coated sphere scattering model which has been used develop spectral libraries of hyperspectral inherent optical properties of roughly 80 species of phytoplankton, covering 15 different classes and nine taxonomic functional types. The EAP model was also used to derive spectral properties of 10 different non-algal particle functional types. Examples of how the SMART-G (Speed-up Monte-carlo Advanced Radiative Transfer using GPU) radiative transfer code is used to model optically complex aquatic signals will be presented and discussed in the context of creating a massive synthetic database which can leverage the full power of next generation machine learning techniques and high end computing for water quality inversion. We will discuss our active investigation in things like appropriate model architectures, dimensionality reduction techniques such as PCA and autoencoders, uncertainty quantification and abstaining, and which variables actually benefit most from hyperspectral information versus multispectral resolution. We are also curious about questions relating to cost/benefit analysis in terms of computation resources, neural network complexity, and data volumes. Answers to these questions will hopefully elaborate on cost efficiency for potential future sensor design considerations.

SWIPE↗

Advanced Analytics and Big Earth Data

NASA's Earth Science Data Systems process, archive and distribute petabytes of Earth Observation data to a variety of end users. These end users will face dramatically increased data size in the near future, bringing about new challenges and opportunities in analyzing those data. One area of particular ferment currently is Machine Learning. Many Machine Learning methods are black boxes, limiting direct insight into the data's properties. However, they can be used for a variety of data enhancement purposes, such as parameter retrieval, data fusion and image classification and segmentation. The Earth Observing System Data and Information System is also evolving to host large data volumes in the cloud, enabling data proximal analysis. As part of this effort, an Analytics framework is being developed to support and enhance user analysis of the data. By using standards based services in the framework, diverse user communities can be served, while also allowing inter-system collaboration in the analysis process.

Cloud Computing↗

Human Factors in Accidents Involving Remotely Piloted Aircraft

This presentation examines human factors that contribute to RPA mishaps and provides analysis of lessons learned. RPA accident data from U.S. military and government agencies were reviewed and analyzed to identify human factors issues. Common contributors to RPA mishaps fell into several major categories: cognitive factors (pilot workload), physiological factors (fatigue and stress), environmental factors (situational awareness), staffing factors (training and crew coordination), and design factors (human machine interface).

Merlin, Peter William↗

SkICAT: A cataloging and analysis tool for wide field imaging surveys

We describe an integrated system, SkICAT (Sky Image Cataloging and Analysis Tool), for the automated reduction and analysis of the Palomar Observatory-ST ScI Digitized Sky Survey. The Survey will consist of the complete digitization of the photographic Second Palomar Observatory Sky Survey (POSS-II) in three bands, comprising nearly three Terabytes of pixel data. SkICAT applies a combination of existing packages, including FOCAS for basic image detection and measurement and SAS for database management, as well as custom software, to the task of managing this wealth of data. One of the most novel aspects of the system is its method of object classification. Using state-of-theart machine learning classification techniques (GID3* and O-BTree), we have developed a powerful method for automatically distinguishing point sources from non-point sources and artifacts, achieving comparably accurate discrimination a full magnitude fainter than in previous Schmidt plate surveys. The learning algorithms produce decision trees for classification by examining instances of objects classified by eye on both plate and higher quality CCD data. The same techniques will be applied to perform higher-level object classification (e.g., of galaxy morphology) in the near future. Another key feature of the system is the facility to integrate the catalogs from multiple plates (and portions thereof) to construct a single catalog of uniform calibration and quality down to the faintest limits of the survey. SkICAT also provides a variety of data analysis and exploration tools for the scientific utilization of the resulting catalogs. We include initial results of applying this system to measure the counts and distribution of galaxies in two bands down to Bj is approximately 21 mag over an approximate 70 square degree multi-plate field from POSS-II. SkICAT is constructed in a modular and general fashion and should be readily adaptable to other large-scale imaging surveys.

Weir, N.↗

Exploring Anomalous PM 2.5 from Wildfires and Dust Storms using Data and Services at NASA GES DISC

The presence of fine particles in the atmosphere with a diameter of less than 2.5 µm, called particulate matter 2.5 (PM 2.5 ), poses a significant threat to human health as a criteria air pollutant. Fortunately, NASA's Goddard Earth Sciences Data and Information Services Center (GES DISC) provides easy access to several PM 2.5 concentration products. These datasets include the reanalysis of global hourly and monthly aerosol components including PM 2.5 data from the Modern-Era Retrospective analysis for Research and Applications, version 2 (MERRA-2), as well as 3-hourly real-time ensemble forecasts of PM 2.5 from the Hazardous Air Quality Ensemble System (HAQES). The HAQES products are developed by the George Mason University Air Quality Laboratory as part of NASA's Health Air Quality Applied Science Team (HAQAST). The GES DISC is actively collaborating with scientists in the HAQAST program to further expand air quality data collections. Two new datasets are currently being archived: one is the machine learning-based global hourly PM 2.5 derived from MERRA-2; the other is the localized data (NO 2 , O 3 , and PM 2.5 ) time series derived from NASA's GEOS Composition Forecasting (GEOS-CF) system. In this presentation, we will explore the spatial patterns and long-distance transport characteristics of elevated PM 2.5 during extreme pollution events, such as the June 2023 Canadian wildfires, which are still active at the time of writing; and severe spring dust storms in 2023 over Asia. To gain comprehensive insights, we will utilize various PM 2.5 data in conjunction with satellite-observed aerosol data from TROPOspheric Monitoring Instrument (TROPOMI) on Sentinel-5P. The primary focus of this presentation will be to demonstrate effective use of data tools and services to visualize and explore extreme air pollution phenomena. Additionally, we will provide guidance on how users can download specific data of interest, facilitating further analysis and research in this critical area.

air quality↗