Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Numerical Control Machine Data Manual

Numerical Control Machine Data Manual provides programmers with specific information for various types and sizes of numerical control machine tools and auxiliary equipment.

Mackey, R. T., Sr.↗

MLtool: Universal Supervised Machine Learning Tool to Model Tabulated Data

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine learning↗

A Robust Schema for Storing and Managing Machine Learning Data and Models

- Machine Learning (ML) has enabled models that can improve efficiency and decrease computational cost - ML models are crucial in enabling Integrated Computational Materials Engineering (ICME) - Large data sets require robust means of storing ML data and models

Brandon L. Hearley↗

Urban land use monitoring from computer-implemented processing of airborne multispectral data

Machine processing techniques were applied to multispectral data obtained from airborne scanners at an elevation of 600 meters over central Indianapolis in August, 1972. Computer analysis of these spectral data indicate that roads (two types), roof tops (three types), dense grass (two types), sparse grass (two types), trees, bare soil, and water (two types) can be accurately identified. Using computers, it is possible to determine land uses from analysis of type, size, shape, and spatial associations of earth surface images identified from multispectral data. Land use data developed through machine processing techniques can be programmed to monitor land use changes, simulate land use conditions, and provide impact statistics that are required to analyze stresses placed on spatial systems.

Todd, W. J.↗

A survey of machine readable data bases

Forty-two of the machine readable data bases available to the technologist and researcher in the natural sciences and engineering are described and compared with the data bases and date base services offered by NASA.

Matlock, P.↗

The TOLNet 2.0 Website: How an API Can Promote Open Science and FAIR Principles

The Tropospheric Ozone Lidar Network (TOLNet) has generated over a decade of ozone vertical profile data products over North America. The science value of the TOLNet data has been demonstrated in numerous peer-reviewed publications on air quality and other ozone relevant research. To support the broad spectrum of data use, the TOLNet team launched a major effort to upgrade the web-based data repository aiming to enhance the data discoverability and to enable machine-to-machine data upload and download processes. Specifically, the TOLNet website included an application programming interface (API), which supports machine-to-machine data search and data download. The API also extracts selected variables from the files, which can be retrieved as JSON objects and used to create data displays without having to download or open the underlying files. The TOLNet science team members can also use the API for automated data upload, including a data file scanning feature to ensure data product integrity. To be presented will include a summary of key features of data repositories, an actual use case of machine-to-machine data access/use, as well as our journey to make TOLNet data more FAIR, i.e., more findable, accessible, interoperable, and (re)usable.

Crystal Gummo↗

Toolsets for Airborne Data (TAD): Improving Machine Readability for ICARTT Data Files

The Atmospheric Science Data Center (ASDC) at NASA Langley Research Center is responsible for the ingest, archive, and distribution of NASA Earth Science data in the areas of radiation budget, clouds, aerosols, and tropospheric chemistry. The ASDC specializes in atmospheric data that is important to understanding the causes and processes of global climate change and the consequences of human activities on the climate. The ASDC currently supports more than 44 projects and has over 1,700 archived data sets, which increase daily. ASDC customers include scientists, researchers, federal, state, and local governments, academia, industry, and application users, the remote sensing community, and the general public.

Early, Amanda Benson↗

ImageLabler: Labeling and Managing Image Data for Machine Learning in the Earth Sciences

While machine learning techniques for image classification have been around for a long time, storing and managing the vast number of images required as training data is still a problem for scientists. This is especially true for the field of Earth science, where only recently have experts begun using machine learning techniques for image-based phenomena classification. Image Labeler, a fast and scalable cloud-based tagging platform for Earth science images, seeks to improve upon existing methods of managing images and associated metadata, such as maintaining categorized folders of images on a local machine, a process that can be cumbersome and difficult to scale. The platform facilitates rapid development of image-based Earth science phenomena training datasets by allowing scientists to upload their existing imagery as well as extract new samples from open satellite imagery services made available through NASA’s Global Imagery Browse Service (GIBS). Image Labeler also supports GeoTIFF data, with capabilities such as displaying GeoTIFFs on an interactive map, drawing shapefiles over them, and tagging them with additional metadata. This allows scientists to perform spatiotemporal subsetting with geographic information and develop training data more quickly. Built using modern web technologies, Image Labeler includes additional capabilities such as team collaboration for large-scale image tagging projects. Users can download their data in a machine-learning-ready format, allowing scientists to spend time on experimentation rather than on the collection of training data. In this presentation, we demonstrate how Image Labeler seeks to become a one-stop image data management solution for machine learning applications in Earth science.

Ashish Acharya↗

High speed machining of space shuttle external tank liquid hydrogen barrel panel

Actual and projected optimum High Speed Machining data for producing shuttle external tank liquid hydrogen barrel panels of aluminum alloy 2219-T87 are reported. The data included various machining parameters; e.g., spindle speeds, cutting speed, table feed, chip load, metal removal rate, horsepower, cutting efficiency, cutter wear (lack of) and chip removal methods.

Hankins, J. D.↗

Sub-Continental-Scale Carbon Stocks of Individual Trees in African Drylands

The distribution of dryland trees and their density, cover, size, mass and carbon content are not well known at sub-continental to continental scales. This information is important for ecological protection, carbon accounting, climate mitigation and restoration efforts of dryland ecosystems. We assessed more than 9.9 billion trees derived from more than 300,000 satellite images, covering semi-arid sub-Saharan Africa north of the Equator. We attributed wood, foliage and root carbon to every tree in the 0–1,000 mm year −1 rainfall zone by coupling field data, machine learning, satellite data and high-performance computing. Average carbon stocks of individual trees ranged from 0.54 Mg C ha −1 and 63 kg C tree −1 in the arid zone to 3.7 Mg C ha −1 and 98 kg tree −1 in the sub-humid zone. Overall, we estimated the total carbon for our study area to be 0.84 (±19.8%) Pg C. Comparisons with 14 previous TRENDY numerical simulation studies23 for our area found that the density and carbon stocks of scattered trees have been underestimated by three models and overestimated by 11 models, respectively. This benchmarking can help understand the carbon cycle and address concerns about land degradation. We make available a linked database of wood mass, foliage mass, root mass and carbon stock of each tree for scientists, policymakers, dryland-restoration practitioners and farmers, who can use it to estimate farmland tree carbon stocks from tablets or laptops.

Compton Tucker↗

Construction of a Fluid Flowfield from Discrete Point Data using Machine Learning

Many verification and validation procedures in aerospace engineering involve the comparison of computational fluid dynamics (CFD) data to experimental results from sources like wind tunnel tests. However, an incongruity exists between the data available from these sources: flow visualization is available by default in computational data, whereas in most experimental setups the available data is far more discrete and far more limited: integrated forces and moments, discrete pressure and temperature probes, etc. When differences exist between quantities of interest like lift and drag coefficients, the lack of full-field flow data from the experiments complicates most attempts to reconcile why the different data sources disagree. To this end, a shallow neural network, constrained by certain fluid flow properties, was trained to approximate flow field snapshots given only discrete data like that available in a wind tunnel test. The constructed snapshots, even for complex incompressible fluid flows, were found to agree at the large scales with the true flow fields. With this tool, researchers can more readily and easily understand why quantities of interest differ between their experimental and computational datasets. This in turn improves the resulting data's uncertainty measures.

Yury Lebedev↗

Construction of a Fluid Flowfield from Discrete Point Data using Machine Learning

Many verification and validation procedures in aerospace engineering involve the comparison of computational fluid dynamics (CFD) data to experimental results from sources like wind tunnel tests. However, an incongruity exists between the data available from these sources: flow visualization is available by default in computational data, whereas in most experimental setups the available data is far more discrete and far more limited: integrated forces and moments, discrete pressure and temperature probes, etc. When differences exist between quantities of interest like lift and drag coefficients, the lack of full-field flow data from the experiments complicates most attempts to reconcile why the different data sources disagree. To this end, a shallow neural network, constrained by certain fluid flow properties, was trained to approximate flow field snapshots given only discrete data like that available in a wind tunnel test. The constructed snapshots, even for complex incompressible fluid flows, were found to agree at the large scales with the true flow fields. With this tool, researchers can more readily and easily understand why quantities of interest differ between their experimental and computational datasets. This in turn improves the resulting data's uncertainty measures.

Yury Lebedev↗

Space Flown Rodent Liver RNA Sequencing Data for Machine Learning in Space Biology Research

High-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. Data analysis has been accelerated in recent years by the adoption of artificial intelligence (AI) and machine learning (ML) techniques by biomedical researchers. In space biology research, RNAseq datasets from space-flown experimental samples are critical for characterizing the gene expression aberrations associated with exposure to spaceflight stressors. However, space biological experiments tend to be very low sample size, so identifying proper AI/ML algorithms for sequencing data analysis is an ongoing challenge since these algorithms typically require large sample size. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML”, focused on creating datasets meant for three main applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. These scientific benchmarks consist of an AI-ready dataset and a reference implementation on a specific scientific question. In this work, we focused on generating standardized datasets to allow the scientific community to benchmark AI/ML algorithms in the domain of space biology. We present here a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data as a collaboration between the NASA AI4LS (Artificial Intelligence for Life Sciences) working group. and NASA’s SMD. This dataset consists of space-flown and ground control mouse liver found in the NASA GeneLab omics database. However, to amplify the small sample number (n=112 samples) for ML purposes, we employ Gaussian noise and a generative adversarial network to extend this dataset to 6,000 synthetic samples, matching the original gene expression characteristics.

James Casaletto↗

Estimation and Bias Correction of Aerosol Abundance using Data-driven Machine Learning and Remote Sensing

Air quality information is increasingly becoming a public health concern, since some of the aerosol particles pose harmful effects to peoples health. One widely available metric of aerosol abundance is the aerosol optical depth (AOD). The AOD is the integrated light extinction coefficient over a vertical atmospheric column of unit cross section, which represents the extent to which the aerosols in that vertical profile prevent the transmission of light by absorption or scattering. The comparison between the AOD measured from the ground-based Aerosol Robotic Network (AERONET) system and the satellite MODIS instruments at 550 nm shows that there is a bias between the two data products. We performed a comprehensive analysis exploring possible factors which may be contributing to the inter-instrumental bias between MODIS and AERONET. The analysis used several measured variables, including the MODIS AOD, as input in order to train a neural network in regression mode to predict the AERONET AOD values. This not only allowed us to obtain an estimate, but also allowed us to infer the optimal sets of variables that played an important role in the prediction. In addition, we applied machine learning to infer the global abundance of ground level PM2.5 from the AOD data and other ancillary satellite and meteorology products. This research is part of our goal to provide air quality information, which can also be useful for global epidemiology studies.

Malakar, Nabin K.↗

Mapping Inundation from Hurricane Florence (2018) with L-Band Synthetic Aperture Radar, Commercial Imagery, and Ancillary Data via Machine Learning Classification

During and after flooding events, mapping the extent of floodwaters aids in the distribution of resources, recovery efforts, and damage assessment practices. Development of a land cover classification system focused on mapping inundation after major hurricane events using synthetic aperture radar (SAR) data could allow for the production of near-real-time inundation mapping, enabling government and emergency response entities to get a preliminary idea of a developing situation. Complimentary optical and SAR images from domestic and foreign entities are brought together through activations of the International Charter: Space and Major Disasters to support response efforts, from true-color, near-infrared, and thermal remote sensing data obtained by NASA, NOAA, and international satellites to the collection of high-resolution true color aerial photography by NOAA and the National Geodetic Survey. In response to Hurricane Florence of 2018, NASA JPL collected numerous swaths of quad-pol L-band SAR data with the Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) instrument observing the record-setting river stages across North and South Carolina. The resulting fully-polarized SAR images allow for mapping of inundation extent at a high spatial resolution with a unique advantage over optical imaging stemming from the sensor’s ability to penetrate cloud cover and dense vegetation. In this study, true-color NOAA aerial and commercial satellite imagery are used in conjunction with four UAVSAR data swaths centered on the Lumberton and Cape Fear River basins in southeastern North Carolina to develop a Random Forest classification model focused on mapping open water and floodwater otherwise obscured by vegetation or lingering cloud cover. Ancillary building footprint, transportation route, and population data will also be incorporated into the classification scheme to estimate the societal impacts of flooding based on the proximity of features to detected inundation. Preliminary results from the Hurricane Florence case study will be discussed in addition to the limitations of available validation data for assessment of the classifier’s accuracy.

Alexander M Melancon↗