Search NASA⌕ Search

SEARCH · Search NASA

Results for “Labeled Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Predicting Drug Effects from High-dimensional Asymmetric Drug Data Sets using Graph Neural Networks: A Comprehensive Analysis of Multi-target Drug Effect Prediction

Graph neural networks (GNNs) have emerged as one of the most effective Machine learning (ML) techniques for drug effect prediction from drug molecular graphs. Despite having immense potential, GNN models lack performance when using data sets that contain high dimensional asymmetrically co-occurrent drug effects as targets with complex correlations between them. Training individual learning models for each drug effect and incorporating every prediction result for a wide spectrum of drug effects is beyond practicality. Such an implication provides a testbed to address this challenge as multi-target prediction problems, aiming to predict all drug effects at a time. We develop standard and hybrid graph neural networks (GNNs)to perform two separate tasks that are multi-regression for continuous values and multi-label classification for categorical values contained in our data sets. Since this step makes the target data even more sparse and introduces asymmetric label co-occurrence, the learning of multi-label classification models becomes difficult and heavily impacts the GNN's performance. To address these challenges, we propose a new data oversampling technique to improve multi-label classification performances on all the given imbalanced molecular graph data sets. Using the technique, we improve the data imbalance ratio of the drug effects better than before while protecting the data set's integrity. Finally, we evaluate multi-label classification performance using the best-performant hybrid GNN model on all the oversampled data sets obtained from the proposed oversampling technique. These results outperform those of other ML models including GNN models when they are trained on the original data sets or oversampled data sets using MLSMOTE (a well-known oversampling technique) in all evaluation metrics precision, recall, and F1 score by a significant margin.

Bose, Avishek [ORNL]↗

A Machine Learning Approach to Objective Identification of Dust in Satellite Imagery

Airborne dust has broad adverse effects on human activity, including aviation, human health, and agriculture. Remote sensing observations are used to detect dust and aerosols in the atmosphere using long established techniques. False color Red-Green-Blue (RGB) imagery using band differences sensitive to dust absorption (Dust RGB) is currently used operationally to assist forecasters and decision-makers in identifying dust at night, but there are still limitations, subjectivity, and nuances to image interpretation making night-time dust identification difficult even for experts. This study applies machine learning to the problem of night-time dust detection with a simple random forest (RF) model using Geostationary Operational Environmental Satellite-16 (GOES-16) Advanced Baseline Imager (ABI) infrared imagery, band differences sensitive to dust absorption, and Dust RGB color components as inputs to the model. The RF model achieves an Area-Under-Curve (AUC) of 0.97 with a standard deviation of 0.04 for dust cases. For images with dust present, the model correctly labels 85% of dust pixels and 99.96% of no-dust pixels for all dust images in the validation data set. The addition of a single null case to the training data set drastically reduces error in labeling no-dust pixels as dust from 45% to 14.5%. Application of the machine learning model to the April 13–14, 2019 dust event demonstrates the ability of the model to identify dust during night-time hours when visual dust detection is limited by the cooling ground surface characteristics.

dust↗

Coincidence anomaly detection for unsupervised locating of edge localized modes in the DIII-D tokamak dataset

Using supervised learning to train a machine learning model to predict an on-coming edge localized mode (ELM) requires a large number of labeled samples. Creating an appropriate data set from the very large database of discharges at a long-running tokamak, such as DIII-D, would be a very time-consuming process for a human. Considering this need and difficulty, we use coincidence anomaly detection, an unsupervised learning technique, to train an ELM-identifier to identify and label ELMs in the DIII-D discharge database. This ELM-identifier shows, simultaneously, a precision of 0.68 and a recall of 0.63 (AUC is 0.73) on identifying ELMs in example time series pulled from thousands of discharges spanning five years. In a test set of 50 discharges, the algorithm finds over 26 thousand ELM candidates, more than 5 times the existing catalog of ELMs labeled by humans.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

User Guide for the Anvil Threat Cooridor Forecast Tool V2.4 for AWIPS

The Anvil Tool GUI allows users to select a Data Type, toggle the map refresh on/off, place labels, and choose the Profiler Type (source of the KSC 50 MHz profiler data), the Date- Time of the data, the Center of Plot, and the Station (location of the RAOB or 50 MHz profiler). If the Data Type is Models, the user selects a Fcst Hour (forecast hour) instead of Station. There are menus for User Profiles, Circle Label Options, and Frame Label Options. Labels can be placed near the center circle of the plot and/or at a specified distance and direction from the center of the circle (Center of Plot). The default selection for the map refresh is "ON". When the user creates a new Anvil Tool map with Refresh Map "ON, the plot is automatically displayed in the AWIPS frame. If another Anvil Tool map is already displayed and the user does not change the existing map number shown at the bottom of the GUI, the new Anvil Tool map will overwrite the old one. If the user turns the Refresh Map "OFF", the new Anvil Tool map is created but not automatically displayed. The user can still display the Anvil Tool map through the Maps dropdown menu* as shown in Figure 4.

Barett, Joe H., III↗

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY↗

Users manual for the US baseline corn and soybean segment classification procedure

A user's manual for the classification component of the FY-81 U.S. Corn and Soybean Pilot Experiment in the Foreign Commodity Production Forecasting Project of AgRISTARS is presented. This experiment is one of several major experiments in AgRISTARS designed to measure and advance the remote sensing technologies for cropland inventory. The classification procedure discussed is designed to produce segment proportion estimates for corn and soybeans in the U.S. Corn Belt (Iowa, Indiana, and Illinois) using LANDSAT data. The estimates are produced by an integrated Analyst/Machine procedure. The Analyst selects acquisitions, participates in stratification, and assigns crop labels to selected samples. In concert with the Analyst, the machine digitally preprocesses LANDSAT data to remove external effects, stratifies the data into field like units and into spectrally similar groups, statistically samples the data for Analyst labeling, and combines the labeled samples into a final estimate.

Horvath, R.↗

The IRGen infrared data base modeler

IRGen is a modeling system which creates three-dimensional IR data bases for real-time simulation of thermal IR sensors. Starting from a visual data base, IRGen computes the temperature and radiance of every data base surface with a user-specified thermal environment. The predicted gray shade of each surface is then computed from the user specified sensor characteristics. IRGen is based on first-principles models of heat transport and heat flux sources, and it accurately simulates the variations of IR imagery with time of day and with changing environmental conditions. The starting point for creating an IRGen data base is a visual faceted data base, in which every facet has been labeled with a material code. This code is an index into a material data base which contains surface and bulk thermal properties for the material. IRGen uses the material properties to compute the surface temperature at the specified time of day. IRGen also supports image generator features such as texturing and smooth shading, which greatly enhance image realism.

Bernstein, Uri↗

Single-cell proteomics of Arabidopsis leaf mesophyll reveals dynamic protein responses to water-deficit stress

Background The application of single-cell omics tools to biological systems can provide unique insights into diverse cellular populations and their heterogeneous responses to internal and external perturbations. Thus far, most single-cell studies in plant systems have been limited to RNA-sequencing approaches, which only provide indirect readouts of cellular functions. Results Here, we present a single-cell proteomics workflow for plant cells that integrates tape-sandwich protoplasting, piezoelectric cell sorting, nanoPOTS sample preparation, and ion mobility-based MS data acquisition method for label-free single-cell proteomics analysis of Arabidopsis leaf mesophyll cells. From a single leaf protoplast, over 3,000 proteins were quantified with high precision. The workflow is demonstrated to identify stress associated changes in protein abundance by analyzing 117 protoplasts from well-watered and water-deficit stressed plants. Additionally, we describe a new approach for constructing covarying protein networks at the single-cell level and demonstrate how single-cell protein covariation analysis can reveal previously unrecognized protein functions while also capturing stress-induced changes in protein–protein dynamics. Conclusions The label-free scProteomic approach presented here represents a significant advance through the demonstration of a facile protoplast isolation method combined with deep and precise proteomic coverage of Arabidopsis leaf mesophyll cell types. We believe this study will serve as an informative reference to future plant scProteomic investigations.

Arabidopsis↗

Dataset 1: A National and City Dataset on Human Factors in Pooled Rideshare, 2021

Dataset 1: A National and City Dataset on Human Factors in Pooled Rideshare, 2021. Dataset Description: Pooled Rideshare Acceptance Survey - Phase 1 (2021, N = 5,385). This dataset captures responses from a nationally representative sample of 5,385 adults across the United States to understand public acceptance, preferences, and behavioral intentions related to pooled rideshare (PR) services. The primary objective of this research is to provide actionable insights to inform the design, deployment, and policy development of sustainable shared mobility systems. Data was collected via an online survey administered through a national panel provider. Participants ranged in age from 18 to 95 years, and representation from all U.S. regions. The survey instrument was designed to explore numerous dimensions related to PR adoption including demographic traits, current travel habits, rideshare familiarity, trust, safety, environmental attitudes, and user experience preferences. Both rideshare users and non-users were included, offering a diverse range of perspectives. - Phase_1_Final - The dataset includes survey items developed from literature reviews, and prior field studies. Each row represents an individual respondent, and each column corresponds to a variable such as willingness to use pooled rideshare, attitudes toward specific service features, and sociodemographic data. The data is available in both .CSV and .SAV formats. - Phase_1_Final_MapFile - The accompanying data dictionary explains all variable labels, response scales, and codes. An .XLSX format of the full survey instrument is also included to support interpretation and reuse of the dataset.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dataset 2: A National Dataset on Human Choices in Pooled Rideshare, 2022

Dataset 2: A National Dataset on Human Choices in Pooled Rideshare, 2022. Dataset Description: Pooled Rideshare Acceptance Survey - Phase 2 (2022, N = 2,884). This dataset captures responses from a nationally representative sample of 2,884 adults across the United States to understand choice behaviors between personal and pooled rideshare services. The primary objective of this research is to investigate choice behaviors in rideshare services and provide insights that inform service design, policymaking, and transportation planning, with the aim of encouraging pooled rideshare adoption and enhancing transportation network energy efficiency. Data was collected via an online survey administered through a national panel provider. Participants ranged in age from 18 to 94 years, and representation from all U.S. regions. The survey was designed with a focus on investigating the stated-preference between personal and pooled rideshare services. Each participant responded to 20 stated-preference questions, where they were presented with a hypothesized situation to choose between a personal rideshare option and a pooled rideshare option to complete a trip. The sociodemographic information and attitudes towards factors of pooled rideshare acceptance were also collected to support the comprehensive investigation of participants’ rideshare choice behaviors. - Phase_2_Final - Each row represents an individual respondent, and each column corresponds to a variable such as stated-preference scenario attributes, stated-preference scenario responses, attitudes toward specific service features, and sociodemographic data. The data is available in both .CSV and .SAV formats. - Phase_2_Final_MapFile - The accompanying data dictionary explains all variable labels, response scales, and codes. An .XLSX format of the full survey instrument is included to support interpretation and reuse of the dataset.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

CDL description of the CDC 6600 stunt box

The CDC 6600 central memory control (stunt box) is described utilizing CDL (Computer Design Language), block diagrams, and text. The stunt box is a clearing house for all central memory references from the 6600 central and peripheral processors. Since memory requests can be issued simultaneously, the stunt box must be capable of assigning priorities to requests, of labeling requests so that the data will be distributed correctly, and of remembering rejected addresses due to memory conflicts.

Hertzog, J. B.↗

A survey of automated remote sensing for agriculture

The state-of-the-art of the technology available to make remote sensing crop production estimates is reviewed with reference to several past and present research projects. In particular, attention is given to Landsat data acquisition, registration and preprocessing, data transformation, data modeling, proportion estimation, and labeling. Development stage models and crop condition models are briefly characterized, and areas where further research is needed are identified.

Hall, F. G.↗

The pH of Mars

The Viking labeled release (LR) experiments provided data that can be used to determine the acid-base characteristics of the regolith. Constraints on the acid-base properties and redox potentials of the Martian surface material would provide additional information for determining what reactions are possible and defining formation conditions for the regolith. Calculations devised to determine the pH of Mars must include the amount of soluble acid species or base species present in the LR regolith sample and the solubility product of the carbonate with the limiting solubility. This analysis shows that CaCO3, either as calcite or aragonite, has the correct K(sub sp) to have produced the Viking LR successive injection reabsorption effects. Thus CaCO3 or another MeCO3 with very similar solubility characteristics must have been present on Mars. A small amount of soluble acid, but no more than 4 micro-mol per sample, could also have been present. It is concluded that the pH of the regolith is 7.2 +/- 0.1.

Plumb, R. C.↗

A Statistical Model to Predict the Extratropical Transition of Tropical Cyclones

This paper introduces a logistic regression model for the extratropical transition(ET) of tropical cyclones in the North Atlantic and the Western North Pacific, using elastic net regularization to select predictors and estimate coefficients.Predictors are chosen from the 1979-2017 best track and reanalysis datasets, and verification is done against the tropical/extratropical labels in the best track data. In an independent test set, the model skillfully predicts ET at lead times up to two days, with latitude and sea surface temperature as its most important predictors. At a lead time of 24 h, it predicts ET with a Matthews correlation coefficient of 0.4 in the North Atlantic, and 0.6 in the Western North Pacific. It identifies 80% of storms undergoing ET in the North Atlantic, and 92% of those in the Western North Pacific. 90% of transition time errors are less than 24 h. Select examples of the model's performance on individual storms illustrate its strengths and weaknesses.Two versions of the model are presented: an "operational model" that may provide baseline guidance for operational forecasts, and a "hazard model"that can be integrated into statistical TC risk models. As instantaneous diagnostics for tropical/extratropical status, both models' zero lead time predictions perform about as well as the widely used Cyclone Phase Space (CPS) in the Western North Pacific and better than the CPS in the North Atlantic, and predict the timings of the transitions better than CPS in both basins.

Melanie Bieli↗

Label-based Virtual Directories In dCache

Traditional filesystems organize data in directories. These directories are typically a collection of files whose grouping is based on a single criterion, e.g., the starting date of an experiment, experiment name, beamline ID, measurement device, or instrument. However, each file in a directory can belong to several logical groups, such as a special event type, experiment condition, or a part of a selected dataset. dCache is a storage system developed to store large amounts of scientific data, used by many HEP and Photon Science experiments. With recent developments in dCache, we have introduced a concept of file tagging, which dynamically groups files with the same label into virtual directories. The file labels can be added, removed, renamed, and deleted through the admin interface or via REST API. The files in virtual directories are exposed through all protocols supported by dCache. This contribution will describe the details of the implementation for file tagging in dCache and present our future development plans on automatic metadata extractions, a feature that will significantly simplify data management. Additionally, we are exploring the future use of virtual directories as a way to translate scientific data catalogs into filesystem views for direct data analysis.

Sahakyan, Marina [DESY]↗

Crop identification studies using Landsat data Separation of barley from other spring small grains and corn and soybean decision logic

Two labeling procedures were developed which identify various agricultural crops through the use of Landsat data. One procedure separates barley from other spring small grains, and the other identifies corn and soybeans. For both procedures, a minimum data set (critical acquisition time) has been designated. Landsat data in both image format and various graphic displays were used along with ancillary data to obtain information which aided in labeling the spectral signatures. The corn and soybean procedure also employed a structured decision logic. Test results for the barley separation procedure emphasized the importance of obtaining a critical acquisition and showed some success especially in areas where spring crops followed the expected growth patterns. Two tests of the corn and soybean procedure produced good labeling accuracies. Problems with the procedure were easy to identify, and some solutions were implemented for the second test. Automation of various parts of the procedure and extension to other crops and regions were recommended.

Dailey, C. L.↗