Search NASA⌕ Search

SEARCH · Search NASA

Results for “labeled data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Sharing the Sun Community Solar Project Data

This database represents a list of community solar projects, complete and pending, identified through various sources. The dataset is updated multiple times per year. The current version is the first file located below. Previous versions of the dataset published before June of 2024 can be found in the dataset below labeled “ARCHIVE_Sharing the Sun Community Solar Project Data_Before 06.24.“ The list has been reviewed but errors may exist, and the list may not be comprehensive. Errors in the sources e.g. press releases may be duplicated in the list. Blank spaces represent missing information. NLR invites input to improve the database including, to correct erroneous information, add missing projects, fill in missing information, and remove inactive projects. Updated information can be submitted to Sudha Kannan ( sudha.kannan@nlr.gov ).

14 SOLAR ENERGY↗

Telemetry-Enhancing Scripts

Scripts Providing a Cool Kit of Telemetry Enhancing Tools (SPACKLE) is a set of software tools that fill gaps in capabilities of other software used in processing downlinked data in the Mars Exploration Rovers (MER) flight and test-bed operations. SPACKLE tools have helped to accelerate the automatic processing and interpretation of MER mission data, enabling non-experts to understand and/or use MER query and data product command simulation software tools more effectively. SPACKLE has greatly accelerated some operations and provides new capabilities. The tools of SPACKLE are written, variously, in Perl or the C or C++ language. They perform a variety of search and shortcut functions that include the following: Generating text-only, Event Report-annotated, and Web-enhanced views of command sequences; Labeling integer enumerations with their symbolic meanings in text messages and engineering channels; Systematic detecting of corruption within data products; Generating text-only displays of data-product catalogs including downlink status; Validating and labeling of commands related to data products; Performing of convenient searches of detailed engineering data spanning multiple Martian solar days; Generating tables of initial conditions pertaining to engineering, health, and accountability data; Simplified construction and simulation of command sequences; and Fast time format conversions and sorting.

Maimone, Mark W.↗

Description of Data Archiving Associated With the NASA Goddard Grant NAGS-9590

Data restoration and archiving activities for this project have resulted in the restoration of 100% of the original Mariner 9 raw data set as well as many of the secondary analysis data sets. These data sets have been submitted to the Planetary Data System (PDS) Atmospheric Node, along with their PDS labels and descriptive metadata. In addition, a useful visualization and analysis tool has also been developed which allows the user to compare these Mariner 1971 Ultraviolet spectral data with several choices of related data sets: Mariner 9 images, USGS geologic data, MGS MOLA topography, Viking images (Viking MDIM) and thermal inertia data (MGS TES).

Simmons, K. E.↗

Machine Learning Prototype App For Recognition of Fruits

As the incidence of obesity and associated negative health consequences is rising, it becomes crucial to monitor the dietary choices of individuals. Unfortunately, traditional methods to collect this information involve collecting food frequency questionnaires from individuals using paper. Electronic food trackers have been developed to collect food data, but they require participants to manually label and describe the content of their meals, and which may be difficult for researchers to interpret in a standardized fashion. Machine learning, however, provides an easy and efficient method for both participants and researchers to label food items with standardized descriptions. This project aims to create a prototype phone application that can identify and label photos of apples. This is done by making a machine learning model through Turicreate, a python module, which is then implemented into an iOS app through Xcode and Swift. The modules used in Swift include CoreML and AVFoundation. This machine learning application will be incorporated with a MealLogger phone app that is also under development. The MealLogger app will be used to keep track of participants' calorie intake and other personal details throughout the sleep study. The machine learning model will present several potential identities of the foods found in the photo, and the user will only need to select the correct option. This will be a user-friendly method for participants to easily log their food consumption without the hard work of manually inputting each and every description. Some limitations to this project include the wide variety of food, including those within different cultures. To deal with this, the model will include the most generic food categories, which the participant may select, and produce a drop-down menu of more specific dishes under that specified category, with the option of self-input. Additional questionnaires may be implemented according to the food type selected This will allow the process to be quick and easy, but also specific for the purpose of analysis. The release of the application will require a much longer process, but the machine learning prototype presents a first step toward an application that may change data analysis for researchers interested in collecting food intake from individuals living in the real world.

Food tracker↗

A photostationary state analysis of the NO2-NO system based on airborne observations from the subtropical/tropical North and South Atlantic

The Chemical Instrumentation Test and Evaluation 3 (CITE 3) NO-NO2 database has provided a unique opportunity to examine important aspects of tropospheric photochemistry as related to the rapid cycling between NO and NO2. Our results suggest that when quantitative testing of this photochemical system is based on airborne field data, extra precautions may need to be taken in the analysis. This was particularly true in the CITE 3 data analysis where different regional environments produced quite different results when evaluating the photochemical test ratio (NO2)(sub expt)/(NO2)(sub calc), designated here as R(sub E)/R(sub C). The quantity (NO2)(sub Calc) was evaluated using the following photostationary state expression: (NO2)(sub Calc) = k(sub 1)(O3) + k(sub 4)(HO2) + k(sub 5)(CH3O2) + k(sub 6)(RO2))(NO)(sub Expt)/J(sub 2). The four most prominent regional environmental data sets identified in this analysis were those labeled here as free-tropospheric northern hemisphere (FTNH), free-tropospheric tropical northern hemisphere (FTTNH), free-tropospheric southern hemisphere (FTSH), and tropical-marine boundary layer (plume) (TMBL(P)). The respective R(sub E)/R(sub C) mean and median values for these four data subsets were 1.74, 1.69; 3.00, 2.79; 1.01, 0.97; and 0.99, 0.94. Of the four data subsets listed, the two that were statistically the most robust were FTNH and FTSH; for these the respective R(sub E)/R(sub C) mean and standard deviation of the mean values were 1.74 +/- 0.07 and 1.01 +/- 0.04. The FTSH observations were in good agreement with theory, whereas those from the FTNH data set were in significant disagreement. An examination of the critical photochemical parameters O3, UV(zenith), NO, NO2, and non-methane hydrocarbons (NMHCs) for these two databases indicated that the most likely source of the R(sub E)/R(sub C) bias in the FTNH results was the presence of a systematic error in the observational data rather than a shortening in our understanding of fundamental photochemical processes. Although neither a chemical nor meteorological analyses of these data identified this error with complete certainty, they did point to the three most likely possibilities: (1) an NO2 interference from a yet unidentified NO(y) species: (2) the presence of unmeasured hydrocarbons, the integrated reactivity of which would be equivalent to approximately 2.7 parts per billion by volume (ppbv) of toluene; or (3) some combination of points (1) and (2). Details concerning hypotheses (1) and (2) as well as possible ways to minimize these problems in future airborne missions are discussed.

Davis, D. D.↗

LCLS RF Station Phase Anomaly Candidate Dataset

A public anomaly detection dataset constructed from RF station faults for phase at SLAC's LCLS (Linac Coherent Light Source). We have compiled a dataset of the RF station diagnostic phase data and the beam-position monitor (BPM) signals, alongside the hand labels, for a labeled study period. The dataset consists of two HDF5 files (one for train and one for test) containing the raw data, two CSV files containing information about the candidates. The CSV file for the test dataset also contains the label.

Liang, Jia [Stanford Univ., CA (United States). In↗

Evaluation of the procedure for separating barley from other spring small grains

The success of the Transition Year procedure to separate and label barley and the other small grains was assessed. It was decided that developers of the procedure would carry out the exercise in order to prevent compounding procedural problems with implementation problems. The evaluation proceeded by labeling the sping small grains first. The accuracy of this labeling was, on the average, somewhat better than that in the Transition Year operations. Other departures from the original procedure included a regionalization of the labeling process, the use of trend analysis, and the removal of time constraints from the actual processing. Segment selection, ground truth derivation, and data available for each segment in the analysis are discussed. Labeling accuracy is examined for North Dakota, South Dakota, Minnesota, and Montana as well as for the entire four-state area. Errors are characterized.

Magness, E. R.↗

Accuracy assessment in the Large Area Crop Inventory Experiment

The Accuracy Assessment System (AAS) of the Large Area Crop Inventory Experiment (LACIE) was responsible for determining the accuracy and reliability of LACIE estimates of wheat production, area, and yield, made at regular intervals throughout the crop season, and for investigating the various LACIE error sources, quantifying these errors, and relating them to their causes. Some results of using the AAS during the three years of LACIE are reviewed. As the program culminated, AAS was able not only to meet the goal of obtaining accurate statistical estimates of sampling and classification accuracy, but also the goal of evaluating component labeling errors. Furthermore, the ground-truth data processing matured from collecting data for one crop (small grains) to collecting, quality-checking, and archiving data for all crops in a LACIE small segment.

Houston, A. G.↗

Time-locked time-histories - A new way of examining eye-movement data

A problem often encountered when eye-movement measurement is conducted is the choice of 'indices' or 'statistics' available to present such information. The present study reports the use of the Time-Locked Time-History as a technique of value in the examination and presentation of eye-movement data. Plots created using this technique are labeled time-locked time-histories as they illustrate subject eye lookpoint during a period of time before and after a certain time-locking event. Events that occur with some degree of repetition, such as the onset or termination of control activities, warning signals, or changes in indicator positions may be utilized as time-locking events. The present study reports the use of this technique in an eye-movement study using a secondary task in which the subject must discriminate specific types of information in the display.

Comstock, J. Raymond, Jr.↗

Effects of perfluorohexane vapor on relative blood flow distribution in an animal model of surfactant-depleted lung injury

OBJECTIVE: To test the hypothesis that treatment with vaporized perfluorocarbon affects the relative pulmonary blood flow distribution in an animal model of surfactant-depleted acute lung injury. DESIGN: Prospective, randomized, controlled trial. SETTING: A university research laboratory. SUBJECTS: Fourteen New Zealand White rabbits (weighing 3.0-4.5 kg). INTERVENTIONS: The animals were ventilated with an FIO(2) of 1.0 before induction of acute lung injury. Acute lung injury was induced by repeated saline lung lavages. Eight rabbits were randomized to 60 mins of treatment with an inspiratory perfluorohexane vapor concentration of 0.2 in oxygen. To compensate for the reduced FIO(2) during perfluorohexane treatment, FIO(2) was reduced to 0.8 in control animals. Change in relative pulmonary blood flow distribution was assessed by using fluorescent-labeled microspheres. MEASUREMENTS AND MAIN RESULTS: Microsphere data showed a redistribution of relative pulmonary blood flow attributable to depletion of surfactant. Relative pulmonary blood flow shifted from areas that were initially high-flow to areas that were initially low-flow. During the study period, relative pulmonary blood flow of high-flow areas decreased further in the control group, whereas it increased in the treatment group. This difference was statistically significant between the groups (p =.02) as well as in the treatment group compared with the initial injury (p =.03). Shunt increased in both groups over time (control group, 30% +/- 10% to 63% +/- 20%; treatment group, 37% +/- 20% to 49% +/- 23%), but the changes compared with injury were significantly less in the treatment group (p =.03). CONCLUSION: Short treatment with perfluorohexane vapor partially reversed the shift of relative pulmonary blood flow from high-flow to low-flow areas attributable to surfactant depletion.

NASA Discipline Cardiopulmonary↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Sim-to-real supervised domain adaptation for radioisotope identification

Machine learning has the potential to improve the speed and reliability of radioisotope identification using gamma spectroscopy. However, meticulously labeling an experimental dataset for training is often prohibitively expensive, while training models purely on synthetic data is risky due to the domain gap between simulated and experimental measurements. In this research, we demonstrate that supervised domain adaptation can substantially improve the performance of radioisotope identification models by transferring knowledge between synthetic and experimental data domains. We consider two domain adaptation scenarios: (1) a simulation-to-simulation adaptation, where we perform multi-label proportion estimation using simulated high-purity germanium detectors, and (2) a simulation-to-experimental adaptation, where we perform multi-class, single-label classification using measured spectra from handheld lanthanum bromide (LaBr) and sodium iodide (NaI) detectors. We begin by pretraining a spectral classifier on synthetic data using a custom transformer-based neural network. After subsequent fine-tuning on just 64 labeled experimental spectra, we achieve a test accuracy of 96% in the sim-to-real scenario with a LaBr detector, far surpassing a synthetic-only baseline model (75%) and a model trained from scratch (80%) on the same 64 spectra. Furthermore, we demonstrate that domain-adapted models learn more human-interpretable features than experiment-only baseline models. Overall, our results highlight the potential for supervised domain adaptation techniques to bridge the sim-to-real gap in radioisotope identification, enabling the development of accurate and explainable classifiers even in real-world scenarios where access to experimental data is limited.

Lalor, Peter W.↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗

Semiempirical Estimate of Aircraft Wing Weight

Computational method estimates weight of aircraft wings from theoretical relationships and empirical data. Permits comparison of alternative materials, methods of construction, and design philosophies. Method used to make tradeoffs in preliminary design phases on basis of simple input data and for more accurate calculations in later phases when more data are available.

York, P.↗

The diageotropica mutant of tomato lacks high specific activity auxin binding sites

Tomato plants homozygous for the diageotropica (dgt) mutation exhibit morphological and physiological abnormalities which suggest that they are unable to respond to the plant growth hormone auxin (indole-3-acetic acid). The photoaffinity auxin analog [3H]5N3-IAA specifically labels a polypeptide doublet of 40 and 42 kilodaltons in membrane preparations from stems of the parental variety, VFN8, but not from stems of plants containing the dgt mutation. In roots of the mutant plants, however, labeling is indistinguishable from that in VFN8. These data suggest that the two polypeptides are part of a physiologically important auxin receptor system, which is altered in a tissue-specific manner in the mutant.

NASA Discipline Number 29-20↗

Data integration in multi-sensor based robotic workstations

A relaxation labeling algorithm is developed. The major advantage of this algorithm over the existing ones is that the mathematic operation is simplified. The simplification eases the analysis of the convergence properties. Both the theoretical and application aspects of the proposed algorithm are investigated. The local convergence properties of a labeling process with n labels and m labels are established. The investigation of the interaction among the nodes in a multinode labeling process reveals some insight into the mathematical issues involved in the relaxation operations.

Chen, Qin↗

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES↗