Search NASA⌕ Search

SEARCH · Search NASA

Results for “labeled data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Augmented Reality Data Generation for Training Deep Learning Neural Network

One of the major challenges in deep learning is retrieving sufficiently large labeled training datasets, which can become expensive and time consuming to collect. A unique approach to training segmentation is to use Deep Neural Network (DNN) models with a minimal amount of initial labeled training samples. The procedure involves creating synthetic data and using image registration to calculate affine transformations to apply to the synthetic data. The method takes a small dataset and generates a highquality augmented reality synthetic dataset with strong variance while maintaining consistency with real cases. Results illustrate segmentation improvements in various target features and increased average target confidence.

Torres, Gil↗

Comparison of GOES Cloud Classification Algorithms Employing Explicit and Implicit Physics

Cloud-type classification based on multispectral satellite imagery data has been widely researched and demonstrated to be useful for distinguishing a variety of classes using a wide range of methods. The research described here is a comparison of the classifier output from two very different algorithms applied to Geostationary Operational Environmental Satellite (GOES) data over the course of one year. The first algorithm employs spectral channel thresholding and additional physically based tests. The second algorithm was developed through a supervised learning method with characteristic features of expertly labeled image samples used as training data for a 1-nearest-neighbor classification. The latter's ability to identify classes is also based in physics, but those relationships are embedded implicitly within the algorithm. A pixel-to-pixel comparison analysis was done for hourly daytime scenes within a region in the northeastern Pacific Ocean. Considerable agreement was found in this analysis, with many of the mismatches or disagreements providing insight to the strengths and limitations of each classifier. Depending upon user needs, a rule-based or other postprocessing system that combines the output from the two algorithms could provide the most reliable cloud-type classification.

EXPLICIT PHYSICS ALGORITHMS↗

Hierarchical modeling for image classification

As part of the California Integrated Remote Sensing System's (CIRSS) San Bernardino County Project, the use of data layers from a geographic information system (GIS) as an integral part of the Landsat image classification process was investigated. Through a hierarchical modeling technique, elevation, aspect, land use, vegetation, and growth management data from the project's data base were used to guide class labeling decisions in a 1976 Landsat MSS land cover classification. A similar model, incorporating 1976-1979 Landsat spectral change data in addition to other data base elements, was used in the classification of a 1979 Landsat image. The resultant Landsat products were integrated as additional layers into the data base for use in growth management, fire hazard, and hydrological modeling.

Likens, W.↗

Methods for segment wheat area estimation

The major research conducted during the three years of LACIE to solve problems associated with segment wheat area estimation is reviewed. Topics covered include proportion estimation, clustering, feature extraction, and signature extension. It would appear that LANDSAT-1 and LANDSAT-2 data do not contain enough information to discriminate between crop types perfectly all the time and, therefore, a basic problem arises when no ground truth data on crop types in the area are available. New approaches are needed to reduce labeling error. Perhaps better use of multiyear LANDSAT data, a more detailed understanding of the cropping practices in the area, better crop calendar prediction, and a better understanding of the limiting sources of error in LANDSAT data related to crop discrimination may provide the insight required to develop improved designs.

Heydorn, R. P.↗

The CanBikeCO Full Pilot: Long-Term Results and Analysis From an E-Bike Program in Colorado, USA

Personal micromobility devices like bicycles, e-bikes, and scooters are low- or zero-energy alternatives to single-occupancy vehicles. However, a lack of data has led to a dearth of data-driven research on personally owned e-bike usage. We present longitudinal findings from the CanBikeCO program, focused on e-bike adoption and use across demographics, trip characteristics, and geographies in the state of Colorado. CanBikeCO recorded travel survey data from low-income individuals provided with personal e-bikes by the Colorado Energy Office in six communities across Colorado from July 2021 to December 2022. The data were collected using a custom instance of the National Renewable Energy Laboratory OpenPATH platform, which combines passive data collection with semantic information such as trip mode and purpose labels. To our knowledge, there are no prior travel survey data on personally owned e-bikes with this range and scope. Insights from this unique dataset include: (i) work trips were 17% more likely than average trips to be taken on an e-bike, (ii) e-bikes were most often reported to replace cars (34% of e-bike trips) and other personal micromobility devices (22%), and (iii) participants favored walking for trips less than 1 mile, e-bikes for trips of 1-3 miles, and e-bikes, cars, or shared rides for trips of 3-20 miles. The data used to generate these results have been made available in the Transportation Secure Data Center. We find e-bike use is appealing across age groups and may be related to characteristics of land use, urban form, occupation, income, and car ownership. We conclude for this population that the energy demand added by e-bike use (induced demand and replacing non-motorized modes) is outweighed by the reduction in energy demand from replacement of single-occupancy vehicle trips with e-bike trips. Our findings suggest considerable potential for energy savings from personal e-bike ownership.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Trapped particle absorption by the Ring of Jupiter

The interaction of trapped radiation with the ring of Jupiter is investigated. Because it is an identical problem, the rings of Saturn and Uranus are also examined. Data from the Pioneer II encounter, deductions for some of the properties of the rings of Jupiter and Saturn. Over a dozen Jupiter magnetic field models are available in a program that integrates the adiabatic invariants to compute B and L. This program is to label our UCSD Pioneer II encounter data with the most satisfactory of these models. The expected effects of absorbing material on the trapped radiation are studied to obtain the loss rate as a function of ring properties. Analysis of the particle diffusion problem rounds out the theoretical end of the ring absorption problem. Other projects include identification of decay products for energetic particle albedo off the rings and moons of Saturn and a search for flux transfer events at the Jovian magnetopause.

Fillius, W.↗

Processing AIRS Scientific Data Through Level 2

The Atmospheric Infrared Spectrometer (AIRS) Science Processing System (SPS) is a collection of computer programs, denoted product generation executives (PGEs), for processing the readings of the AIRS suite of infrared and microwave instruments orbiting the Earth aboard NASA s Aqua spacecraft. AIRS SPS at an earlier stage of development was described in "Initial Processing of Infrared Spectral Data' (NPO-35243), NASA Tech Briefs, Vol. 28, No. 11 (November 2004), page 39. To recapitulate: Starting from level 0 (representing raw AIRS data), the PGEs and their data products are denoted by alphanumeric labels (1A, 1B, and 2) that signify the successive stages of processing. The cited prior article described processing through level 1B (the level-2 PGEs were not yet operational). The level-2 PGEs, which are now operational, receive packages of level-1B geolocated radiance data products and produce such geolocated geophysical atmospheric data products such as temperature and humidity profiles. The process of computing these geophysical data products is denoted "retrieval" and is quite complex. The main steps of the process are denoted microwave-only retrieval, cloud detection and cloud clearing, regression, full retrieval, and rapid transmittance algorithm.

Oliphant, Robert↗

Generative learning for slow manifolds and bifurcation diagrams

In dynamical systems characterized by separation of time scales, the approximation of so called “slow manifolds”, on which the long term dynamics lie, is a useful step for model reduction. Initializing on such slow manifolds is a useful step in modeling, since it circumvents fast transients, and is crucial in multiscale algorithms (like the equation-free approach) alternating between fine scale (fast) and coarser scale (slow) simulations. In a similar spirit, when one studies the infinite time dynamics of systems depending on parameters, the system attractors (e.g., its steady states) lie on bifurcation diagrams (curves for one-parameter continuation, and more generally, on manifolds in state parameter space. Sampling these manifolds gives us representative attractors (here, steady states of ODEs or PDEs) at different parameter values. Algorithms for the systematic construction of these manifolds (slow manifolds, bifurcation diagrams) are required parts of the “traditional” numerical nonlinear dynamics toolkit. In more recent years, as the field of Machine Learning develops, conditional score-based generative models (cSGMs) have been demonstrated to exhibit remarkable capabilities in generating plausible data from target distributions that are conditioned on some given label. It is tempting to exploit such generative models to produce samples of data distributions (points on a slow manifold, steady states on a bifurcation surface) conditioned on (consistent with) some quantity of interest (QoI, observable). In this work, we present a framework for using cSGMs to quickly (a) initialize on a low-dimensional (reduced-order) slow manifold of a multi-time-scale system consistent with desired value(s) of a QoI (a “label”) on the manifold, and (b) approximate steady states in a bifurcation diagram consistent with a (new, out-of-sample) parameter value. This conditional sampling can help uncover the geometry of the reduced slow-manifold and/or approximately “fill in” missing segments of steady states in a bifurcation diagram. Finally, the quantity of interest, which determines how the sampling is conditioned, is either known a priori or identified using manifold learning-based dimensionality reduction techniques applied to the training data.

Dynamical systems↗

EOS MLS Level 1B Data Processing, Version 2.2

A computer program performs level- 1B processing (the term 1B is explained below) of data from observations of the limb of the Earth by the Earth Observing System (EOS) Microwave Limb Sounder (MLS), which is an instrument aboard the Aura spacecraft. This software accepts, as input, the raw EOS MLS scientific and engineering data and the Aura spacecraft ephemeris and attitude data. Its output consists of calibrated instrument radiances and associated engineering and diagnostic data. [This software is one of several computer programs, denoted product generation executives (PGEs), for processing EOS MLS data. Starting from level 0 (representing the aforementioned raw data, the PGEs and their data products are denoted by alphanumeric labels (e.g., 1B and 2) that signify the successive stages of processing.] At the time of this reporting, this software is at version 2.2 and incorporates improvements over a prior version that make the code more robust, improve calibration, provide more diagnostic outputs, improve the interface with the Level 2 PGE, and effect a 15-percent reduction in file sizes by use of data compression.

Perun, Vincent↗

Ask-The-Expert: Minimizing Human Review for Big Data Analytics Through Active Learning

In this CIF project, we worked toward semi-automating knowledge discovery from anomaly detection algorithms through the use of active learning. Active learning is an area of research within machine learning that uses an "expert in the loop" to learn from large data sets that have very few annotations or labels available, and where providing such labels is expensive. In our case, the task can be defined as the identification of safety events from flight operational data. Since traditional anomaly detection algorithms cannot differentiate between operationally relevant and irrelevant statistical anomalies, Subject Matter Experts (SMEs) have a lengthy and expensive burden of investigating every example identified by the detection algorithm, classifying and labeling them as relevant or irrelevant. Active learningidentifies the unlabeled example for which a label would most improve the classifier, asks the domain expert for a label, and repeats this process until there are no more resources (time, budget) available for labeling or a minimum required performance is reached. A positive label indicates an operationally significant safety event whereas a negative label indicates otherwise. Based on these few labels we propose to build an active learning system that utilizes the SME's time in the most effective manner by iteratively asking for labels for as few informative instances as possible. Our work was proposed to be a stepping stone toward implementation and deployment of the system with user interface to be pursued by the Aviation Operations and Safety Program (AOSP) given its interest in safety monitoring and discovery of safety incidents.

aviation safety↗

Labeled trees and the efficient computation of derivations

The effective parallel symbolic computation of operators under composition is discussed. Examples include differential operators under composition and vector fields under the Lie bracket. Data structures consisting of formal linear combinations of rooted labeled trees are discussed. A multiplication on rooted labeled trees is defined, thereby making the set of these data structures into an associative algebra. An algebra homomorphism is defined from the original algebra of operators into this algebra of trees. An algebra homomorphism from the algebra of trees into the algebra of differential operators is then described. The cancellation which occurs when noncommuting operators are expressed in terms of commuting ones occurs naturally when the operators are represented using this data structure. This leads to an algorithm which, for operators which are derivations, speeds up the computation exponentially in the degree of the operator. It is shown that the algebra of trees leads naturally to a parallel version of the algorithm.

Grossman, Robert↗

An Open-Access Repository of Synchrophasor Data Quality Examples: Curation and Example Applications

Synchrophasor measurements are critical in providing wide-area situational awareness to power system operators. However, data artifacts may be introduced due to various issues such as loss of communication, loss of GPS signal, internal clock error, and vendor-specific implementation of phasor estimation algorithms. Tools designed to provide actionable insights from synchrophasor data, hence, must be designed to be robust to these data quality issues. In this work, two years of synchrophasor data sourced from multiple electric utilities in the United States were analyzed to identify examples of data quality problems. These examples were then labeled and published in the Grid Event Signature Library, a publicly available repository of power system measurements hosted by the Oak Ridge National Laboratory. This paper describes the data curation process, and illustrates two application use cases where the dataset can be valuable to the research community. In the first use case, a random forest classifier is trained to distinguish power system disturbance signatures from data anomalies introduced in synchrophasor measurements due to clock errors. The second use case studies the impact of data quality issues on an example synchrophasor application (specifically, event start time determination). The choice of data quality problems investigated is informed by the examples in the repository curated in this work.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Explainable AI for Multivariate Time Series Pattern Exploration: Latent Space Visual Analytics With Temporal Fusion Transformer and Variational Autoencoders in Power Grid Event Diagnosis

Detecting and analyzing complex patterns in multivariate time-series data is crucial for decision-making in urban and environmental system operations. However, challenges arise from the high dimensionality, intricate complexity, and interconnected nature of complex patterns, which hinder the understanding of their underlying physical processes. Existing AI methods often face limitations in interpretability, computational efficiency, and scalability, reducing their applicability in real-world scenarios. This paper proposes a novel visual analytics framework that integrates two generative AI models, Temporal Fusion Transformer (TFT) and Variational Autoencoders (VAEs), to reduce complex patterns into lower-dimensional latent spaces and visualize them in 2D using dimensionality reduction techniques such as PCA, t-SNE, and UMAP with DBSCAN. These visualizations, presented through coordinated and interactive views and tailored glyphs, enable intuitive exploration of complex multivariate temporal patterns, identifying patterns’ similarities and uncover their potential correlations for a better interpretability of the AI outputs. The framework is demonstrated through a case study on power grid signal data, where it identifies multi-label grid event signatures, including faults and anomalies with diverse root causes. Additionally, novel metrics and visualizations are introduced to validate the models and assess the performance, efficiency, and consistency of latent maps generated by VAE, which have been utilized in prior studies for latent space cartography and used as a benchmark in this study, and the emerging TFT architecture under various configurations. These analyses provide actionable insights for model parameter tuning and reliability improvements. Comparative results highlight that TFT achieves shorter run times and superior scalability to diverse time-series data shapes compared to VAE. This work advances fault diagnosis in multivariate time series, fostering explainable AI to support critical system operations.

Explainable AI↗

Steady-State Cycle Deck Launcher Developed for Numerical Propulsion System Simulation

One of the objectives of NASA's High Performance Computing and Communications Program's (HPCCP) Numerical Propulsion System Simulation (NPSS) is to reduce the time and cost of generating aerothermal numerical representations of engines, called customer decks. These customer decks, which are delivered to airframe companies by various U.S. engine companies, numerically characterize an engine's performance as defined by the particular U.S. airframe manufacturer. Until recently, all numerical models were provided with a Fortran-compatible interface in compliance with the Society of Automotive Engineers (SAE) document AS681F, and data communication was performed via a standard, labeled common structure in compliance with AS681F. Recently, the SAE committee began to develop a new standard: AS681G. AS681G addresses multiple language requirements for customer decks along with alternative data communication techniques. Along with the SAE committee, the NPSS Steady-State Cycle Deck project team developed a standard Application Program Interface (API) supported by a graphical user interface. This work will result in Aerospace Recommended Practice 4868 (ARP4868). The Steady-State Cycle Deck work was validated against the Energy Efficient Engine customer deck, which is publicly available. The Energy Efficient Engine wrapper was used not only to validate ARP4868 but also to demonstrate how to wrap an existing customer deck. The graphical user interface for the Steady-State Cycle Deck facilitates the use of the new standard and makes it easier to design and analyze a customer deck. This software was developed following I. Jacobson's Object-Oriented Design methodology and is implemented in C++. The AS681G standard will establish a common generic interface for U.S. engine companies and airframe manufacturers. This will lead to more accurate cycle models, quicker model generation, and faster validation leading to specifications. The standard will facilitate cooperative work between industry and NASA. The NPSS Steady-State Cycle Deck team released a batch version of the Steady-State Cycle Deck in March 1996. Version 1.1 was released in June 1996. During fiscal 1997, NPSS accepted enhancements and modifications to the Steady-State Cycle Deck launcher. Consistent with NPSS' commercialization plan, these modifications will be done by a third party that can provide long-term software support.

VanDrei, Donald E.↗

Onboard Classifiers for Science Event Detection on a Remote Sensing Spacecraft

Typically, data collected by a spacecraft is downlinked to Earth and pre-processed before any analysis is performed. We have developed classifiers that can be used onboard a spacecraft to identify high priority data for downlink to Earth, providing a method for maximizing the use of a potentially bandwidth limited downlink channel. Onboard analysis can also enable rapid reaction to dynamic events, such as flooding, volcanic eruptions or sea ice break-up. Four classifiers were developed to identify cryosphere events using hyperspectral images. These classifiers include a manually constructed classifier, a Support Vector Machine (SVM), a Decision Tree and a classifier derived by searching over combinations of thresholded band ratios. Each of the classifiers was designed to run in the computationally constrained operating environment of the spacecraft. A set of scenes was hand-labeled to provide training and testing data. Performance results on the test data indicate that the SVM and manual classifiers outperformed the Decision Tree and band-ratio classifiers with the SVM yielding slightly better classifications than the manual classifier.

classification↗

Heating effects on jack pine pyrogenic organic matter properties from a pyrocosm study in 2022

This dataset contains data associated with the preprint “Fire removes preexisting pyrogenic organic matter from the ecosystem through the mechanisms of both direct combustion and increasing mineralizability” (Luo et al., 2025b), which is the complementary study to the published paper “Reburning pyrogenic organic matter: a laboratory method for dosing dynamic heat fluxes from above” (Luo et al., 2025a). We designed a full-factorial experiment with different burial depths of jack pine (Pinus banksiana Lamb) pyrogenic organic matter (PyOM) (Surface, 1 cm, and 5 cm) and different heat-flux profiles (High, Low, and Control) to examine how subsequent fires affect the properties of preexisting PyOM. We measured total carbon (C), pH, dissolved organic carbon (DOC), dissolved inorganic carbon (DIC), and mineralized C (as CO₂-C, from a 12-week incubation).We found that high heat flux and/or surface placement resulted in substantial direct C losses through combustion. Intermediate heat exposure produced both combustion losses and increases in DOC and mineralizability, which may have complex long-term implications: an increased dissolved fraction of PyOM may promote downward transport into mineral soils and potentially contribute to deeper, longer-term C storage, but it may also make PyOM more susceptible to microbial decomposition. Under the lowest heat flux and deepest burial, most PyOM was retained, and changes in DOC and C mineralization were minimal. Finally, PyOM pH, an important chemical property, decreased under low-temperature heating but increased under higher temperatures.We uploaded pH data for all samples (“pH_of_all_samples.csv”); pH and temperature-related data (peak temperature and degree hours) for samples in High and Low heat-flux treatments (“pH_vs_peakT_and_degree_hours_only_for_heated_samples.csv”); total C data (“CN_pct_C_stock_C_loss_in_samples.csv”); DOC and DIC data (“doc_dic.csv”); and mineralized C (CO₂-C) data (“CO2-C_all_original.csv”). Additional details can be found in the Methods & Sampling section.All datasets uploaded to ESS-DIVE are clearly labeled, cleaned, and include both raw and derived data, ready for reuse in other analyses. All analysis code and raw datasets are also available on GitHub: https://github.com/MengmengLuo/Fire-removes-preexisting-pyrogenic-organic-matter-from-the-ecosystem.

54 ENVIRONMENTAL SCIENCES↗

Fluorescent Applications to Crystallization

By covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and tests with model proteins have shown that labeling u to 5 percent of the protein molecules does not affect the X-ray data quality obtained . The presence of the trace fluorescent label gives a number of advantages. Since the label is covalently attached to the protein molecules, it "tracks" the protein s response to the crystallization conditions. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a darker background. Non-protein structures, such as salt crystals, do not show up under fluorescent illumination. Crystals have the highest protein concentration and are readily observed against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. Preliminary tests, using model proteins, indicates that we can use high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that more rapid amorphous precipitation kinetics may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Experiments are now being carried out to test this approach using a wider range, of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.↗

An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry

Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy.

Wang, Tianle [Brookhaven National Laboratory (BNL)↗