Search NASA⌕ Search

SEARCH · Search NASA

Results for “Labeled Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Statistical and Neural Network for Real Sensor-Data-Driven Anomaly Detection in Nuclear Applications

Anomaly detection (AD) in sensor data is critical to ensure uninterrupted functionality of nuclear power plants (NPPs). Consequently, validation of AD models through real-world sensor data is important for their application in nuclear facilities. In this paper, we propose an Autoencoder (AE)— a multi-layered neural network, for AD in sensor data from an operational NPP testbed. Since the dataset lacks labels for irregularities, we introduce random noise and label them to effectively train our model. The proposed AE model assigns a higher reconstruction error to the abnormal samples that deviate from those encountered during the training phase and uses the reconstruction loss to detect anomalies in a representative imbalanced dataset. We also introduce an analytical solution—seasonal trend decomposition (STD) — as another AD scheme for identifying irregularities within the same time-series dataset. In contrast to the AE model which relies on reconstruction loss, the STD scheme decomposes the entire dataset into its trend, seasonality, and residual components to pinpoint irregularities. Our findings indicate that the proposed AE and STD models individually achieve recall scores of 97% and 92%, respectively. We also validate the performance of the two models on both balanced and imbalanced data. We further solidify the results by picking the combined selected anomalies of the two solutions with an "AND" operator for more reliable predictions.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Using active learning to improve quasar identification for the DESI spectra processing pipeline

The Dark Energy Spectroscopic Instrument (DESI) survey uses an automatic spectral classification pipeline to classify spectra. QuasarNET is a convolutional neural network used as part of this pipeline originally trained using data from the Baryon Oscillation Spectroscopic Survey (BOSS). In this paper we implement an active learning algorithm to optimally select spectra to use for training a new version of the QuasarNET weights file using only DESI data, with the goal of improving classification accuracy. This active learning algorithm includes a novel outlier rejection step using a Self-Organizing Map to ensure we label spectra representative of the larger quasar sample observed in DESI. We perform two iterations of the active learning pipeline, assembling a final dataset of 5600 labeled spectra, a small subset of the approximately 1.3 million quasar targets in DESI's Data Release 1. When splitting the spectra into training and validation subsets we achieve similar performance to the previously trained weights file in completeness and purity calculated on the validation dataset but do so with less than one tenth of the amount of training data. The new weights also more consistently classify objects in the same way when used on unlabeled data compared to the old weights file. In the process of improving QuasarNET's classification accuracy we discovered a systemic error in QuasarNET's redshift estimation and used our findings to improve our understanding of QuasarNET's redshifts.

Machine learning↗

On the Prospect of Chemically Transferable Coarse-Grained Electronic Models for Soft Materials

Electronic coarse-graining (ECG) methods predict quantum-mechanical electronic properties directly from coarse-grained (CG) molecular configurations, enabling electronic predictions at mesoscale length scales. Here, we present a diagnostic assessment of the feasibility of chemically transferable ECG models across a broad polymer-relevant chemical space using all-atom, united-atom, and Martini-scale representations. While high-resolution ECG models achieve near-quantitative accuracy, we show that chemically transferable ECG at the Martini resolution fails because the CG force field does not sample the same configurational distribution of local molecular structure as that underlying the DFT-parameterized ECG model. We demonstrate that our proposed Element-Count-Label (ECL) representation, which augments Martini beads with explicit stoichiometric data, significantly improves chemical generalization across diverse polymer chemistries. However, we find that even with improved chemical resolution, the model cannot recover electronic property distributions that are absent from the configurational space sampled by the CG force field. These results demonstrate that chemically transferable ECG requires future Martini-like force fields to explicitly preserve quantum chemistry–compatible local molecular structure in addition to thermodynamic and structural fidelity.

Kidder, Katherine M [Department of Chemistry; Univ↗

Determining Stellar Elemental Abundances from DESI Spectra with the Data-driven Payne

Abstract Stellar abundances for a large number of stars provide key information for the study of Galactic formation history. Large spectroscopic surveys such as the Dark Energy Spectroscopic Instrument (DESI) and LAMOST take median-to-low-resolution (R≲ 5000) spectra in the full optical wavelength range for millions of stars. However, the line-blending effect in these spectra causes great challenges for elemental abundance determination. Here we employDD-Payne, a data-driven method regularized by differential spectra from stellar physical models, to the DESI early data release spectra for stellar abundance determination. Our implementation delivers 15 labels, including effective temperatureT eff , surface gravity log g , microturbulence velocityv mic , and the abundances for 12 individual elements, namely C, N, O, Mg, Al, Si, Ca, Ti, Cr, Mn, Fe, and Ni. Given a spectral signal-to-noise ratio of 100 per pixel, the internal precisions of the label estimates are about 20 K forT eff , 0.05 dex for log g , and 0.05 dex for most elemental abundances. These results agree with the theoretical limits from the Crámer–Rao bound calculation within a factor of 2. The majority of the accreted halo stars contributed by the Gaia–Enceladus–Sausage are discernible from the disk and in situ halo populations in the resultant [Mg/Fe]–[Fe/H] and [Al/Fe]–[Fe/H] abundance spaces. We also provide distance and orbital parameters for the sample stars, which spread over a distance out to ∼100 kpc. The DESI sample has a significantly higher fraction of distant (or metal-poor) stars than the other existing spectroscopic surveys, making it a powerful data set for studying the Galactic outskirts. The catalog is publicly available.

Astronomy & Astrophysics↗

MolViewSpec: a Mol* extension for describing and sharing molecular visualizations

Data visualization is a pivotal component of a structural biologist’s arsenal. The Mol* Viewer makes molecular visualizations available to broader audiences via most web browsers. While Mol* provides a wide range of functionality, it has a steep learning curve and is only available via a JavaScript interface. To enhance the accessibility and usability of web-based molecular visualization, we introduce MolViewSpec (molstar.org/mol-view-spec), a standardized approach for defining molecular visualizations that decouples the definition of complex molecular scenes from their rendering. Scene definition can include references to commonly used structural, volumetric, and annotation data formats together with a description of how the data should be visualized and paired with optional annotations specifying colors, labels, measurements, and custom 3D geometries. Developed as an open standard, this solution paves the way for broader interoperability and support across different programming languages and molecular viewers, enabling more streamlined, standardized, and reproducible visual molecular analyses. MolViewSpec is freely available as a Mol* extension and a standalone Python package.

Midlik, Adam [European Bioinformatics Institute (U↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Solar PV

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) research platform. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence data centers and other variable loads. This dataset entry describes the behavior of a 1.25-MW proton exchange membrane MC250 electrolyzer system, manufactured by Nel Hydrogen , [1] when fed historical data generated by the 430-kW, fixed-axis solar photovoltaic (PV) array located at NLR’s Flatirons Campus. (While the electrolyzer balance of plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack.) Solar PV power output data for the 2020 calendar year were categorized on a daily basis by total energy generation and standard deviation. Each day was then ranked by these metrics, and the 25th, 50th, and 100th percentiles were selected. The 75th percentile day did not exhibit sufficient variability to make for a valuable experiment. A similar process was used for the related historical wind dataset . [2] The historical days in 2020 that represented these percentiles are Dec. 19, March 29, and May 4, respectively. The entire solar day’s power profile was then fed through the MC250 electrolyzer. Due to its length, the 100th percentile day experiment was split into two parts, and the final 3 hours of the solar day were not captured. These final 3 hours contained no spikes or dips of interest and simply represented a slow decay of input solar power. Also, a single timestamp (13:13:47 on Jan. 14, 2026) was lost in the hydrogen system supervisory control and data acquisition. Finally, during the 25th percentile experiment (solar day Dec. 19, 2020) data recording was lost from 11:00:13 to 11:14:45. The roughly 15 minutes of the solar profile were rerun at the end of the experiment and spliced into this time slot during post-processing. The electrolysis system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operation of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical solar profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. For more details on the statistical analysis process, see the slide deck “Public Reference Data for Megawatt-Scale Hydrogen Electrolysis: NLR Historical Solar PV Analysis and Profile Generation” accessible with this data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single solar PV electrolysis experiment and is formatted as: {technology}_{percentile}_{scaling factor} For instance, “solarPV-430kW_25_2x.zip” reports the experiment using the 25th percentile solar data from the historical 2020 solar PV dataset, scaled to 200%. Scaling factors were applied to the generated solar PV power output files to more closely match the 1.25-MW capacity of the electrolyzer. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and solar power input. A PDF file detailing the historical solar data statistical analysis used to generate the solar profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all experiments combined into one dataset labeled "combined_solarPV_experiments.csv". [1] nelhydrogen.com/product/mc-series-electrolyser . [2] data.nlr.gov/submissions/316 .

08 HYDROGEN↗

Evaluation of Digital Nautical Chart data for confirmation and expansion of GeoNames data

Here, this work examines how Digital Nautical Chart (DNC) data may contribute to the evolution and refinement of GeoNames data for near-shore features. GeoNames features are point data with one or more possible place names. DNC Earth Cover Text (ECRText) objects are map labels positioned nearby their real word counterpart. ECRText feature map position strikes a compromise between association with real features and cartographic readability. This work explores whether ECRText features can confirm (or expand names for) existing locations or contribute new locations through data conflation. Due to name variations and spatial position, conflating these data are nontrivial. Previous work engaged in a brief examination using the trigram string matching algorithm under coarse proximity constraints, indicating that ECRText could provide additional value to GeoNames. This work builds on that study, by engaging in a deeper examination of spatial proximity and exploring conflation agreement across an ensemble of string matching approaches. The result finds strong ensemble agreement about ECRText features which already exist in GeoNames but mixed results about which features contribute new information, as well as exploring why some of these matching techniques fail. With an eye toward automation, computational efficiency was found not to be a constraint in sustaining updates.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Creating a Training Dataset for Semantic Segmentation of Canal Networks for Irrigation Modernization

Canal infrastructure has provided critical irrigation water to the western United States for over a century. To continue providing vital water resources to the semi-arid West, irrigation systems must undergo maintenance and modernization. Many canal companies are resource-constrained, and because funding opportunities often require detailed knowledge of existing infrastructure, they can struggle to secure financial capital. We address this problem by creating training data for a semantic segmentation deep learning model to map canal networks throughout the western United States. To create a diverse and robust training dataset, we labelled 1-m NAIP imagery with the locations of no canals, wet canals, and dry/vegetated canals. Since creating these datasets is time consuming, we first developed a preprocessing methodology to identify canals within our four study areas. We used NAIP imagery and provided canal centerline data to buffer, standardize, and cluster the imagery, automating the labeling process as much as possible. However, this still required manual cleaning and manual classification of canal type. Challenges arose when canals were interrupted (e.g., road culverts or piped sections) or when nearby features shared similar characteristics (e.g., irrigated fields, trees, and shadows). Combining automated preprocessing with manual refinement produced four detailed canal masks to be used in the semantic segmentation model developed by Richard Tapia.

13 - HYDRO ENERGY↗

Chlamydomonas reinhardtii responses to Fe-excess, Fe-deficiency, and Fe-limitation in either photoautotrophic or mixotrophic growth

A systems level analysis of Chlamydomonas reinhardtii grown photoautotrophically or mixotrophically with a reduced carbon source, acetate, under four different defined Fe stages of Fe-replete, Fe-deficient, Fe-limited, or Fe-excess. Samples were digested with trypsin, labeled with TMT 10-Plex, then analyzed by LC-MS/MS. Data was searched with MS-GF+ using PNNL's DMS Processing pipeline. [doi:10.25345/C5707X12X] [dataset license: CC0 1.0 Universal (CC0 1.0)]

59 BASIC BIOLOGICAL SCIENCES↗

Iron-starvation induces photosystem I antenna remodeling in green algae

Dunaliella salina and Dunaliella tertiolecta are extremophile, marine algae that can survive in very low Fe conditions. In this study, we used TMT-proteomics to compare the Fe starvation responses to the Fe replete responses. Samples were digested with trypsin, labeled with TMT 10-Plex, then analyzed by LC-MS/MS. Data was searched with MS-GF+ using PNNL's DMS Processing pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Automated 3D cytoplasm segmentation in soft X-ray tomography

Cells’ structure is key to understanding cellular function, diagnostics, and therapy development. Soft X-ray tomography (SXT) is a unique tool to image cellular structure without fixation or labeling at high spatial resolution and throughput. Fast acquisition times increase demand for accelerated image analysis, like segmentation. Currently, segmenting cellular structures is done manually and is a major bottleneck in the SXT data analysis. This paper introduces ACSeg, an automated 3D cytoplasm segmentation model. ACSeg is generated using semi-automated labels and 3D U-Net and is trained on 43 SXT tomograms of immune T cells, rapidly converging to high-accuracy segmentation, therefore reducing time and labor. Furthermore, adding only 6 SXT tomograms of other cell types diversifies the model, showing potential for optimal experimental design. ACSeg successfully segmented unseen tomograms and is published on Biomedisa, enabling high-throughput analysis of cell volume and structure of cytoplasm in diverse cell types.

59 BASIC BIOLOGICAL SCIENCES↗

Power System Feature-Based Event Classification by Means of Multiple PMU Data

Abstract—Phasor Measurement Units (PMUs) provide time synchronized measurements across the power grid, enabling data driven event detection and classification for enhanced system monitoring and situational awareness. However, variations in event duration, spatial extent, and severity, along with coincident events, pose challenges for conventional classification models that require fixed-size inputs. This paper presents a feature-based framework that aggregates diverse attributes from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, and Multilayer Perceptron. A probabilistic post-processing scheme is further introduced to enable multi-label classification in the presence of overlapping events. Experiments using real-world PMU data demonstrate that the Random Forest model achieves 95% accuracy, while the proposed post-processing method yields an additional 3% improvement.

Nematirad, Reza↗

A Novel Method to Train Classification Models for Structure Detection in In Situ Spacecraft Data

We present a method for creating spacecraft-like data which can be used to train Machine Learning (ML) models to detect and classify structures in in situ spacecraft data. First, we use the Grad-Shafranov equation to numerically solve for several magnetohydrostatic equilibria which are variations on a known analytic equilibrium. These equilibria are then used as the initial conditions for Particle-In-Cell simulations in which the structures of interest are observed and labeled. We then take one-dimensional slices through the simulations to replicate what a spacecraft collecting data from the simulation would observe. This sliced data then can be used as training data for the initial training of ML models intended for use on spacecraft data. We demonstrate the method applied to the problem of detecting small-scale plasmoids in the magnetotail, which is important for understanding complex magnetotail reconnection dynamics. The simple 1D classifier we train is able to detect more than 70% of the plasmoid points in the data set but also produces a large number of false positives. Our further work on this example problem is detailed, and further potential uses of the method are discussed.

79 ASTRONOMY AND ASTROPHYSICS↗

Graph learning for particle accelerator operations

Particle accelerators play a crucial role in scientific research, enabling the study of fundamental physics and materials science, as well as having important medical applications. This study proposes a novel graph learning approach to classify operational beamline configurations as good or bad. By considering the relationships among beamline elements, we transform data from components into a heterogeneous graph. We propose to learn from historical, unlabeled data via our self-supervised training strategy along with fine-tuning on a smaller, labeled dataset. Additionally, we extract a low-dimensional representation from each configuration that can be visualized in two dimensions. Leveraging our ability for classification, we map out regions of the low-dimensional latent space characterized by good and bad configurations, which in turn can provide valuable feedback to operators. This research demonstrates a paradigm shift in how complex, many-dimensional data from beamlines can be analyzed and leveraged for accelerator operations.

43 PARTICLE ACCELERATORS↗

MISIP: a data standard for the reuse and reproducibility of any stable isotope probing-derived nucleic acid sequence and experiment

DNA/RNA-stable isotope probing (SIP) is a powerful tool to link in situ microbial activity to sequencing data. Every SIP dataset captures distinct information about microbial community metabolism, process rates, and population dynamics, offering valuable insights for a wide range of research questions. Data reuse maximizes the information derived from the labor and resource-intensive SIP approaches. Yet, a review of publicly available SIP sequencing metadata showed that critical information necessary for reproducibility and reuse was often missing. Here, we outline the Minimum Information for any Stable Isotope Probing Sequence (MISIP) according to the Minimum Information for any (x) Sequence (MIxS) framework and include examples of MISIP reporting for common SIP experiments. Our objectives are to expand the capacity of MIxS to accommodate SIP-specific metadata and guide SIP users in metadata collection when planning and reporting an experiment. The MISIP standard requires 5 metadata fields—isotope, isotopolog, isotopolog label, labeling approach, and gradient position—and recommends several fields that represent best practices in acquiring and reporting SIP sequencing data (e.g., gradient density and nucleic acid amount). The standard is intended to be used in concert with other MIxS checklists to comprehensively describe the origin of sequence data, such as for marker genes (MISIP-MIMARKS) or metagenomes (MISIP-MIMS), in combination with metadata required by an environmental extension (e.g., soil). The adoption of the proposed data standard will improve the reuse of any sequence derived from a SIP experiment and, by extension, deepen understanding of in situ biogeochemical processes and microbial ecology.

Simpson, Abigayle↗

Bayesian SegNet for Semantic Segmentation with Improved Interpretation of Microstructural Evolution During Irradiation of Materials

Understanding the relationship between the evolution of microstructures of irradiated LiAlO2pellets and tritium diffusion, retention and release could improve predictions of tritium performance. Given expert-labeled segmented images of irradiated and unirradiated pellets, we trained Deep Convolutional Neural Networks to segment images into defect, grain, and boundary classes. Qualitative microstructural information was calculated from these segmented images to facilitate the comparison of unirradiated and irradiated pellets. We tested modifications to improve the sensitivity of the model, including incorporating meta-data into the model and utilizing uncertainty quantification. The predicted segmentation was similar to the expert-labeled segmentation for most methods of microstructural qualification, including pixel proportion, defect area, and defect density. Overall, the high performance metrics for the best models for both irradiated and unirradiated images shows that utilizing neural network models is a viable alternative to expert-labeled images.

Oostrom, Marjolein T.↗

Image masks of global ship tracks for NASA MODIS data products

Ship tracks, long thin artificial cloud features formed from the pollutants in ship exhaust, are satellite-observable examples of aerosol-cloud interactions (ACI) that can lead to increased cloud albedo and thus increased solar reflectivity, phenomena of interest in solar radiation management. In addition to ship tracks being of interest to meteorologists and policy makers, their observed cloud perturbations provide benchmark evidence of ACI that remain poorly captured by climate models. To broadly analyze the effects of ship tracks, high-resolution satellite imagery data highlighting their presence are required. To support this, we provide a hand labelled dataset to serve as a benchmark for a variety of subsequent analyses. Established from a previous dataset that identified ship track presence using NASA’s MODIS Aqua satellite imager, our first-of-its-kind dataset is comprised of image masks: capturing full ship track regions, including their contours, emission points and dispersive patterns. In total, 300 images, or around 2,500 masked ship tracks, observed under varying conditions are provided, and may facilitate training of machine learning algorithms to automate extraction.

Atmospheric dynamics↗

Detecting Living-off-the-land Attacks Using K-means And Graph Convolutional Networks

The code ingests Zeek logs derived from network packet captures and goes through data preprocessing before it gets passed into a K-Means model that labels each device as either a client or server. Graph Convolutional Network (GCN) model is used to obtain the embeddings to represent the features in lower dimension. Last, K-means cluster analysis is used to cluster the embeddings for each class.

Quach, Anna [Idaho National Laboratory (INL), Idah↗