Search NASA⌕ Search

SEARCH · Search NASA

Results for “Supervised Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Autoencoder-Based Anomaly Detection System for Online Data Quality Monitoring of the CMS Electromagnetic Calorimeter

The CMS detector is a general-purpose apparatus that detects high-energy collisions produced at the LHC. Online data quality monitoring of the CMS electromagnetic calorimeter is a vital operational tool that allows detector experts to quickly identify, localize, and diagnose a broad range of detector issues that could affect the quality of physics data. A real-time autoencoder-based anomaly detection system using semi-supervised machine learning is presented enabling the detection of anomalies in the CMS electromagnetic calorimeter data. A novel method is introduced which maximizes the anomaly detection performance by exploiting the time-dependent evolution of anomalies as well as spatial variations in the detector response. The autoencoder-based system is able to efficiently detect anomalies, while maintaining a very low false discovery rate. The performance of the system is validated with anomalies found in 2018 and 2022 LHC collision data. In addition, the first results from deploying the autoencoder-based system in the CMS online data quality monitoring workflow during the beginning of Run 3 of the LHC are presented, showing its ability to detect issues missed by the existing system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Hydropower potential derived from streamflow extremes for Alaska, USA

Alaska is an expansive region known for its abundant natural resources, including thousands of miles of streams and rivers. These rivers represent potential opportunities for future hydropower development that could provide reliable energy supply for local communities. There is limited long-term high temporal resolution streamflow data available for the region, making data-driven estimates of potential hydropower and its variability across the state challenging. This study provides a novel data-driven approach for hydropower capacity estimation across Alaska. We use supervised machine learning to develop a relationship between the daily and peak flow duration curves in order to augment the size of our dataset from 44 sites to 67 sites. We perform a stochastic hydropower estimation across the 67 sites and identify approximately 1000 MW of total potential hydropower capacity distributed across these sites. Our study provides the first step towards more comprehensive hydropower estimation for this critical region, highlighting the need for future work integrating high-resolution spatial data, community needs, and economic constraints in estimates of potential hydropower development in Alaska.

Hydropower↗

deadtrees.earth — An open-access and interactive database for centimeter-scale aerial imagery to uncover global tree mortality dynamics

Excessive tree mortality is a global concern and remains poorly understood as it is a complex phenomenon. We lack global and temporally continuous coverage on tree mortality data. Ground-based observations on tree mortality, e.g., derived from national inventories, are very sparse, and may not be standardized or spatially explicit. Earth observation data, combined with supervised machine learning, offer a promising approach to map overstory tree mortality in a consistent manner over space and time. However, global-scale machine learning requires broad training data covering a wide range of environmental settings and forest types. Low altitude observation platforms (e.g., drones or airplanes) provide a cost-effective source of training data by capturing high-resolution orthophotos of overstory tree mortality events at centimeter-scale resolution. Here, we introduce deadtrees.earth, an open-access platform hosting more than two thousand centimeter-resolution orthophotos, covering more than 1,000,000 ha, of which more than 58,000 ha are manually annotated with live/dead tree classifications. This community-sourced and rigorously curated dataset can serve as a comprehensive reference dataset to uncover tree mortality patterns from local to global scales using space-based Earth observation data and machine learning models. This will provide the basis to attribute tree mortality patterns to environmental changes or project tree mortality dynamics to the future. The open nature of deadtrees.earth, together with its curation of high-quality, spatially representative, and ecologically diverse data will continuously increase our capacity to uncover and understand tree mortality dynamics.

Citizen science↗

Neural Network‐Based Methods for Ocean Surface Wave Measurement Using Submarine Distributed Acoustic Sensing (DAS)

Two new data-driven models for estimating ocean surface waves from distributed acoustic sensing (DAS) submarine cable strain rate are developed using supervised machine learning on a 10-day data set collected offshore of Oliktok Point, Alaska. The new models were trained on target data from seafloor pressure moorings at three sites spaced evenly along 27.1 km of cable and were benchmarked against an empirical transfer function method previously used to estimate waves from DAS. A model which uses convolutional neural networks to transform 2-km frequency-wavenumber strain spectra to seafloor pressure spectra outperforms the benchmark in wave height prediction (RMSE of 0.15 vs. 0.41 m) and period prediction (0.29 vs. 0.37 s) when evaluated on a held-out test data set. When applied to a DAS data set collected on the same cable 2 years prior, the CNN-based model maintained similar significant wave height performance (RMSE = 0.23 m) relative to available satellite altimetry data. A two-hidden-layer, fully connected neural network which transforms 1-D strain spectra to seafloor pressure spectra also outperforms the benchmark in wave height prediction (RMSE of 0.19 vs. 0.41 m), but does not generalize as well to the prior data. Regression-based machine learning is useful for estimating waves from DAS data when the pressure-strain relationship varies temporally and spatially across different wave conditions. Models can be applied to DAS data to measure waves with higher spatial resolution and longer temporal coverage than traditional methods, which often measure waves only at a single point.

Davis, Jacob R. [Univ. of Washington, Seattle, WA ↗

Label-free nanoscopy of cell metabolism by ultrasensitive reweighted visible stimulated Raman scattering

Nanoscopic imaging of cell metabolism is hindered by the incompatibility of small metabolites with fluorescent dyes and the limited resolution of imaging mass spectrometry. We present ultrasensitive reweighted visible stimulated Raman scattering (URV-SRS), a label-free vibrational nanoscopy technique that enables multiplexed detection of metabolic nanostructures within cells. We developed an extensively chirped spectral focusing visible SRS microscope that achieves a detection limit of 4,000 molecules and introduced a self-supervised learning-based denoiser to robustly suppress non-independent SRS noise by over 7.2 dB. The instrumentation-based signal enhancement and computation-based noise suppression synergistically improved the detection sensitivity by 50 times over near-infrared SRS. Leveraging this enhanced sensitivity, we further pushed the resolution to nanoscopic levels by introducing Fourier reweighting to amplify sub-100 nm spatial frequencies previously overwhelmed by noise. Validated by Fourier ring correlation, URV-SRS achieves a lateral resolution of 86 nm in cellular imaging. Here, we applied URV-SRS to elucidate the reprogramming of metabolic nanostructures associated with virus replication in Vero E6 host cells and to compositionally delineate subcellular fatty acid synthesis in engineered Escherichia coli, demonstrating its capability towards nanoscopic spatial metabolomics.

59 BASIC BIOLOGICAL SCIENCES↗

Waveform retrieval for ultrafast applications based on convolutional neural networks

Electric field waveforms of light carry rich information about dynamical events on a broad range of timescales. The insight that can be reached from their analysis, however, depends on the accuracy of retrieval from noisy data. In this article, we present a novel approach for waveform retrieval based on supervised deep learning. We demonstrate the performance of our model by comparison with conventional denoising approaches, including wavelet transform and Wiener filtering. The model leverages the enhanced precision obtained from the nonlinearity of deep learning. The results open a path toward an improved understanding of physical and chemical phenomena in field-resolved spectroscopy.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Extended Fayans energy density functional: optimization and analysis

The Fayans energy density functional (EDF) has been very successful in describing global nuclear properties (binding energies, charge radii, and especially differences of radii) within nuclear density functional theory. In a recent study, supervised machine learning methods were used to calibrate the Fayans EDF. Building on this experience, in this work we explore the effect of adding isovector pairing terms, which are responsible for different proton and neutron pairing fields, by comparing a 13D model without the isovector pairing term against the extended 14D model. At the heart of the calibration is a carefully selected heterogeneous dataset of experimental observables representing ground-state properties of spherical even–even nuclei. To quantify the impact of the calibration dataset on model parameters and the importance of the new terms, we carry out advanced sensitivity and correlation analysis on both models. The extension to 14D improves the overall quality of the model by about 30%. The enhanced degrees of freedom of the 14D model reduce correlations between model parameters and enhance sensitivity.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Morphological Model to Separate Resolved–Unresolved Sources in the DESI Legacy Surveys: Application in the LS4 Alert Stream

Separating resolved and unresolved sources in large imaging surveys is a fundamental step to enable downstream science, such as searching for extragalactic transients in wide-field time-domain surveys. Here we present our method to effectively separate point sources from the resolved, extended sources in the Dark Energy Spectroscopic Instrument (DESI) Legacy Surveys (LS). We develop a supervised machine learning model based on the Gradient Boosting algorithm XGBoost. The features input to the model are purely morphological and are derived from the tabulated LS data products. We train the model using ∼2 × 10 5 LS sources in the COSMOS field with HST morphological labels and evaluate the model performance on LS sources with spectroscopic classification from the DESI Data Release 1 (∼2 × 10 7 objects) and the Sloan Digital Sky Survey Data Release 17 (∼3 × 10 6 objects), as well as on ∼2 × 10 8 Gaia stars. A significant fraction of LS sources are not observed in every LS filter, and we therefore build a “Hybrid” model as a linear combination of two XGBoost models, each containing features combining aperture flux measurements from the “blue” (gr) and “red” (iz) filters. The Hybrid model shows a reasonable balance between sensitivity and robustness, and achieves higher accuracy and flexibility compared to the LS morphological typing. With the Hybrid model, we provide classification scores for ∼3 × 10 9 LS sources, making this the largest ever machine learning catalog separating resolved and unresolved sources. The catalog has been incorporated into the real-time pipeline of the La Silla Schmidt Southern Survey (LS4), enabling the identification of extragalactic transients within the LS4 alert stream.

astrostatistics↗

Optimal transport for 𝑒/𝜋 0 particle classification in LArTPC neutrino experiments

The efficient classification of electromagnetic activity from 𝜋 0 and electrons remains an open problem in the reconstruction of neutrino interactions in liquid argon time projection chamber (LArTPC) detectors. We address this problem using the mathematical framework of optimal transport (OT), which has been successfully employed for event classification in other high energy physics contexts and is ideally suited to the high-resolution calorimetry of LArTPCs. Using a publicly available simulated dataset from the MicroBooNE Collaboration, we show that OT methods achieve state-of-the-art reconstruction performance in 𝑒/𝜋 0 classification. The success of this first application indicates the broader promise of OT methods for LArTPC-based neutrino experiments.

Neutrino detection↗

VoroClust: Scalable Clustering for Remote Sensing

Although supervised machine learning provides a powerful framework for image classification and segmentation, it requires comprehensive consistent datasets, which are not available for many remote-sensing applications. Remote-sensing datasets are expensive to collect, and each is acquired under different environmental conditions or with significant variations in system operating parameters. Unsupervised clustering algorithms analyze the structure of each dataset independently, rather than drawing on similarities with existing “training” examples, and are thus well suited for practical remote-sensing applications. We introduce VoroClust, a fast density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. VoroClust runs as fast as distance-based clustering methods, while capturing complex regional geometries at least as well as current-density-based methods. It uses a data-centered sphere cover to reduce computational demands, while still capturing data topology. It then propagates clusters outward from local peaks in density. We show that VoroClust provides fast state-of-the-art clustering for both high-resolution polarimetric synthetic aperture radar and high-dimensional hyperspectral imaging datasets.

42 ENGINEERING↗

Community‐Level Metabolic Shifts Following Land Use Change in the Amazon Rainforest Identified by a Supervised Machine Leaning Approach

ABSTRACT The Amazon rainforest has been subjected to high rates of deforestation, mostly for pasturelands, over the last few decades. This change in plant cover is known to alter the soil microbiome and the functions it mediates, but the genomic changes underlying this response are still unresolved. In this study, we used a combination of deep shotgun metagenomics complemented by a supervised machine learning approach to compare the metabolic strategies of tropical soil microbial communities in pristine forests and long‐term established pastures in the Amazon. Machine learning‐derived metagenome analysis indicated that microbial community structures (bacteria, archaea and viruses) and the composition of protein‐coding genes were distinct in each plant cover type environment. Forest and pasture soils had different genomic diversities for the above three taxonomic groups, characterised by their protein‐coding genes. These differences in metagenome profiles in soils under forests and pastures suggest that metabolic strategies related to carbohydrate and energy metabolisms were altered at community level. Changes were also consistent with known modifications to the C and N cycles caused by long‐term shifts in aboveground vegetation and were also associated with several soil physicochemical properties known to change with land use, such as the C/N ratio, soil temperature and exchangeable acidity. In addition, our analysis reveals that these alterations in land use can also result in changes to the composition and diversity of the soil DNA virome. Collectively, our study indicates that soil microbial communities shift their overall metabolic strategies, driven by genomic alterations observed in pristine forests and long‐term established pastures with implications for the C and N cycles.

carbon and nitrogen cycles↗

RGM: Random Geological Model Generation Package

This Fortran code is to accompany a manuscript to be submitted to Computers & Geosciences, a high-impact, peer-reviewed journal in computer methods for geosciences research. This Fortran code focuses on generation of synthetic geological models using a multi-randomization strategy. Generating high-fidelity synthetic geological models, including realistic seismic reflector migration images, faults, salt bodies, and relative geological time images, is the key for many supervised machine learning methods that aim to delineate faults and other geological properties of interest from seismic migration images. Our package contains two major functionalities: generating 2D synthetic random geological models and generating 3D synthetic random geological models. In each step of the generation process, we set random values for key properties of a geological model to improve the fidelity of the resulting geological model. The package also includes example codes on how to use the random geological model generation subroutines. We name this package RGM – Random Geological Model generation package.

Gao, Kai↗

NuclPred v1

This tool takes a genome assembly as input and predicts per-site nucleosome occupancy as output. Trained on physical maps of nucleosome binding preferences across the fungal kingdom, NuclPred can be applied broadly across fungi (and other eukaryotes). This breadth, combined with its accuracy, means it could have both basic and applied biological implications, for example in understanding eukaryotic gene regulation and genetic engineering. Almost universally across eukaryotes, nucleosomes - each wrapping ~150 base pairs of DNA - serve to package DNA inside the nucleus, with major consequences on DNA access, gene activity and DNA integration. NuclPred was generated using a supervised deep learning approach combining convolutional and recurrent neural networks to take DNA features (nucleotides, GC content and structural information) as input, then use that information to predict the physical attractiveness DNA sequences might have for forming nucleosomes. With this information at hand, researchers can design more efficient CRISPR constructs, explore the interplay between DNA signatures and other regulators impact nucleosome locations, predict expression patterns, etc. This tool will be published as part of a manuscript currently under revision at iScience (draft attached).

Mondo, Stephen↗

ldrd_virus_work

This is a Python code base that takes openly-available genetic information on known viruses and performs supervised machine learning and feature importance analysis on the relationship of the viral genomes to the competence to infect humans or bind to a specific host cell receptor.

Reddy, Tyler [LANL]↗

Spectral Data Fusion From Handheld Laser-Induced Breakdown Spectroscopy (LIBS) and X-ray Fluorescence (XRF) Analyzers for Improved Detection of Cerium in a Simulated Dispersal Accident

Here, this work implements a mid-level data fusion methodology on spectral data from handheld X-ray fluorescence and laser-induced breakdown spectroscopy analyzers to quantify plutonium surrogate (CeO 2 ) contamination in soil samples for the first time. Spectral data from each analyzer were used independently to train supervised machine learning regressions to predict Ce concentration. Fused features from both data sets were then used to train the same models, comparing prediction performance by evaluating model precision and sensitivity. Fusing principal component scores from the two sensors yielded an order of magnitude improvement in precision and sensitivity of predictions made with an artificial neural network, compared to predictions made by models trained on independent sensor data. As a result, a boosted ensemble trained on the fused spectral features yielded an ideal predictor with root-mean-squared error on the order of 10 –6 and calculated limit of detection order 10 –5 wt %.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Real-time neutron multiplicity and source localization for criticality safety during fuel debris removal

Advancing neutron detection and analysis techniques for complex radiation environments is an ongoing focus in nuclear instrumentation and monitoring. This proposal presents research and development of a generalized real-time neutron monitoring and analysis system, applicable to any detector capable of producing time-tagged neutron count data. While the work is demonstrated using the Neutron Multiplication Analysis Detector (NoMAD), a modular 15-tube helium-3 (He-3) array, due to its availability, spatial resolution, and flexible deployment, the methods developed are extensible to other systems, including organic scintillators and fast digital detectors. This research investigates two complementary analytical techniques for real-time characterization of neutron emitting sources: neutron multiplicity estimation based on the Hage-Cifarelli formalism and spatial localization using supervised machine learning applied to spatial count rate patterns. These methods are designed to operate under dynamic, evolving conditions such as fuel debris retrieval or reactor startup, where neutron-emitting material geometries may be partially unknown or changing over time. By integrating statistical neutron emission data with spatial localization, this research aims to develop and evaluate methods for real time neutron monitoring, source characterization, and material verification. Key contributions include implementation of a low-latency data pipeline for continuous neutron multiplicity analysis, development and validation of machine learning models for spatial inference, and experimental evaluation of system performance under variable measurement conditions. The outcomes are intended to support applications in nuclear safeguards, verification, emergency response, and reactor startup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Subcellular Feature-Based Classification of α and β Cells Using Soft X-ray Tomography

The dysfunction of α and β cells in pancreatic islets can lead to diabetes. Many questions remain on the subcellular organization of islet cells during the progression of disease. Existing three-dimensional cellular mapping approaches face challenges such as time-intensive sample sectioning and subjective cellular identification. To address these challenges, we have developed a subcellular feature-based classification approach, which allows us to identify α and β cells and quantify their subcellular structural characteristics using soft X-ray tomography (SXT). We observed significant differences in whole-cell morphological and organelle statistics between the two cell types. Additionally, we characterize subtle biophysical differences between individual insulin and glucagon vesicles by analyzing vesicle size and molecular density distributions, which were not previously possible using other methods. These sub-vesicular parameters enable us to predict cell types systematically using supervised machine learning. We also visualize distinct vesicle and cell subtypes using Uniform Manifold Approximation and Projection (UMAP) embeddings, which provides us with an innovative approach to explore structural heterogeneity in islet cells. This methodology presents an innovative approach for tracking biologically meaningful heterogeneity in cells that can be applied to any cellular system.

3D cell mapping↗

Neural network training by integration of adjoint systems of equations forward in time

A method and apparatus for supervised neural learning of time dependent trajectories exploits the concepts of adjoint operators to enable computation of the gradient of an objective functional with respect to the various parameters of the network architecture in a highly efficient manner. Specifically, it combines the advantage of dramatic reductions in computational complexity inherent in adjoint methods with the ability to solve two adjoint systems of equations together forward in time. Not only is a large amount of computation and storage saved, but the handling of real-time applications becomes also possible. The invention has been applied it to two examples of representative complexity which have recently been analyzed in the open literature and demonstrated that a circular trajectory can be learned in approximately 200 iterations compared to the 12000 reported in the literature. A figure eight trajectory was achieved in under 500 iterations compared to 20000 previously required. The trajectories computed using our new method are much closer to the target trajectories than was reported in previous studies.

Toomarian, Nikzad↗