Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Collection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Measurement of the muon anomalous precession frequency in Runs 4, 5, and 6 of the Muon g-2 experiment at Fermilab

The Fermilab E989 Muon g − 2 experiment measures the muon’s anomalous magnetic moment to a precision of 127 parts per billion, as reported in June 2025. The value is proportional to the difference between the muon’s cyclotron frequency and the spin precession frequency in the presence of a uniform magnetic field, for muons contained within the g − 2 storage ring. Spin precession frequency is extracted from the time distribution of the muon’s decay positrons recorded by 24 electromagnetic calorimeters positioned around the inner circumference of the storage ring. The anomalous precession frequency is one of the primary experimental inputs necessary to estimate the anomalous magnetic moment, the other being the measurement of the magnetic field. This dissertation details the anomalous precession frequency extraction, including reconstruction, time-distribution fitting, and treatment of systematic uncertainties for the final three data-collection runs: Run-4, Run-5, and Run-6. This data represents a fourfold increase in statistics over the previous analysis release, halving the statistical uncertainty. The residual slow term from previous analyses is now well understood and documented in a systematic treatment. As of the writing of this dissertation, the theoretical prediction for the SM estimate of the muon’s anomalous magnetic moment is under debate, with two competing prediction methods, so a definitive comparison with theory is not available. The results submitted for experimental release use the kernel-ratio asymmetry method, contributing 115 parts per billion to the statistical uncertainty and 34 parts per billion to the systematic uncertainty. When combined with the previous analyses in earlier data runs, this thereby improves the measurement beyond the experimental goal and sets the world’s most precise measurement of the muon’s anomalous magnetic moment.

Israel, Scott Nathan [Boston U.]↗

Observation of $t\bar{t}γγ$ production at $\sqrt{s} = 13$ TeV with the ATLAS detector

This paper presents the first observation of top-quark pair production in association with two photons ($t\bar{t}γγ$). The measurement is performed in the single-lepton decay channel using proton-proton collision data collected by the ATLAS detector at the Large Hadron Collider. The data correspond to an integrated luminosity of 140 fb −1 recorded during Run 2 at a centre-of-mass energy of 13 TeV. The $t\bar{t}γγ$ production cross section, measured in a fiducial phase space based on particle-level kinematic criteria for the lepton, photons, and jets, is found to be 2.42$_{− 0.53}^{+ 0.58}$ fb , corresponding to an observed significance of 5.2 standard deviations. Additionally, the ratio of the production cross section of $t\bar{t}γγ$ to top-quark pair production in association with one photon is determined, yielding (3.30$_{− 0.65}^{+ 0.70}$) × 10 −3 .

Aad, G. (ORCID:0000000266654934)↗

PV backsheets survey protocol: A framework for geo-spatial field surveys for bulk material characterization and reliability analysis applied across 41 PV systems

As widespread adoption of photovoltaic (PV) technologies continues, understanding the lifetime of modules is paramount to the viability of the industry as an environmentally conscious alternative to traditional energy generation. Although power degradation can affect the total energy production of a module over its lifetime, module safety failures necessitate the removal of a module leading to a loss of not only the particular asset, but the earning potential of the device. Therefore, it is critical to ensure that the components that provide essential safety functions for PV module operate for their entire rated lifetime. PV backsheets provide necessary electrical insulation to the completed device and failure of this component is cause for a immediate removal of the module. Degradation of the PV module backsheet has led to module safety failures in large-scale installations, costing millions of dollars in damages and lost potential revenue. The spatio-temporal degradation of fielded PV modules is important to study in order to identify which modules within installations are experiencing the greatest exposure conditions and in turn have the highest chance of failure. This paper describes a comprehensive field survey protocol developed for monitoring PV module backsheet performance using solely non-destructive methods in commercial PV fields. The protocol establishes a field naming convention, sampling method, data handling requirements, and measurement procedures. By ensuring consistent data collection practices, the field survey protocol enables research groups to obtain data of uniform quality on backsheet performance over multiple years and locations. In this study, the developed protocol was implemented at forty-one PV sites. Eight different types of airside layer backsheet materials including poly(vinylidene fluoride) (PVDF), acrylic PVDF, poly(tetrafluoroethylene-co-hexafluoropropylene-co-vinylidene fluoride) (THV), poly(vinyl fluoride) (PVF), poly(ethylene terephthalate) (PET), fluoroethylene vinyl ether (FEVE), polyethylene naphthalate (PEN), and glass were identified using attenuated total reflection Fourier transform infrared (ATR-FTIR) spectroscopy. The field survey results show that the spatial distribution of degradation indicators are non-uniform within a particular module, individual site, and across site locations. The degradation of PV modules increased in severity for modules mounted at the edge of rows (across a field) and near the junction box (within a module). This study demonstrates the sensitivity of material performance to exposure length across different materials and climates.

14 SOLAR ENERGY↗

Performance of heavy-flavour jet identification in Lorentz-boosted topologies in proton-proton collisions at √(s) = 13 TeV

Measurements in the highly Lorentz-boosted regime provoke increased interest in probing the Higgs boson properties and in searching for particles beyond the standard model at the LHC. In the CMS Collaboration, various boosted-object tagging algorithms, designed to identify hadronic jets originating from a massive particle decaying to bb̅ or cc̅, have been developed and deployed across a range of physics analyses. This paper highlights their performance on simulated events, and summarizes novel calibration techniques using proton-proton collision data collected at √(s) = 13 TeV during the 2016–2018 LHC data-taking period. Three dedicated methods are used for the calibration in multijet events, leveraging either machine learning techniques, the presence of muons within energetic boosted jets, or the reconstruction of hadronically decaying high-energy Z bosons. The calibration results, obtained through a combination of these approaches, are presented and discussed.

Pattern recognition↗

Insights into the year-round vertical distribution of chlorophyll concentration in high-latitude Arctic Ocean: implications for primary production

Climate-induced rapid changes in the Arctic Ocean, such as decreasing sea ice extent and increasing water temperature, are altering nutrient and light availability, profoundly impacting primary producer growth. However, access to the high-latitude Arctic Ocean is limited, and satellite data are primarily available only during summer, making continuous in-situ data collection challenging. We collected year-round chlorophyll-a (Chl-a) concentration data in high-latitude regions using a mooring system and performed a comparative analysis with reanalysis data. Unlike previous satellite-based studies, which typically rely on surface measurements, we used the annual vertical distribution of Chl-a. These data were applied to the vertically generalized production model to accurately estimate annual primary production. The moored Chl-a concentration data showed that phytoplankton exhibited a typical subsurface chlorophyll maximum (SCM) layer as sea ice retreated in June. Contrary to the gradually deepening SCM distribution predicted by model-based reanalysis data, the SCM layer persisted for approximately 4 months. This indicates that light and nutrient conditions within the SCM layer remained stable, sustaining continuous phytoplankton growth. Annual primary production, reflecting this vertical distribution of Chl-a concentration, was 6.85 gC m −2 yr −1 . This exceeded satellite-based estimates by at least two-fold, highlighting the significant underestimation of primary production by satellite approaches. Estimating primary production while accounting for the vertical distribution of phytoplankton and light is essential for improving ecological models to better understand carbon cycle and food web changes in the Arctic Ocean, with important implications for climate change predictions.

Arctic Ocean↗

Comprehensive framework for assessing and optimizing existing research networks

Conservation, monitoring, and research networks, or collections of ecological research sites unified under a common mission of data collection or a research mission, are essential infrastructure for understanding large landscapes. However, most networks developed opportunistically over decades rather than through systematic design, creating potential limitations in the ability to address conservation challenges across entire regions. We developed a framework to evaluate how well an existing research network represents the environmental conditions its members study and devised an approach to rank sites of priority for strategic expansion. Our approach measures performance through environmental representativeness, geographic coverage, and adequacy for scientific inference and thus optimizes limited monitoring resources to maximize scientific impact. We demonstrated this approach with the U.S. Department of Agriculture (USDA) Forest Service Experimental Forests and Ranges Network (EFRN), a 79‐site network across the United States that grew opportunistically over a century. At the national scale, the network effectively captured high‐biomass forests important for carbon cycle research; 82% of forest biomass was in well‐represented areas. Some areas in Texas, Florida, the Rocky Mountains, and the West Coast had no relevant EFRN sites, which limits the ability to make regional inferences. A fundamental challenge for the EFRN was that sites improving regional extent coverage sometimes provided minimal national benefits, which can create conflicts between local and global priorities. Adding the highest‐ranked candidate site provided a relevant site for 17% of currently poorly represented 1‐km pixel cells nationally, but regional and national site rankings varied considerably due to nested spatial inference. This framework provides quantitative tools for strategic infrastructure decision‐making, ensures that limited monitoring resources maximize conservation impact, and can be applied broadly to address the widespread challenge of optimizing conservation and monitoring networks worldwide.

additional site↗

Five Years of Dissolved Oxygen, Temperature, Salinity, Depth, Weather Data from a Transitioning Wetland at Beaver Creek, Washington, USA

Groundwater dissolved oxygen (DO) variability in coastal system remains poorly understood despite its importance for biogeochemical cycling and ecosystem modeling. Here we investigate the temporal variability in groundwater DO and its hydro-climatic drivers across hourly to seasonal timescales in a transitioning wetland at Beaver Creek, Washington, USA. The site is transitioning from a freshwater forest to a brackish tidal wetland following removal of a barrier in 2014 that prevented tides from accessing the freshwater creek. By utilizing novel optical dissolved oxygen instrumentation (Opti O2, LLC) we obtained continuous, high-frequency (5-minute), in-situ measurements of DO from the flood-plain from June 26th, 2019 through September 30th, 2024. This 63 month dataset is comprised of groundwater dissolved oxygen, temperature, water level and salinity timeseries from the floodplain. This dataset also includes rainfall, air pressure, air temperature, and solar radiation data collected with a co-located Campbell ClimaVUE50 weather sensor. All data is contained within a single csv (2019-06-26 to 2024-09-30 Beaver Creek DO, saln, BGS, temp, weather.csv) that can easily be viewed either using software such as Excel or using any text editor.

54 ENVIRONMENTAL SCIENCES↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Neutron detector response modeling in NOvA

Neutrons can present a significant challenge for neutrino experiments in which energy reconstruction is critical. With the ability to escape detection completely and with a weak correlation between their kinetic energy and any eventual energy deposition, it is difficult to fully account for neutrons produced in neutrino interactions. This in turn leads to significant model dependence when evaluating neutron-related systematic uncertainties. The NOvA experiment is a long-baseline neutrino oscillation experiment with a high-statistics sample of antineutrino data collected by its near detector. We report an excess relative to data of simulated neutron candidates with low energy depositions when using standard Geant4 physics lists. The simulation excess is traced to an overabundance of secondary photons produced from interactions of neutrons with kinetic energy greater than \SI{20}{\mega\eV}. Improved agreement with data is obtained by applying the data-driven neutron-on-carbon \menate model for neutrons between \SI{20}{\mega\eV} and ${\sim}$\SI{100}{\mega\eV}. With \menate, the residual oversimulation is more uniform across the calorimetric neutron energy spectrum, suggesting possible overproduction of primary neutrons by the GENIE neutrino interaction generator. These results motivate the adoption of \menate-supplemented Geant4 simulation as the nominal simulation in the production of future \nova simulation.

Abubakar, S.↗

Daily operational impacts on battery degradation in heavy-duty electric drayage trucks

Battery aging is a critical factor influencing the performance, longevity, and cost of ownership of battery electric trucks (BETs). This paper presents a comprehensive evaluation of battery aging for two Li-ion battery chemistries, Nickel-Manganese-Cobalt (NMC) and Lithium-Iron-Phosphate (LFP), accounting for both cycling and calendar aging. In contrast to traditional methods that rely on simplified linear degradation models based on manufacturer-provided data, this study employs semi-empirical aging models calibrated to experimentally collected data. The models are integrated into a detailed vehicle simulation environment, enabling a comprehensive assessment of battery degradation under realistic operating conditions. A case study focusing on heavy-duty electric drayage truck operations in the Port of Savannah, GA, is presented to illustrate the impact on battery pack lifespan of: seasonal variations, daily operational activities, charging strategies, and battery storage conditions. The results illuminate the significance of the battery pack’s state of charge during stationary periods, such as overnight storage or weekend parking, on battery degradation and its potential implications for long-term vehicle viability. Additionally, the study explores how different operational and environmental factors affect battery degradation, offering critical insights into best battery charging and storage practices. Our results demonstrate that LFP outperforms NMC in terms of years of useful life; however, by utilizing charging strategies that minimize the amount of time the battery spends resting at high levels of state-of-charge, the lifespan of the battery pack that uses NMC can nonetheless be increased by more than a factor of two.

25 ENERGY STORAGE↗

Measurements of Lund subjet multiplicities in 13 TeV proton-proton collisions with the ATLAS detector

This Letter presents a differential cross-section measurement of Lund subjet multiplicities, suitable for testing current and future parton shower Monte Carlo algorithms. This measurement is made in dijet events in 140 fb -1 of $\sqrt{s}$ =13 TeV proton–proton collision data collected with the ATLAS detector at CERN's Large Hadron Collider. The data are unfolded to account for acceptance and detector-related effects, and are then compared with several Monte Carlo models and to recent resummed analytical calculations. The experimental precision achieved in the measurement allows tests of higher-order effects in QCD predictions. Most predictions fail to accurately describe the measured data, particularly at large values of jet transverse momentum accessible at the Large Hadron Collider, indicating the measurement's utility as an input to future parton shower developments and other studies probing fundamental properties of QCD and the production of hadronic final states up to the TeV-scale.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

MiniDAQ-3: Providing concurrent independent subdetector data-taking on CMS production DAQ resources

The data acquisition (DAQ) of the Compact Muon Solenoid (CMS) experiment at CERN, collects data for events accepted by the Level-1 Trigger from the different detector systems and assembles them in an event builder prior to making them available for further selection in the High Level Trigger, and finally storing the selected events for offline analysis. In addition to the central DAQ providing global acquisition functionality, several separate, so-called “MiniDAQ” setups allow operating independent data acquisition runs using an arbitrary subset of the CMS subdetectors. During Run 2 of the LHC, MiniDAQ setups were running their event builder and High Level Trigger applications on dedicated resources, separate from those used for the central DAQ. This cleanly separated MiniDAQ setups from the central DAQ system, but also meant limited throughput and a fixed number of possible MiniDAQ setups. In Run 3, MiniDAQ-3 setups share production resources with the new central DAQ system, allowing each setup to operate at the maximum Level-1 rate thanks to the reuse of the resources and network bandwidth. Configuration management tools had to be significantly extended to support the synchronization of the DAQ configurations needed for the various setups. We report on the new configuration management features and on the first year of operational experience with the new MiniDAQ-3 system.

Amoiridis, Vassileios↗

Stochastic Approximation for Multi-period Simulation Optimization with Streaming Input Data

We consider a continuous-valued simulation optimization (SO) problem, where a simulator is built to optimize an expected performance measure of a real-world system while parameters of the simulator are estimated from streaming data collected periodically from the system. At each period, a new batch of data is combined with the cumulative data and the parameters are re-estimated with higher precision. The system requires the decision variable to be selected in all periods. Therefore, it is sensible for the decision-maker to update the decision variable at each period by solving a more precise SO problem with the updated parameter estimate to reduce the performance loss with respect to the target system. We define this decision-making process as the multi-period SO problem and introduce a multi-period stochastic approximation (SA) framework that generates a sequence of solutions. Two algorithms are proposed: Re-start SA (ReSA) reinitializes the stepsize sequence in each period, whereas Warm-start SA (WaSA) carefully tunes the stepsizes, taking both fewer and shorter gradient-descent steps in later periods as parameter estimates become increasingly more precise. We show that under suitable strong convexity and regularity conditions, ReSA and WaSA achieve the best possible convergence rate in expected sub-optimality either when an unbiased or a simultaneous perturbation gradient estimator is employed, while WaSA accrues significantly lower computational cost as the number of periods increases. In addition, we present the regularized ReSA, which obviates the need to know the strong convexity constant and achieves the same convergence rate at the expense of additional computation.

Computer Science↗

Regulatory Testing of WTP HL W Glasses for Compliance with Delisting Requirements, VSL-03R3780-1, Rev. 1

(Part of the data collected for this work and discussed in this report was subject to data quality and bias issues. All the affected tests have subsequently been repeated, as directed by the Waste Treatment Plant Project, and new statistical analyses have been performed on the revised data set. A subsequently-issued report supercedes this report and describes the revised data, together with the revised composition-property models (Kot et al. 2004). The reader should refer to the new report for discussion of the revised data set and composition-property (TCLP cadmium release) models. None of the data for the spike glasses, designed for Case 1 and Case 2 COPCs testing, were affected by this issue.)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Search for dark matter produced in association with a Higgs boson decaying to a τ lepton pair in proton-proton collisions at $\sqrt{s}=13$ TeV

A search for dark matter particles produced in association with a Higgs boson decaying into a pair of τ leptons is performed using data collected in proton-proton collisions at a center-of-mass energy of 13 TeV with the CMS detector. The analysis is based on a data set corresponding to an integrated luminosity of 101 fb −1 collected in 2017–2018. No significant excess over the expected standard model background is observed. This result is interpreted within the frameworks of the 2HDM+a and baryonic Z′ benchmark simplified models. The 2HDM+a model is a type-II two-Higgs-doublet model featuring a heavy pseudoscalar with an additional light pseudoscalar. Upper limits at 95% confidence level are set on the product of the production cross section and the branching fraction for each of these two simplified models. Heavy pseudoscalar boson masses between 400 and 700 GeV are excluded for a light pseudoscalar mass of 100 GeV. For the baryonic Z′ model, a statistical combination is made with an earlier search based on a data set of 36 fb −1 collected in 2016. In this model, Z′ boson masses up to 1050 GeV are excluded for a dark matter particle mass of 1 GeV.

Dark Matter↗

Location Identifiers, Metadata, and Map for Field Measurements at the East-Taylor Watershed Community Observatory, Colorado, USA (Version 3.3)

This dataset contains identifiers, metadata, and a map of the locations where field measurements have been conducted at the East-Taylor Watershed Community Observatory located in the Upper Colorado River Basin, United States. This is version 3.3 of the dataset and replaces the prior version 3.2 (see below for details on changes between the versions). Dataset description: The East River-Taylor Watershed is the primary field site of the Watershed Function Scientific Focus Area (WFSFA) and the Rocky Mountain Biological Laboratory. Researchers from several institutions generate highly diverse hydrological, biogeochemical, climate, vegetation, geological, remote sensing, and model data at the East-Taylor Watershed in collaboration with the WFSFA. Thus, the purpose of this dataset is to maintain an inventory of the field locations and instrumentation to provide information on the field activities in the East-Taylor Watershed and coordinate data collected across different locations, researchers, and institutions. The dataset contains (1) a README file with information on the various files, (2) three csv files describing the metadata collected for each surface point location, plot and region registered with the WFSFA, (3) csv files with metadata and contact information for each surface point location registered with the WFSFA, (4) a csv file with with metadata and contact information for plots, (5) a csv file with metadata for geographic regions and sub-regions within the watershed, (6) a compiled xlsx file with all the data and metadata which can be opened in Microsoft Excel, (7) a kml map of the locations plotted in the watershed which can be opened in Google Earth, (8) a jpg image of the kml map which can be viewed in any photo viewer, and (9) a zipped file with the registration templates used by the SFA team to collect location metadata. The zipped template file contains two csv files with the blank templates (point and plot), two csv files with instructions for filling out the location templates, and one compiled xlsx file with the instructions and blank templates together. Additionally, the templates in the xlsx include drop down validation for any controlled metadata fields. Persistent location identifiers (Location_ID) are determined by the WFSFA data management team and are used to track data and samples across locations. Dataset uses: This location metadata is used to update the Watershed SFA’s publicly accessible Field Information Portal (an interactive field sampling metadata exploration tool; https://wfsfa-data.lbl.gov/watershed/), the kml map file included in this dataset, and other data management tools internal to the Watershed SFA team. Version Information: The latest version of this dataset publication is version 3.3. This version contains 167 new point locations, 1 new plot, and 2 new geographic regions. Overall, there are a total of 1439 point locations, 75 plots, and 54 geographic regions. Additionally, the kml map of locations and image now includes two boundaries (Upper Ohio Creek (UO) and Carbon Creek (CA)) outside of the East River watershed (USGS HUC-10) and accompanying stream network that represents areas of focus. Refer to methods for further details on the version history. This dataset will be updated on a periodic basis with new measurement location information. Researchers interested in having their East-Taylor Watershed measurement locations added to this list should reach out to the WFSFA data management team at wfsfa-data@googlegroups.com. Acknowledgments: Please cite this dataset if using any of the location metadata in other publications or derived products. If using the location metadata for the 2018 NEON hyperspectral campaign, additionally cite Chadwick et al. (2020). doi:10.15485/1618130. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

2018 NEON and 2025 CHESS Campaigns↗

Contrastive Machine Learning with Gamma Spectroscopy Data Augmentations for Detecting Shielded Radiological Material Transfers

Data analysis techniques can be powerful tools for rapidly analyzing data and extracting information that can be used in a latent space for categorizing observations between classes of data. Machine learning models that exploit learned data relationships can address a variety of nuclear nonproliferation challenges like the detection and tracking of shielded radiological material transfers. The high resource cost of manually labeling radiation spectra is a hindrance to the rapid analysis of data collected from persistent monitoring and to the adoption of supervised machine learning methods that require large volumes of curated training data. Instead, contrastive self-supervised learning on unlabeled spectra can enhance models that are built on limited labeled radiation datasets. This work demonstrates that contrastive machine learning is an effective technique for leveraging unlabeled data in detecting and characterizing nuclear material transfers demonstrated on radiation measurements collected at an Oak Ridge National Laboratory testbed, where sodium iodide detectors measure gamma radiation emitted by material transfers between the High Flux Isotope Reactor and the Radiochemical Engineering Development Center. Label-invariant data augmentations tailored for gamma radiation detection physics are used on unlabeled spectra to contrastively train an encoder, learning a complex, embedded state space with self-supervision. A linear classifier is then trained on a limited set of labeled data to distinguish transfer spectra between byproducts and tracked nuclear material using representations from the contrastively trained encoder. The optimized hyperparameter model achieves a balanced accuracy score of 80.30%. Any given model—that is, a trained encoder and classifier—shows preferential treatment for specific subclasses of transfer types. Regardless of the classifier complexity, a supervised classifier using contrastively trained representations achieves higher accuracy than using spectra when trained and tested on limited labeled data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗