Search NASASearch

SEARCH · Search NASA

Results for “Machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

High-level hadronic tau lepton triggers of the CMS experiment in proton-proton collisions at √(s) = 13.6 TeV

The trigger system of the CMS detector is pivotal in the acquisition of data for physics measurements and searches. Studies of final states characterized by hadronic decays of tau leptons require the reconstruction and the identification of genuine tau leptons against quark- and gluon-initiated jets at the trigger level. This is a difficult task, particularly as improvements to the LHC have resulted in an increased number of interactions per bunch crossing in recent years. To address this challenge, a series of machine-learning algorithms with high identification efficiency and low computational cost have been incorporated into the high-level trigger for hadronically decaying tau leptons. In this paper, these developments and the trigger performance are summarized using data collected by the CMS experiment in proton-proton collisions at √(s) = 13.6 TeV in 2022–2023, corresponding to an integrated luminosity of 62 fb -1 .

Particle identification methods

Deep probabilistic direction prediction in 3D with applications to directional dark matter detectors

Abstract We present the first method to probabilistically predict 3D direction in a deep neural network model. The probabilistic predictions are modeled as a heteroscedastic von Mises-Fisher distribution on the sphere S 2 , giving a simple way to quantify aleatoric uncertainty. This approach generalizes the cosine distance loss which is a special case of our loss function when the uncertainty is assumed to be uniform across samples. We develop approximations required to make the likelihood function and gradient calculations stable. The method is applied to the task of predicting the 3D directions of electrons, the most complex signal in a class of experimental particle physics detectors designed to demonstrate the particle nature of dark matter and study solar neutrinos. Using simulated Monte Carlo data, the initial direction of recoiling electrons is inferred from their tortuous trajectories, as captured by the 3D detectors. For 40 keV electrons in a 70% He 30% CO 2 gas mixture at STP, the new approach achieves a mean cosine distance of 0.104 (26 ∘ ) compared to 0.556 (64 ∘ ) achieved by a non-machine learning algorithm. We show that the model is well-calibrated and accuracy can be increased further by removing samples with high predicted uncertainty. This advancement in probabilistic 3D directional learning could increase the sensitivity of directional dark matter detectors.

Computer Science

Characterization of the astrophysical diffuse neutrino flux using starting track events in IceCube

In this article, a measurement of the diffuse astrophysical neutrino spectrum is presented using IceCube data collected from 2011-2022 (10.3 years). We developed novel detection techniques to search for events with a contained vertex and exiting track induced by muon neutrinos undergoing a charged-current interaction. Searching for these starting track events allows us to not only more effectively reject atmospheric muons but also atmospheric neutrino backgrounds in the southern sky, opening a new window to the sub-100 TeV astrophysical neutrino sky. The event selection is constructed using a dynamic starting track veto and machine learning algorithms. We use this data to measure the astrophysical diffuse flux as a single power law flux (SPL) with a best-fit spectral index of γ=2.58$_{-0.09}^{+0.10}$ and per-flavor normalization of $\phi$$_{per-flavor}^{Astro}$=1.68$_{-0.22}^{+0.19}$×10 -18 ×GeV -1 cm -2 s -1 sr -1 (at 100 TeV). The sensitive energy range for this dataset is 3-550 TeV under the SPL assumption. This data was also used to measure the flux under a broken power law, however we did not find any evidence of a low energy cutoff.

79 ASTRONOMY AND ASTROPHYSICS

Observation of 𝑡⁢𝑊⁢𝑍 Production at the CMS Experiment

The first observation of single top quark production in association with a 𝑊 and a 𝑍 boson in proton-proton collisions is reported. The analysis uses data at center-of-mass energies of 13 and 13.6 TeV recorded with the CMS detector at the CERN LHC, corresponding to a total integrated luminosity of 200 fb −1 . Events with three or four charged leptons, which can be electrons or muons, are selected. Advanced machine-learning algorithms and improved reconstruction methods, compared to an earlier analysis, result in an unprecedented sensitivity to 𝑡⁢𝑊⁢𝑍 production. The measured cross sections for 𝑡⁢𝑊⁢𝑍 production are 248 ± 52 fb and 242 ± 77 fb for $\sqrt{s}$ =13 and 13.6 TeV, respectively. The signal is established with a statistical significance of 5.8 standard deviations, with 3.5 expected, compared to the background-only hypothesis.

Hayrapetyan, Aram [Yerevan Physics Institute]

Novel CHI3L1 ‐Associated Angiogenic Phenotypes Define Glioma Microenvironments: Insights From Multi‐Omics Integration

ABSTRACT The CHI3L1 signaling pathway significantly influences glioma angiogenesis, but its role in the tumor microenvironment (TME) remains elusive. We propose a novelCHI3L1‐associated vascular phenotype classification for glioma through integrative analyses of multiple datasets with bulk and single‐cell transcriptome, genomics, digital pathology, and clinical data. We investigated the biological characteristics, genomic alterations, therapeutic vulnerabilities, and immune profiles within these phenotypes through a comprehensive multi‐omics approach. We constructed the vascular‐related risk (VR) score based onCHI3L1‐associated vascular signatures (CAVS) identified by machine learning algorithms. Utilizing unsupervised consensus clustering, gliomas were stratified into three distinct vascular phenotypes: Cluster A, marked by high vascularization and stromal activation with a relatively low levels of tumor‐infiltrating lymphocytes (TILs); Cluster B, characterized by moderate vascularization and stromal activity, coupled with a high density of TILs; and Cluster C, defined by low vascularization and sparse immune cell infiltration. We observed that the CAVS effectively indicated glioma‐associated angiogenesis and immune suppression by single‐cell RNA‐seq analysis. Moreover, the high‐VR‐score group exhibited enhanced angiogenic activity, reduced immune response, resistance to immunotherapy, and poorer clinical outcomes. The VR score independently predicted glioma prognosis and, combined with a nomogram, provided a robust clinical decision‐making tool. Potential drug prediction based on transcription factors for high‐risk patients was also performed. Our study reveals thatCHI3L1‐associated vascular phenotypes shape distinct immune landscapes in gliomas, offering insights for optimizing therapeutic strategies to improve patient outcomes.

Oncology

Sensitive detection of structural dynamics using a statistical framework for comparative crystallography

Chemical and conformational changes are crucial to protein function and its pharmacological control. X-ray crystallography can reveal these changes in atomic detail, but standard analysis methods, which refine separate datasets, often overlook differences that are subtle or arise in only a subset of molecules. Direct comparison of crystallographic datasets is, in principle, more powerful, but systematic errors (“scales”) often mask changes in the crystallographic observables (“structure factors”). Machine learning algorithms that jointly estimate scales and structure factors can address this limitation. Here, we augment this approach with multivariate, structured priors derived from crystallographic theory, implemented in the variational deep learning framework Careless. Doing so strongly improves the detection of protein dynamics, element-specific anomalous signals, and the binding of drug candidates, offering a robust approach to comparative crystallography and, potentially, to detection of protein dynamics by other structure determination methods.

Hekstra, Doeke R. [Harvard Univ., Cambridge, MA (U

Scalable multiplexed machine learning gas sensor chips for food classification

Multiplexed gas sensor arrays combined with machine learning have unlocked previously inaccessible applications for scent-based sensing. Current platforms are limited by overlapping sensing materials with similar compositions, leading to highly correlated responses, or multistep deposition processes that hinder scalability. In this work, we developed a 16-element monolithic chip with fully distinct sensing layers, enabling a truly heterogeneous array. The system consists of highly sensitive carbon nanotube field effect transistors that are functionalized through a single-step microdispensing method compatible with automated pipetting systems. The resulting chip produces characteristic signal patterns in response to object-specific scent profiles and, when combined with machine learning algorithms, can perform automated object identification. We demonstrate the classification of 16 different objects, including food spoilage and nut allergens, with a 92.6% overall prediction accuracy.

Bassil, Carla [University of California, Berkeley,

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR

PyJMAK: An Open-Source Python Toolkit for Modeling Solid-State Metallurgical Phase Transformations

Accurate prediction of metallurgical phase transformations is an essential basis for autonomous optimization and rapid part qualification. Several methods can be used to estimate the evolution of phase fractions such as JMAK kinetics-based models, phase-field models, thermodynamic models, and data-driven machine learning models. Thermodynamic and phase-field-based methodologies solve multiphysics equations requiring numerous calibration parameters and significant computational resources. As a result, the computation domain is limited to a point or on order of micron-meters. The data-driven models rely on large datasets from experiments and simulations. While the JMAK model only provides information about phase fraction evolution, it can predict this evolution in near real-time using thermal history and thermodynamic data without restriction on the domain. JMAK models have been popularly used by researchers to model phase transformations occuring during additive manufacturing or over arbitrary temperature profiles. Commercial proprietary software such as Abaqus and Ansys or closed-source in-house implementations offer the ability to model JMAK based kinetics to predict phase transformation. However, these software packages are not open-source or freely available for use and development in conjunction with manufacturing machines, sensors, and machine learning algorithms. In addition, the use of the model is restricted by a license token. In contrast, given temperature profiles at multiple points in the domain, this Python-based PyJMAK model can compute phase evolution in parallel due to its stand-alone modular, voxel-based structure, and it can be executed on high-performance computing resources without any license restrictions.

Prabhune, Bhagya [Oak Ridge National Laboratory (O

The Role of Internal Variability in Springtime Arctic Amplification from 1980 to 2022

Arctic amplification (AA) refers to the enhanced warming of the Arctic relative to the global average due to rising greenhouse gases, measured as the ratio of Arctic-mean to global-mean surface air temperature (SAT) trends. From 1980 to 2022, annual-mean AA reached 4.2 (Arctic defined as north of 70°N). Climate models simulate AA but fail to reproduce its magnitude. Sweeney et al. attributed much of this model–observation discrepancy to internal variability. AA shows seasonality and so does the discrepancy. Spring (March–May) shows the largest gap: Observed AA is 4.2, while the multimodel mean is 2.7. This raises several questions: 1) What role does internal variability play in observed spring AA? 2) How does simulated spring AA compare to observations when internal variability is removed? 3) If internal variability is significant, what mechanisms drive it? To address these, we adapted the machine learning algorithm from Sweeney et al., training on simulated multidecadal spring SAT and sea level pressure (SLP) trend maps. Our results show that internal variability enhanced spring Arctic warming by 37% and reduced global warming by 10%. Removing internal variability reconciles the spring AA discrepancy. The estimated internal contribution to Arctic spring warming is supported by an independent dynamical adjustment approach. We identify an atmospheric circulation pattern in observations associated with this internal warming. Observed internal Siberian SAT and SLP trends follow the simulated SAT–SLP relationship but lie at the distribution’s extreme, suggesting models generally underestimate internal variability unless the observed configuration reflects a rare real-world realization.

Arctic

CalderaCast Web Interface

CalderaCast may be accessed as a web-based tool at the first link in the references section of this dataset. All of the necessary datasets to run the tool are built into the simulation software running behind the web interface. These input datasets are referenced by the additional links in the references section below. Many of those datasets are taken into machine-learning algorithms by the Caldera team and heavily processed to create internal datasets, which are then relied upon by the simulation to produce individual results. These internal datasets are not accessible and are not necessary for use of the CalderaCast tool.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Data and Scripts associated with a manuscript on ecosystem responses to wildfires in the Columbia River Basin

This data package is associated with the publication “Ecosystem leaf area, gross primary production, and evapotranspiration responses to wildfire in the Columbia River Basin” submitted to Biogeosciences (Shi et al., 2024; doi: 10.22541/au.171053013.30286044/v1). In this research, data products, leaf area index (LAI), gross primary production (GPP), and evapotranspiration (ET), from the Moderate Resolution Imaging Spectroradiometer (MODIS) are used to quantify the resistance and resilience of different ecosystem types in the Columbia River Basin (CRB). A machine learning algorithm, random forest (RF), was used to examine the impacts of precipitation, vapor pressure deficit (VPD), and burn severity from Monitoring Trends in Burn Severity (MTBS) on ecosystem resilience. The data package includes the processed MODIS data products, precipitation, VPD, and burn severity in 138 fire regions in CRB and the input files for RF model training. This data package includes six folders. The MODIS products are included in three MODIS_* folders with shell scripts for data clipping and *ncl files for data processing: (1) “/MODIS_LAI_CRB”; (2) “/MODIS_GPP_CRB”; and (3) “/MODIS_ET_CRB”. All the processed data for each fire event are NetCDF formatted. The MTBS burn severity data and the shell and *ncl scripts used for data processing are in the folder named (4) “MTBS_fire”. The ERA meteorological fields and the data processing scritps are in (5) “ERA_Var_CR”. All the scripts for figure development are in the format of *ncl and in the folder (6) “paper_scripts”. See the file ending in “flmd.csv” for a list of all files contained in this data package and descriptions for each. Tabular column headers and units are described in the data dictionary file ending in “dd.csv”.

54 ENVIRONMENTAL SCIENCES

High-Fidelity Building Emulator

This dataset provides high-fidelity time series data for an emulated commercial office building sited in the Chicago, IL area during a Typical Meteorological Year (TMY). This dataset consists of air-side HVAC measurements and control inputs, and it includes normal operations as well as various implemented faults (with associated ground truth measurements) implemented on selected days. This data could be used to quantify and compare the impacts of different faults, and it could also be used as training or validation data for machine learning algorithms (e.g., reduced-order modelling, fault detection and diagnosis).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Solid-State Mixed-Potential Electrochemical Sensors for Natural Gas Leak Detection and Quality Control (Final Technical Report)

Mitigation of methane emissions are a critical factor to limiting the impact of the natural gas industry on global climate change. Throughout the period of 2020-2024, the University of New Mexico and its commercialization partner and subcontractor, SensorComm Technologies, Inc. (SCT), have worked together to develop a low-cost Artificial Intelligence (AI)-driven Internet of Things (IoT)-based multi-gas sensor platform for methane emissions detection. In the final year of the project, we extended this work to include hydrogen detection in support of a transition to a hydrogen economy where hydrogen could be transported through existing natural gas infrastructure. Mixed potential electrochemical sensors were first prototyped by ceramic additive manufacturing and then transitioned to conventional ceramic manufacturing tape casting and screen-printing technologies in preparation for mass production. Demonstrated limits of detection of 5 ppm of methane in natural gas and 1 ppm of hydrogen were measured. These limits of detection are among the lowest of solid-state electrochemical sensors that have been reported in the literature or available in the industry. Machine learning algorithms were developed to identify natural gas mixtures with > 98% accuracy level and quantify methane concentrations at 97% accuracy. The presence of hydrogen could also be identified, and its concentration quantified at these accuracy levels. These algorithms were optimized for running on portable computing hardware which enabled > 1 Hz processing rates. A portable packaged IoT system was integrated with the electrochemical sensor in collaboration with SCT. The package consists of readout electronics with < 1 mV resolution, sensor temperature control, and data transmission over cellular wireless and/or Wi-Fi networks. Field testing was performed in two rounds at Colorado State University’s Methane Emissions Technology Evaluation Center (CSU METEC). The first round of testing demonstrated successful measurements of methane from an underground natural gas leak of 20 standard liters per minute (SLPM), which agreed with previously published literature using more sophisticated and expensive analytical equipment. The second round of testing showed that an above ground leak of 2 SLPM of hydrogen could be detected at 32 ft. This project has resulted in six published peer reviewed journal articles, over ten presentations at professional conferences, and one full patent application filed in 2023. Future work on this project includes increased sensitivity, higher production yields, and applications in the hydrogen safety and flare emissions monitoring spaces.

03 NATURAL GAS

Reconstruction of the 4D beam matrix

The widely used transverse parameters characterizing particle beams are the Twiss parameters. These parameters can be measured experimentally but they do not fully characterize the beam since they do not account for possible correlations in particle distribution between two transverse coordinates. These correlations may occur due to uncompensated magnetic field at the cathode or misalignment of focusing quadrupoles in the transport beamline. We test a novel diagnostic for diagnosing full 4D beam matrix which may be used to identify such imperfections. The diagnostic is based on transporting the beam through the beamline which includes a quadrupole and a skew quadrupole magnets and measuring the resulting 2D beam distribution at the screen downstream. Such a measurement can be viewed as measuring a 2D projection of the 4D distribution. Different settings of the quads provide measurements of different slices of the phase space. The reconstruction of the original beam matrix from a number of measurements is done using machine learning algorithm, which provides a fast and reliable way of reconstruction for an arbitrary configuration of the scanning beamline. In August 2024, we set up the diagnostic beamline to perform a quadrupole scan of the beam. The setup includes a skew quadrupole, a regular quadrupole, and a screen. The images on the screen were post-processed to remove experimental artifacts and enhance contrast by eliminating background noise outside the core of the distribution=. The rms parameters of the distribution were then calculated and used as inputs for the reconstruction algorithm. This algorithm attempts to determine the initial beam matrix that produces expected images on the screen closely matching the observed images across all quadrupole settings. The algorithm found a solution in which the expected rms parameters closely align with the observations. Validation of the results is planned for FY25.

43 PARTICLE ACCELERATORS

Novel experimental probes of QCD in SIDIS and e + e - annihilation (Final Technical Report)

The research addressed with this award seeks to advance our understanding of the structure and dynamics underlying the properties of visible matter. In our current understanding the nucleons (protons and neutrons) are not fundamental but are comprised of quarks and gluons, which are collectively called partons. A static quark picture fails to explain the properties of the nucleons, such as their mass and intrinsic spin, which are thought to emerge dynamically from the quark-gluon interactions via the strong force. In this work, novel observables employing correlations of particles produced in the scattering of high energy electrons off protons at the CLAS12 experiment at Jefferson Lab were analyzed to probe quark-gluon interactions. The ultimate goal of this line of inquiry is to be able to describe the properties of protons and neutrons from first principles, similar to how studying the hydrogen atom has led to the formulation of the theory of Quantum Electrodynamics. Because quarks cannot be observed directly but only as part of more complex composite particles, a smaller, complimentary part of this work was the analysis of particle production in electron-positron annihilation at the Belle II experiment to understand the production of particles from initial quarks. This takes advantage of the fact that in e + e - annihilation the initial quark dynamics is known, unlike in the scattering of nucleons. Several new applications using Machine Learning algorithms for event tagging and reconstruction to support this program were developed as part of this award. In addition, we had a significant role in the development of the physics program for the future Electron-Ion Collider, which is a new collider to be build in the US within the next decade.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Clean Water Production in Cooling Towers

This project developed and demonstrated a novel technology that produces clean water from cooling tower recirculating water by using the natural evaporation and condensation cycle inside cooling towers. The system captures the escaping plume and converts blowdown quality water into high purity water suitable for on-site reuse such as boiler feed. The technology uses electric fields to ionize exhaust plumes, charge the entrained droplets, and direct them toward collection electrodes where they coalesce and flow downward. This allows water recovery at a low energy cost while reducing visible plume emissions. In addition, we developed a complementary software platform that improves overall cooling tower performance. The system uses wireless sensors and physics-based machine learning algorithms to optimize key parameters of the cooling process. For power generation facilities, this increases the thermal efficiency of the cooling loop and condenser, resulting in measurable cycle efficiency gains. Improvements of one percent or more can deliver significant increases in electricity production for the same fuel input.

01 COAL, LIGNITE, AND PEAT

Battery Life Prediction Using Reduced-Order Physics Models and Machine Learning (CRADA Final Report)

Phase 1 (Original CRADA, plus no-cost extension modifications #1-3, 6/1/2017 to 3/13/2021): The Australian Department of Defence (AUDoD) is performing accelerated aging tests of Li-ion batteries to benchmark their reliability and degradation characteristics. Using its previously developed battery lifetime predictive model framework, the National Laboratory of the Rockies (NLR) will develop analytical models based the AUDoD data to predict lifetime of the multiple Li-ion battery chemistries under real-world use scenarios of interest to AUDoD. The NLR model is based on physical degradation mechanisms encountered by Li-ion batteries and has been previously validated. Phase 2 (CRADA modification #4, plus no-cost extension modification #5, 2/22/2021 to 3/30/2025): Train and support Australian Department of Defence personnel to use NLR software for model-based estimation of Li-ion battery lifetime using accelerated battery aging data collected by the Australian Department of Defence. Under separate DOE funding from 2019 to 2021, NLR enhanced its battery life-prediction software using machine learning algorithms to automate portions of the model-fitting process, requiring significantly less labor and expert judgment and also adding uncertainty quantification, increasing statistical rigor. Under Phase 2, NLR will customize NLR Software and provide it to AuDoD. NLR will enhance its NLR Model to capture aging modes of AuDoD's multi-cell modules, including cell-balancing effects. NLR will develop example single-cell and multi-cell models based on one AuDoD battery aging dataset. NLR will train AuDoD personnel on NLR Software. By the conclusion of the project, NLR will have provided AuDoD the training materials, a user manual and software needed to perform their own analysis of additional and/or future battery aging datasets.

33 ADVANCED PROPULSION SYSTEMS