Search NASASearch

SEARCH · Search NASA

Results for “decision trees”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie

Biomass for Carbon Removal and Storage (BiCRS) Counterfactual Decision Tree

Counterfactual is the term used to describe a "business-as-usual" scenario which used as a baseline to compare against a new project, allowing the calculation of net impacts for a life cycle analysis (LCA). The choice of counterfactual is critical for determining the results from LCA and must be carefully justified to ensure a fair and accurate comparison. Using forest residues as an example, this decision tree illustrates decision points to be considered for sustainable biomass sourcing and provides a framework for estimating the carbon emissions or storage under the "business-as-usual” scenarios for biomass otherwise destined for use in Biomass for Carbon Removal and Storage (BiCRS) projects.

09 BIOMASS FUELS

Using Boosted Decision Trees to Select High Quality Measurements in the Mu2e Experiment at Fermilab

This thesis presents the implementation and evaluation of a Boosted Decision Tree (BDT) model to improve the selection of high-quality track measurements in the Mu2e experiment at Fermilab. The Mu2e experiment is a high-energy physics experiments seeking to observe a rare theoretical physics process known as Charged Lepton Flavor Violation. A significant challenge faced by the Mu2e experiment are so-called background events, which are events whose data mimics that of the rare physics process the experiment seeks to observe. Without a mechanism to reduce background, it would be impossible to know whether Charged Lepton Flavor Violation occurred or not. To this end, high-quality track measurements must be distinguished from low-quality track measurements. A track can be conceived of as the reconstructed path of a particle that traveled through the Mu2e detector. In addition to other data, data about such tracks is stored using a C++-based framework, specific to the domain of high-energy physics, known as ROOT. A boosted decision tree model was trained using ROOT’s Toolkit For Multivariate Analysis by leveraging variables ancillary to track quality. In evaluation, the BDT achieves a ROC-AUC of 0.927 in discriminating good-quality tracks from poor-quality tracks. Such a score is indicative of both strong discrimination and strong generalization. Subsequently, it is shown that applying a BDT-based quality cut to the distribution of particle momenta significantly enhances the signal-to-background distinction for signal electrons, paving the way for improved sensitivity to Charged Lepton Flavor Violation.

Mullany, Brendan T. [Drew U.] (ORCID:0009000818888

Decision Tree for Variable Selection vs. Impact on Durability for Biomass and Biochar Burial Pathways [Slides]

Quantifying durability for lower-TRL BiCRS pathways has been challenging as limited data are available from real-world projects and long-term experiments, resulting in an overall lack of scientific consensus. We develop a decision tree that aims to summarize the current scientific understanding and state-of-the-art project experience. The decision tree can be used to (1) guide the selection of key variables and evaluate their relative impact on durability, (2) identify data and knowledge gaps for future research.

09 BIOMASS FUELS

Hyperplane decision trees as piecewise linear surrogate models for chemical process design

Recent trends in chemical engineering research point towards an increasing reliance on data-driven modeling approaches. Neural networks, for instance, have proven to be accurate when data is plentiful and high-dimensional, but in many cases, they require computationally-intensive training procedures. Here, in this work, we describe hyperplane decision trees (HT) as a highly expressive and low-compute machine learning model architecture. These models are locally linear and have linear decision boundaries, resulting in a piecewise linear model of the data. This property allows them to be converted into mixed-integer linear constraints which can be globally optimized. Our open-source PyTorch implementation of this method is a fast, flexible, and accessible way to build accurate piecewise linear models of data.

Decision trees

Boosted decision tree reweighting of simulated neutrino interactions for O ( 1 ) GeV neutrino cross-section measurements

This paper illustrates a generic method for multidimensional reweighting of O ( 1 ) GeV neutrino interaction Monte Carlo samples. The reweighting is based on a boosted decision tree algorithm trained on high-dimensional space in detector final-state observables. This enables one generator’s events to be reweighted so that its reconstructed particle content and kinematics distributions, as well as detector efficiency, match those of a target model. The approach establishes an efficient way to reuse legacy Monte Carlo data, avoiding regeneration. As an example, we test its use in a measurement of transverse kinematic imbalance of the μ - and proton in charged-current quasielastic like ν μ events from the MINERvA experiment.

Lin, Z. [Rochester U.] (ORCID:0009000188903698)

Decision-tree structures utilizing a phase-transition material

The rich internal physics due to competing electronic phases present in phase-transition materials such as VO2 offer the potential for compact building block design for emerging non-von Neumann computing technologies. Here, based on the relaxation dynamics of an insulator-metal phase transition, we demonstrate experimentally a decision-tree classifier embedded within a single volatile resistive switching device. The tree is constructed by the combination of the voltage pulse and relaxation time and can adapt to different tasks. We use machine learning to analyze the relaxation process, enabling a predictive voltage-relaxation time phase diagram for the electrical resistance state. Classification of the etiology of the chronic cough is presented as a proof-of-principle use case. Further, our approach can be generalized to broader classes of solid-state and solid-liquid interfacial systems that demonstrate a variety of phase relaxations.

36 MATERIALS SCIENCE

Dataset for Top Model Decision Tree: Selecting Segmentation Models for Reliable Quantitative Analysis in Low- and Ultralow-Dose CryoEM

Motivation Multiple deep learning model architectures can be used to segment bacterial membranes in cryoEM images. However, an AI-based tool advancement is often presented with only a single segmentation model for broad use, and this single model may show inconsistent results across datasets from different users. Here, we present the Top Model Decision Tree, a model screening framework to screen for the best model to generate bacterial inner and outer membrane masks based on user priorities. We use pre-trained segmentation models from YOLOv11, YOLO26, U-Net, Detectron2 and SAM3 fine-tuned on bacterial inner and outer membranes imaged with cryoEM. Run the Framework This notebook must be opened in Google Colab. Mount Google Drive and run with a GPU-based runtime. Open the notebook and follow steps to git clone in folders and files within this repository. There will be a repeating top_model_decision_tree.ipynb (notebook clone) that will not be used. Save your .png binary mask files and .csv table outputs within your Google Drive or download before closing the notebook. The models and all analysis/training scripts are available at [GitHub: https://github.com/Lynnicia/CryoEM_membranes_top_model_decision_tree and https://github.com/Sireesiru/Semantic-Segmentation-of-bacterial-cell-envelope-using-U-Nets.

59 BASIC BIOLOGICAL SCIENCES

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions

A procedure for rule extraction from a Self-Organising plasma disruption predictor for JET

In a previous paper, a Self-Organizing Map had proven to be able to identify the regions of the plasma operative space characterizing the pre-disruptive phase at JET without relying on any a priori information. One of the strengths of this disruption predictor lies in its inherent self-organization capability. The Self-Organizing Map discovers non-trivial relationships and captures the complicated interplay of device diagnostics on the internal plasma states directly from the experimental data. Moreover, the provided model allows the visualization of high-dimensional plasma parameters and facilitates easy interrogation of the model to understand the reasons behind its correlations. In this paper, an additional step is taken towards the interpretability of models for predicting disruptions by training a Decision Tree to classify the plasma states according to the interpretation provided by the Self-Organizing Map (stable or at high risk of disruptions). The Decision tree provides a set of rules which describe the transition of the plasma towards the pre-disruptive phase as visualized in the Self-Organizing Map. The obtained rules for the database explored in the study identify four regions in the map, two of which are at risk of disruption. These regions correspond to partitions of a 3D space based on the peaking factors of the core and divertor radiation, as well as the Locked Mode. The agreement between the Self-Organizing Map answers and the rules supplied by the Decision Tree is confirmed by the comparison of the performance exhibited by the two models in the prediction of disruptions.

Setzu, Samuele [Univ. of Cagliari, Monserrato, Cag

Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI

A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.

Harwood, Thomas V. [USDOE Joint Genome Institute (

Multi-parametric analysis for mixed integer linear programming: An application to transmission upgrade and congestion management

Upgrading the capacity of existing transmission lines is essential for meeting the growing energy demands, facilitating the integration of renewable energy, and ensuring the security of the transmission system. This study focuses on the selection of lines whose capacities and by how much should be expanded from the perspective of the Independent System Operators (ISOs) to minimize the total system cost. We employ advanced multi-parametric programming and an enhanced branch-and-bound algorithm to address complex mixed-integer linear programming (MILP) problems, considering multi-period time constraints and physical limitations of generators and transmission lines. To characterize the various decisions in transmission expansion, we model the increased capacity of existing lines as parameters within a specified range. This study first relaxes the binary variables to continuous variables and applies the Lagrange method and Karush-Kuhn-Tucker (KKT) conditions to obtain optimal solutions and identify critical regions associated with active and inactive constraints. Moreover, we extend the traditional branch-and-bound (B&B) method by determining the problem’s upper and lower bounds at each node of the B&B decision tree, helping to manage computational challenges in large-scale MILP problems. Here, we compare the difference between the upper and lower bounds to obtain an approximate optimal solution within the decision-makers’ tolerable error range. In addition, the first derivative of the objective function on the parameters of each line is used to inform the selection of lines for easing congestion and maximizing social welfare. Finally, the capacity upgrades are selected by weighing the reductions in system costs against the expense of upgrading line capacities. The findings are supported by numerical simulations and provide transmission-line planners with decision-making guidance.

24 POWER TRANSMISSION AND DISTRIBUTION

Search for low-mass hidden-valley dark showers with non-prompt muon pairs in proton-proton collisions at $\sqrt{s}=13$ TeV

A search for signatures of a dark analog to quantum chromodynamics is performed. The analysis targets long-lived dark mesons that decay into standard-model particles, with a high branching fraction of the dark mesons decaying into muons. The dark mesons are formed by the hadronisation of dark partons, which are produced by a decay of the Higgs boson. The search is performed using a data set corresponding to an integrated luminosity of 41.6 fb −1 , which was collected in proton-proton collisions at $\sqrt{s}=13$ TeV by the CMS experiment at the CERN LHC in 2018 using non-prompt muon triggers. The search is based on resonant muon pair signatures. Machine-learning techniques are employed in the analysis, utilising boosted decision trees to discriminate between signal and background. No significant excess is observed above the standard model expectation. Upper limits on the branching fraction of the Higgs boson decaying to dark partons are determined to be as low as 10−4 at 95% confidence level, surpassing and extending the existing limits on models with dark $\tilde{ω}$ mesons for mean proper decay lengths of less than 500 mm and for $\tilde{ω}$ masses down to 0.3 GeV. First limits are set for extended dark-shower models with two dark flavours that contain dark photons, probing their masses down to 0.33 GeV.

Beyond Standard Model

Searches for direct slepton production in the compressed-mass corridor in $\sqrt{\textrm{s}}$ = 13 TeV pp collisions with the ATLAS detector

This paper presents searches for the direct pair production of charged light-flavour sleptons, each decaying into a stable neutralino and an associated Standard Model lepton. The analyses focus on the challenging ``corridor'' region, where the mass difference, $Δm$, between the slepton ($\tilde{e}$ or $\tildeμ$) and the lightest neutralino ($\tildeχ^{0}_{1}$) is less or similar to the mass of the $W$ boson, $m(W)$, with the aim to close a persistent gap in sensitivity to models with $Δm \lesssim m(W)$. Events are required to contain a high-energy jet, significant missing transverse momentum, and two same-flavour opposite-sign leptons ($e$ or $μ$). The analysis uses $pp$ collision data at $\sqrt{s} = 13$ TeV recorded by the ATLAS detector, corresponding to an integrated luminosity of 140 fb$^{-1}$. Several kinematic selections are applied, including a set of boosted decision trees. These are each optimised for different $Δm$ to provide expected sensitivity for the first time across the full $Δm$ corridor. The results are generally consistent with the Standard Model, with the most significant deviations observed with a local significance of 2.0 $σ$ in the selectron search, and 2.4 $σ$ in the smuon search. While these deviations weaken the observed exclusion reach in some parts of the signal parameter space, the previously present sensitivity gap to this corridor is largely reduced. Constraints at the 95% confidence level are set on simplified models of selectron and smuon pair production, where selectrons (smuons) with masses up to 300 (350) GeV can be excluded for $Δm$ between 2 GeV and 100 GeV.

hadron-hadron scattering

Search for supersymmetry using vector boson fusion signatures and missing transverse momentum in pp collisions at $\sqrt{s}$ = 13 TeV with the ATLAS detector

This paper presents a search for supersymmetric particles in models with highly compressed mass spectra, in events consistent with being produced through vector boson fusion. The search uses 140 fb −1 of proton-proton collision data at $\sqrt{s}$ = 13 TeV collected by the ATLAS experiment at the Large Hadron Collider. Events containing at least two jets with a large gap in pseudorapidity, large missing transverse momentum, and no reconstructed leptons are selected. A boosted decision tree is used to separate events consistent with the production of supersymmetric particles from those due to Standard Model backgrounds. The data are found to be consistent with Standard Model predictions. The results are interpreted using simplified models of R-parity-conserving supersymmetry in which the lightest supersymmetric partner is a bino-like neutralino with a mass similar to that of the lightest chargino and second-to-lightest neutralino, both of which are wino-like. Lower limits at 95% confidence level on the masses of next-to-lightest supersymmetric partners in this simplified model are established between 117 and 120 GeV when the lightest supersymmetric partners are within 1 GeV in mass.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS