Search NASASearch

SEARCH · Search NASA

Results for “DECISION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Typology of Decision-Making Tasks for Visualization

Despite decision-making being a vital goal of data visualization, little work has been done to differentiate decision-making tasks within the field. While visualization task taxonomies and typologies exist, they often focus on more granular analytical tasks that are too low-level to describe large complex decisions, which can make it difficult to reason about and design decision-support tools. In this paper, we contribute a typology of decision-making tasks that were iteratively refined from a list of design goals distilled from a literature review. Our typology is concise and consists of only three tasks: CHOOSE, ACTIVATE, and CREATE. Although decision types originating in other disciplines exist, we provide definitions for these tasks that are suitable for the visualization community. Our proposed typology offers two benefits. First, the ability to compose and hierarchically organize the tasks enables flexible and clear descriptions of decisions with varying levels of complexities. Second, the typology encourages productive discourse between visualization designers and domain experts by abstracting the intricacies of data, thereby promoting clarity and rigorous analysis of decision-making processes. We demonstrate the benefits of our typology through four case studies, and present an evaluation of the typology from semi-structured interviews with experienced members of the visualization community who have contributed to developing or publishing decision support systems for domain experts. Our interviewees used our typology to delineate the decision-making processes supported by their systems, demonstrating its descriptive capacity and effectiveness. Finally, we present preliminary findings on the usefulness of our typology for visualization design.

97 MATHEMATICS AND COMPUTING

First-of-a-Kind Risk-Informed Digital Twin for Operational Decision Making

A digital twin (DT) is a digital model or a collection of models of a physical entity. DTs in the nuclear arena can be used from plant design through decommissioning. Decisions are typically a priori or made offline. Risk-informed decision making is identifying what can go wrong, its frequency, and the consequences of its failure. Ideally risk-informed decision making reflects the current state of the plant and provides a decision in real time. Traditionally, probabilistic risk assessments (PRAs) evaluate the failures of safety systems, the risk of core damage, and the offsite dose as the consequence. However, this DT evaluates the decisions on the control side rather than the protection side. It uses the same risk methods to probabilistically inform the decision-making process but in a different way. Rather than evaluating the risk of core damage, this DT evaluates the likelihood of avoiding a trip set point while maintaining plant safety. Performance-based assessments are identified via its probabilistic evaluation of operational alternatives based on system status. Because the purpose of the control system is to maintain system variables within prescribed operating ranges, upsets or challenges that can exceed a trip set point resulting in a plant transient and a challenge to plant mitigating systems based on actual plant conditions, are evaluated to safely maintain the plant within the operating ranges. The probabilistic portion of the model is autonomously and automatically adjusted, and the metric of interest (i.e. likelihood of avoiding a trip set point) is recalculated. The digital representation of the physical system (i.e. the DT) performs a deterministic performance–based assessment of the probabilistically identified alternatives identified to validate the probabilistic assessment. A decision-making algorithm selects the appropriate option based on the probabilistic and deterministic assessments and transmits a control signal to a component(s) to initiate a corrective action or informs an operator of its decision.

digital twin

Reimagining How Flood Warnings Can Inform Decision‐Making and Community Actions

Society faces increasingly severe flood hazards, intensifying demand for flood early warning systems (FEWS) that deliver accurate and actionable information. However, most existing FEWS remain prediction‐centric, treating decision‐making as a downstream consumer of hazard forecasts while offering limited support for uncertainty interpretation, risk communication, and real‐world response. This Perspective presents a vision and blueprint for a novel inland FEWS‐decision‐making (FEWS‐DM) framework that repositions decision‐making as an equal partner in the forecasting process—not a passive recipient of its outputs. The framework is built on three tightly coupled, co‐evolving thrusts: Physical Science (T1), which advances flood prediction with quantified uncertainty informed by decision relevance; Human Science (T2), which incorporates psychology, behavior, and cultural and institutional context; and Decision Science (T3), which unifies physical predictions and human factors through principled, utility‐based decision support with end‐to‐end uncertainty management. Rather than treating T1 as a solved problem, FEWS‐DM recognizes that forecast development itself must be shaped by decision needs through continuous bidirectional feedback. We identify key scientific, behavioral, and operational challenges limiting such integration and discuss the enabling role of AI, while emphasizing human‐centered design and community feedback as essential for building trust and improving flood risk management.

54 ENVIRONMENTAL SCIENCES

A Tool to Incorporate Non-Energy Impacts in Energy Efficiency Investment Decision Making for Firms

Energy efficiency is a key demand-side strategy for sustainability, recently identified by the United States Department of Energy as a pillar of industrial decarbonization. The increased focus on decarbonization and the requirement for efficiency to enable electrification, another decarbonization pillar, due to the spark spread between natural gas and electricity prices, make energy efficiency increasingly relevant. Still, industries face challenges in adopting energy efficiency measures. Researchers have long found a gap in adoption of even those measures with a profitable net present value, attributed to lack of strategic value among other barriers (see for rigorous exploration and taxonomy). One solution to facilitate energy efficiency projects is the inclusion of non-energy impacts, as this has been shown to double potential deployment of such projects at system level. Energy efficiency can provide valuable benefits outside of simple operating cost reductions, from decreased pollution to enhanced productivity. The inclusion of these benefits in decision making assessments faces hurdles due to inconsistency of ancillary benefits across projects, difficulties in quantifying impacts and the need for additional measurement to quantify them. The decision-making tools to support this have been designed primarily for the European context. We begin with a stakeholder engagement process to better characterize the U.S. decision making process surrounding adoption of energy efficiency investments. Characterization of non-energy impacts has developed substantially over recent decades. Cagno et al. provided a framework for studying the applicability of these impacts to energy efficiency projects, listing 120 key performance indicators focused mainly on reductions of costs/harms. Other researchers have included impacts on the strategic and revenue side that can be merged into this framework as well. We seek a tractable set of impacts that can be included in a decision-making tool in the US, and as such are well suited to US industry, management and decision making processes. We also seek to understand how to best quantify or characterize these impacts. This work will demonstrate the results of a survey conducted among US manufacturing industry decision makers to assess the decision making landscape of stakeholders as well as the most relevant performance indicators for energy efficiency projects.

ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATION,

Multi‐Model Ensembles in Ecosystem Modeling: Challenges and Best Practices for Decision‐Making

Ecosystem models are increasingly central to the decision-making for environmental policy, conservation planning, and climate-related investments. Yet, the growing reliance on Multi-Model Ensembles (MMEs) of ecosystem models by practitioners and policymakers, sometimes under tight timelines and imperfect information, has frequently outpaced the scientific rigor required to ensure ensemble reliability. Here, MMEs refer to approaches that combine targeted predictions from multiple models with the expectation of improving robustness and quantifying predictive uncertainty. Poorly designed MMEs may create a false sense of confidence and lead to suboptimal policy and market decisions. This perspective argues that robust decision-making-relevant MMEs must be grounded on two pillars: (1) rigorous Model Intercomparison Projects (MIPs), which identify inter-model agreement and disagreement, characterize model uncertainties, and evaluate robustness with observationally based benchmarks—MIPs' diagnostic evaluation is so critical that it must be needed to drive MME's decision in model selection and weighting, especially when only a limited number of models available; and (2) co-design by both stakeholders and scientists to ensure that scenarios, metrics and uncertainty requirements provide decision-relevant information. Building upon the past success and lessons from the existing MIPs-MMEs efforts (e.g., climate/Earth system/crop), we derived the theoretical basis for MMEs, addressed their specific challenges in ecosystem modeling, and highlighted proper consideration of model numbers and diversity, risk of model inter-dependence, effective calibration of model parameters, possible overdue of some ecosystem model development, critical roles of open benchmark data across a wide range of conditions, and suggested use of Artificial Intelligence to support MIPs-MMEs. We highlighted the under-recognized opportunity for MIPs and MMEs to drive scientific progress and innovation through identifying better performing models, systematic benchmarking, feedback loops, and targeted model improvement. By following actionable best practice guidelines, MMEs can evolve from ad hoc aggregation of models into a trusted backbone of environmental policy and decision-making.

ecosystem modeling

Decision Support System

Forest-based value chains involve decisions that begin at the landscape level and extend through processing, product manufacturing, and end-use markets. However, these decisions are often made independently across sectors, with limited visibility into how upstream resource conditions, incentives, and land management choices influence downstream production systems. In forested regions of the United States, wildfire risk, fragmented ownership, and uncertain markets for low-value residues complicate efforts to align extraction, processing, and utilization decisions. Without tools that link these stages, stakeholders may overlook opportunities to improve resource utilization or inadvertently shift impacts elsewhere in the value chain. This repository introduces a decision support system (DSS) that applies a system-impact-analysis approach to forest biomass residues and co-products. The framework integrates forest inventory data, geospatial resource assessments, and economic modeling to evaluate how biomass extraction decisions influence downstream product pathways. By linking regional feedstock avail- ability with market incentives and processing options—such as fuels, wood products, or soil amendments like biochar—the tool allows decision-makers to compare value chain outcomes across multiple utilization strategies.

Davis, Maggie [Oak Ridge National Laboratory (ORNL

Virtual Reality for Shoot/No-Shoot Decision Training in Law Enforcement: A Literature Review and Research Agenda

Virtual reality (VR) can materially improve “shoot / no-shoot” (SNS) training by giving officers realistic, repeatable practice making high-stakes decisions under pressure. Traditional tools—live-fire ranges and video simulators—build basics, but they cannot adapt to each officer in real time or fully mirror the complexity of the field. VR closes that gap by creating immersive scenarios that are safer, more flexible, easier to scale across units, and able to capture objective performance data. SNS decisions are not just about marksmanship; they rely on perception, judgment, memory, and the ability to hold fire when a threat is uncertain. Effective training therefore needs realism, decision complexity, and branching outcomes that reflect the true consequences of choices. These elements strengthen recognition of hostile intent while reducing false positives and building the self-control required in ambiguous situations. VR brings specific advantages: dynamic environments, full-body interaction, and the ability to measure performance with precision—enabling targeted feedback and better transfer of learning to the street. At the same time, responsible deployment must address scenario quality (credible environments and behaviors), lawful decision models, and user wellbeing (appropriate stress levels, comfort, and safety). Sandia’s VIPER Lab is positioned to lead this work. The team combines human-performance science, AI/ML, and VR/AR development with a deep equipment bench (e.g., omnidirectional treadmill, eye-tracking, haptics, multiple HMDs). This ecosystem supports building and validating next-generation SNS training that is immersive, measurable, and trustworthy. Bottom line: Investment in VR-enabled SNS training that blends evidence-based design with careful validation and legal safeguards is expected to pay off in safer, more consistent decision-making and improved community trust, delivered through training that is practical to deploy at scale.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF

A comprehensive academic and industrial survey of blockchain technology for the energy sector using fuzzy Einstein decision-making

The global energy sector is undergoing a significant transformation driven by decarbonization and digitalization, leading to the emergence of Distributed Ledger Technology (DLT) — particularly blockchain — as a promising tool for enhancing transparency, security, and efficiency in modern power systems. This study aims to provide a comprehensive academic and industrial survey of blockchain applications in the energy sector and develop a robust decision-making framework to identify and prioritize the most promising real-world use cases based on multidisciplinary criteria. A three-stage methodology was adopted: (i) a literature and market review encompassing over 300 academic publications and commercial blockchain initiatives in energy, (ii) an in-depth evaluation of the evolution and viability of blockchain initiatives in energy with the help of expert surveys, and (iii) a novel decision-making model using a q-rung orthopair fuzzy Multi-Attributive Border Approximation (q-ROF-MABAC) method under the Einstein operator. The results were compared with existing decision models to validate consistency and robustness. Nine key blockchain use case categories were identified and ranked based on technical, economic, and governance dimensions. The results demonstrated that integrating expert insights into a fuzzy logic framework helps filter out overhyped claims in the literature and prioritize realistic and high-impact applications such as green certificates, grid services, and peer-to-peer energy trading. The model’s rankings remained stable across varying weight configurations, confirming the robustness of the methodology. This study provides an evidence-based decision-support tool for researchers, industry stakeholders, and policymakers to better understand, evaluate, and adopt blockchain technologies in the energy sector.

29 ENERGY PLANNING, POLICY, AND ECONOMY

A dataset for understanding self-reported patterns influencing residential energy decisions

Household occupant behavior and decision-making dynamics substantially impact technology uptake and residential building energy performance. Although significant research underscores the importance of social science in energy studies, few public data with representative samples on household energy decision-making patterns are available. The dataset (UPGRADE-E: Understanding Patterns Guiding Residential Adoption and Decisions about Energy Efficiency) presents 9,919 responses from U.S. residents of single-family and small multifamily homes. Derived from a national-scale internet survey, the dataset contains 391 variables: demographics, building characteristics, home modifications, willingness to adopt new technologies, motivations for making changes, barriers, program participation, trusted information sources, and energy scenarios. Responses were validated via internal consistency checks and comparison with other U.S. national scale datasets. UPGRADE-E advances knowledge of household energy related decision-making, tying demographics, home modifications, and self-reported cognitive drivers together at a scale and breadth that has not been previously achieved. Policymakers and researchers at local, regional, and national levels may leverage this dataset to understand drivers influencing the adoption of key technologies in U.S. homes.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Persistent Classification: Understanding Adversarial Attacks by Studying Decision Boundary Dynamics

ABSTRACT There are a number of hypotheses underlying the existence of adversarial examples for classification problems. These include the high‐dimensionality of the data, the high codimension in the ambient space of the data manifolds of interest, and that the structure of machine learning models may encourage classifiers to develop decision boundaries close to data points. This article proposes a new framework for studying adversarial examples that does not depend directly on the distance to the decision boundary. Similarly to the smoothed classifier literature, we define a (natural or adversarial) data point to be ( γ , σ)‐stable if the probability of the same classification is at least for points sampled in a Gaussian neighborhood of the point with a given standard deviation . We focus on studying the differences between persistence metrics along interpolants of natural and adversarial points. We show that adversarial examples have significantly lower persistence than natural examples for large neural networks in the context of the MNIST and ImageNet datasets. We connect this lack of persistence with decision boundary geometry by measuring angles of interpolants with respect to decision boundaries. Finally, we connect this approach with robustness by developing a manifold alignment gradient metric and demonstrating the increase in robustness that can be achieved when training with the addition of this metric.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

The relative influences of hydrologic information and dams’ hydropower scheduling decisions on electricity price forecasts

Price dynamics in wholesale electricity markets are driven by supply and demand. In markets with hydroelectric dams, the timing and amount of hydropower offered can influence prices in similar ways to wind and solar power. Unlike variable renewable energy, however, the supply of hydropower in wholesale markets is a function of both water availability and operational decisions at dams. Dam operators maximize revenues in wholesale markets by aligning generation with the periods of highest expected prices, and these scheduling decisions may in turn influence prices. Here, we examine the relative importance of two types of information in predicting forward electricity prices: a) water availability at dams, in the form of short-to-medium-range hydrological forecasts; and b) hourly scheduling decisions at dams. Using softly coupled hydrologic, hydropower scheduling, and power systems models spanning the U.S. Western Interconnection, we quantify the importance of hydrologic forecast accuracy in correctly predicting wholesale electricity prices and compare this with the influence of dam operators’ own hourly scheduling decisions on realized market prices. We find that aligning hydropower generation schedules with the periods of high forecasted prices causes larger, inadvertent price forecast errors than imperfect hydrologic forecasts. This suggests that knowledge of how water is managed by dam operators within the week is more important than weekly inflow forecast errors when predicting forward electricity prices. Our findings have implications for optimal hydropower scheduling by region. Specifically, accounting for price effects is critical in markets dominated by hydropower capacity.

Electricity markets

Comparative evaluation and selection of heat exchangers using multicriteria decision-making

Here, this study presents a well-structured method for comparing and selecting Heat Exchanger (HE) technologies for Integrated Energy Systems (IES). The decision to select a HE for a particular IES configuration can vary greatly depending not only on engineering requirements but also on customer’s specific demand. In other words, the HE selection for IES requires a multicriteria decision-making approach, taking into account diverse technical, economic, and safety aspects, as well as the relative priorities considered by energy users. This study employs a HE evaluation approach combining multicriteria decision-making techniques widely used in various industries: quality function deployment (QFD) and analytic hierarchy process (AHP) techniques. Of particular interest is the use of the proposed method to select a high-temperature HEs that couples advanced nuclear reactors and industrial processes. To build a practical basis for comparing HEs within the proposed framework, efforts were made to identify the various HEs requirements for IES purposes. In addition, leveraging the insights obtained from the literature review and the market survey of commercial HE suppliers, a knowledge base was built to facilitate the comparison of each requirement across various HE designs. Also, evaluation metrics were identified for HE requirements with robust rational to enhance the quality of decisions made throughout the proposed evaluation process. The evaluation procedure and knowledge base described in this study can provide a useful basis for those interested in screening the appropriate HE designs for various IES scenarios.

Analytic Hierarchy Process (AHP)

A path to intelligent watersheds: coordinating the data to decision pipeline

Operations of multi-reservoir systems are challenged in-part by the interplay of complex physical processes functioning within the watershed. The employment of intelligent systems can be of aid by linking environmental sensing, information technology, data analytics, simulation and decision support to achieve a data-to-decision flow of information. A further challenge is that watershed resources are managed for multiple purposes requiring some level of coordination among numerous resource managers, asset operators and users. System intelligence in this context relies on shared community platforms (data portals, community models), and coordinated communication between decision makers. Opportunities to enrich watershed intelligence has been the subject of a roadmapping exercise for the Department of Energy’s Water Power Technologies Office which has relied on broad stakeholder engagement. Initial phases of engagement involved personal interviews and a series of virtual group meetings, which focused on identifying opportunities to improve the intelligence of the physical infrastructure within our watersheds—examples of feedback include improved sensing of snowpack and runoff, data standards for facilitated data sharing, and better forecasting tools. The latter phase of engagement involved the conduct of a case study in the Upper Colorado River basin where key stakeholders were interviewed to map how their decisions are informed by intelligence from other basin stakeholders. Our presentation will highlight the interdisciplinary flow of information in complex watershed systems and identify physical and institutional opportunities toward the strategic operation of water infrastructure.

Colorado River

A Neural Optimizer With Decision-Focused Learning for Optimal Energy Storage Operation

Here, this article introduces a neural optimizer-based framework for optimizing battery energy storage system (BESS) control for grid services, including demand charge and energy cost reduction. By leveraging decision-focused learning (DFL), the proposed framework ensures seamless integration and adaptation, significantly enhancing control performance. A patch time-series transformer is employed for peak load forecasting, incorporating aleatoric uncertainty quantification to account for forecasting uncertainties within the decision-making process. The framework utilizes a solver-in-the-loop approach to generate optimal BESS actions, which are then used to train the neural optimizer-based agent. By co-optimizing both BESS operational modes and output power within the NN, the system achieves improved performance and robustness. After initial training, the forecasting and control models are jointly fine-tuned to account for forecasting errors, further improving decision precision and efficiency through DFL. Case studies are performed to validate the performance of the framework using multiple real-world datasets, demonstrating superior performance in monthly peak load forecasting compared to state-of-the-art models. In addition, the results are compared against existing decision-making approaches. The results demonstrate a reduction in monthly peak forecasting error by approximately 15% across various performance measures and achieve an optimization gap for BESS operation that is about three times smaller compared to existing methods.

Kim, Hyeonjin [Pacific Northwest National Laborato

Decarbonization Dilemmas: Deliberating Difficult Decisions in Laboratory Design

Dive into the depths of design and decision-making, while we discuss the decarbonization of laboratory buildings! Delve into the dense domain of laboratory design and operations, where every development presents a diverse array of dilemmas and delights. Join us for these dynamic sessions focused on decoding the secrets of sustainable success. Dig deep into the dynamic world of heat pump designs and the delicate balance of heating and cooling loads. Debate between constant and variable fume hood designs, where these decisions determine outcomes. Discover the divergent paths of HVAC system implementation, from the deployment of chilled beams to the diverse array of different terminal unit types. But don't delay; decisive action is demanded for these goals! Dare to dream of decarbonization as we direct discussions on retrofitting existing building stock versus innovative new design approaches. Delve into the depths of debate and emerge with a decisive strategy for sustainable success. Discuss recent discoveries in development from experts associated with existing laboratory buildings with decarbonization goals. These insights and lessons learned will help determine the path forward in our industry's drive for decarbonization designs. Decarbonization is no easy task, but with determination, dedication, and devotion, we can defy the odds and forge a brighter future for laboratory design and operations. Let's dare to decarbonize together!

decarbonization

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING

Scenario Generation for Built Environment Decision Support under Uncertainty: Case Studies of Airflow Modeling and Climate-Resilient Infrastructure System Design

When confronted with unforeseen challenges, practicing informed decision making is crucial for enhancing resilience in the built environment. While scan-to-building information modeling (BIM) is a well-established approach for creating detailed digital representations of physical assets, its application in assessing and improving infrastructure resilience remains underexplored. This study addresses this gap by proposing a novel application of scan-to-BIM, namely, scan-to-BIM-to-digital twin (S-BIM-DT) workflow. By integrating reality capture and digital twin technologies, this workflow creates continuously updated and accurate digital representations of physical assets, enabling the generation of various scenarios. Unlike traditional methods, the S BIM-DT workflow facilitates continuous model refinement, supporting informed resilience strategies. By combining these technologies into a cohesive process, the workflow facilitates decision making under uncertainty, enabling stakeholders to evaluate and respond to various scenarios effectively. We demonstrate the implementation of the S-BIM-DT workflow through two use cases that highlight its capability to enhance resilience at different scales. The first use case involves the Combined Transportation, Emergency, and Communications Center (CTECC) in Austin, Texas. BIM-enriched computational fluid dynamics (CFD) modeling simulates airflow and develops alternative scenarios for optimizing the heating, ventilation, and air conditioning (HVAC) systems. This approach enhances resilience against airborne health threats in a postCOVID context. The second use case focuses on designated areas within Beaumont, Texas, as part of the Southeast Texas Urban Integrated Field Laboratory (SETx-UIFL) research. By developing inundation maps to assess extreme weather events, this modeling aids in preparedness efforts and informs the development of climate-resilient infrastructure in vulnerable neighborhoods. Results indicate that the S-BIM-DT workflow effectively generates scenarios that enhance resilience in the built environment by facilitating informed decision making. Furthermore, this study serves as a bridge between advanced scan-to-BIM methodologies and the practical strategies needed to improve built infrastructure resilience.

Built environment

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie