Search NASA⌕ Search

SEARCH · Search NASA

Results for “trees”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Direct evaluation of fault trees using object-oriented programming techniques

Object-oriented programming techniques are used in an algorithm for the direct evaluation of fault trees. The algorithm combines a simple bottom-up procedure for trees without repeated events with a top-down recursive procedure for trees with repeated events. The object-oriented approach results in a dynamic modularization of the tree at each step in the reduction process. The algorithm reduces the number of recursive calls required to solve trees with repeated events and calculates intermediate results as well as the solution of the top event. The intermediate results can be reused if part of the tree is modified. An example is presented in which the results of the algorithm implemented with conventional techniques are compared to those of the object-oriented approach.

Patterson-Hine, F. A.↗

Fault-Tree Compiler

Fault-Tree Compiler (FTC) program, is software tool used to calculate probability of top event in fault tree. Gates of five different types allowed in fault tree: AND, OR, EXCLUSIVE OR, INVERT, and M OF N. High-level input language easy to understand and use. In addition, program supports hierarchical fault-tree definition feature, which simplifies tree-description process and reduces execution time. Set of programs created forming basis for reliability-analysis workstation: SURE, ASSIST, PAWS/STEM, and FTC fault-tree tool (LAR-14586). Written in PASCAL, ANSI-compliant C language, and FORTRAN 77. Other versions available upon request.

Butler, Ricky W.↗

Cloud Detection from Satellite Imagery: A Comparison of Expert-Generated and Automatically-Generated Decision Trees

Automated cloud detection and tracking is an important step in assessing global climate change via remote sensing. Cloud masks, which indicate whether individual pixels depict clouds, are included in many of the data products that are based on data acquired on- board earth satellites. Many cloud-mask algorithms have the form of decision trees, which employ sequential tests that scientists designed based on empirical astrophysics studies and astrophysics simulations. Limitations of existing cloud masks restrict our ability to accurately track changes in cloud patterns over time. In this study we explored the potential benefits of automatically-learned decision trees for detecting clouds from images acquired using the Advanced Very High Resolution Radiometer (AVHRR) instrument on board the NOAA-14 weather satellite of the National Oceanic and Atmospheric Administration. We constructed three decision trees for a sample of 8km-daily AVHRR data from 2000 using a decision-tree learning procedure provided within MATLAB(R), and compared the accuracy of the decision trees to the accuracy of the cloud mask. We used ground observations collected by the National Aeronautics and Space Administration Clouds and the Earth s Radiant Energy Systems S COOL project as the gold standard. For the sample data, the accuracy of automatically learned decision trees was greater than the accuracy of the cloud masks included in the AVHRR data product.

Shiffman, Smadar↗

(abstract) Characterization of Tree Water Status and Dielectric Constant Changes of North American Boreal Forests in Combination with Synthetic Aperture Radar Remote Sensing

The occurrence and magnitude of temporal and spatial tree water status changes in the boreal environment were studied in a floodplain forest in Alaska and in four forest types of Central Canada. Under limited water supply conditions from the rooted soil zone in early spring (freeze/thaw transition) and during summer, trees show declining water potentials. Coincidental change in tree water potential, tree transpiration and tree dielectric constant had been observed in previous studies performed in Mediterranean ecotones. If radar is sensitive to chances in tree water status as reflected through changes in dielectric constant, then radar remote sensing could be used to monitor the water status of forests. The SAR imagery is examined to determine the response of the radar backscatter to the ground based observations of the water status of forest canopies. Comparisons are made between stands and also along the large North-South gradient between sites. Data from SAR are used to examine the radar response to canopy physiological state as related to vegetation freeze/thaw and growing season length.

remote sensing imaging↗

Microwave Soil Moisture Retrieval Under Trees

Soil moisture is recognized as an important component of the water, energy, and carbon cycles at the interface between the Earth's surface and atmosphere. Current baseline soil moisture retrieval algorithms for microwave space missions have been developed and validated only over grasslands, agricultural crops, and generally light to moderate vegetation. Tree areas have commonly been excluded from operational soil moisture retrieval plans due to the large expected impact of trees on masking the microwave response to the underlying soil moisture. Our understanding of the microwave properties of trees of various sizes and their effect on soil moisture retrieval algorithms at L band is presently limited, although research efforts are ongoing in Europe, the United States, and elsewhere to remedy this situation. As part of this research, a coordinated sequence of field measurements involving the ComRAD (for Combined Radar/Radiometer) active/passive microwave truck instrument system has been undertaken. Jointly developed and operated by NASA Goddard Space Flight Center and George Washington University, ComRAD consists of dual-polarized 1.4 GHz total-power radiometers (LH, LV) and a quad-polarized 1.25 GHz L band radar sharing a single parabolic dish antenna with a novel broadband stacked patch dual-polarized feed, a quad-polarized 4.75 GHz C band radar, and a single channel 10 GHz XHH radar. The instruments are deployed on a mobile truck with an 19-m hydraulic boom and share common control software; real-time calibrated signals, and the capability for automated data collection for unattended operation. Most microwave soil moisture retrieval algorithms developed for use at L band frequencies are based on the tau-omega model, a simplified zero-order radiative transfer approach where scattering is largely ignored and vegetation canopies are generally treated as a bulk attenuating layer. In this approach, vegetation effects are parameterized by tau and omega, the microwave vegetation opacity and single scattering albedo. One goal of our current research is to determine whether the tau-omega model can work for tree canopies given the increased scatter from trees compared to grasses and crops, and. if so, what are effective values for tau and omega for trees.

O'Neill, P.↗

ANTLR Tree Grammar Generator and Extensions

A computer program implements two extensions of ANTLR (Another Tool for Language Recognition), which is a set of software tools for translating source codes between different computing languages. ANTLR supports predicated- LL(k) lexer and parser grammars, a notation for annotating parser grammars to direct tree construction, and predicated tree grammars. [ LL(k) signifies left-right, leftmost derivation with k tokens of look-ahead, referring to certain characteristics of a grammar.] One of the extensions is a syntax for tree transformations. The other extension is the generation of tree grammars from annotated parser or input tree grammars. These extensions can simplify the process of generating source-to-source language translators and they make possible an approach, called "polyphase parsing," to translation between computing languages. The typical approach to translator development is to identify high-level semantic constructs such as "expressions," "declarations," and "definitions" as fundamental building blocks in the grammar specification used for language recognition. The polyphase approach is to lump ambiguous syntactic constructs during parsing and then disambiguate the alternatives in subsequent tree transformation passes. Polyphase parsing is believed to be useful for generating efficient recognizers for C++ and other languages that, like C++, have significant ambiguities.

Craymer, Loring↗

An end-to-end deep learning solution for automated LiDAR tree detection in the urban environment

Cataloging and classifying trees in the urban environment is a crucial step in urban and environmental planning; however, manual collection and maintenance of this data is expensive and time-consuming. Although algorithmic approaches that rely on remote sensing data have been developed for tree detection in forests, they generally struggle in the more varied urban environment. This work proposes a novel end-to-end deep learning method for the detection of trees in the urban environment from remote sensing data. Specifically, we develop and train a novel PointNet-based neural network architecture to predict tree locations directly from LiDAR data augmented with multi-spectral imagery. We compare this model to a number of high-performing baselines on a large and varied dataset in the Southern California region, and find that our method outperforms all baselines in terms of tree detection ability (75.5% F-score) and positional accuracy (2.28 meter root mean squared error), while being highly efficient. We then analyze and compare the sources of errors, and how these reveal the strengths and weaknesses of each approach. Our results highlight the importance of fusing spectral and structural information for remote sensing tasks in complex urban environments.

54 ENVIRONMENTAL SCIENCES↗

Sap Velocity Data for Urban Trees in Chicago, Illinois (2024-2025)

This dataset contains uncorrected sap velocity measurements using the heat ratio method (HRM) collected using ICT International SFM1x sensors at five urban sites in Chicago, Illinois, as part of the DOE CROCUS project. The data includes continuous monitoring of sap velocity from various tree species, including Maples (Acer spp.): Sugar Maple (Acer saccharum), Silver Maple (Acer saccharinum), and Red Maple (Acer rubrum); Oaks (Quercus spp.): Swamp White Oak (Quercus bicolor); American Elm (Ulmus americana); Honey Locust (Gleditsia triacanthos); Cottonwood (Populus deltoides); and Tree of Heaven (Ailanthus altissima) across Chicago State University (CSU), Northeastern Illinois University (NEIU), Northwestern University (NU), University of Illinois Chicago (UIC), and West Woodlawn "Blacks in Green" (BIG). These include both street trees and those in urban park locations. Measurements were collected at 15-20 minute intervals, depending on the sensor, and transmitted via Long Range Wide Area Network (LoRaWAN) protocols. The wireless data was collected by Sage Network (https://sagecontinuum.org/) nodes. The dataset includes sensor ID, Global Positioning System (GPS) coordinates, tree species (common and scientific names), tree identification number, diameter at breast height (DBH in cm), uncorrected sap velocity measurements (cm/hr) from both inner and outer probes, and Sage Node identifiers so the data can be mapped to related variables such as air quality and wind speed that were collected on the Sage nodes. All timestamps are in local Chicago time (CDT/CST). Quality control flags are provided using a 3-bit binary system indicating physical range violations (< -10 or > 60 cm/hr), step spikes (absolute difference > 36 cm/hr), and stuck sensor conditions (> 10 consecutive identical values). These are raw data, not corrected for wood anatomy or species-specific characteristics. Data is provided in comma separated (CSV) format. This dataset is part of a larger collection of CROCUS environmental monitoring data, including linked datasets from Air Quality Transmitter (AQT) sensors, Weather Transmitter (WXT) sensors, and Multi-Function Research LoRaWAN (MFR) Nodes. DOIs for the supporting data are provided as part of this data package.

Chicago↗

Tree-level carbon stock estimations across diverse species using multi-source remote sensing integration

Forests are critical carbon sinks, and remote sensing has been increasingly widely used for forest monitoring and biomass estimations. However, species-specific tree-level studies remain limited. In this study, we demonstrated the feasibility of integrating UAV-based LiDAR with high-resolution optical satellite imagery (0.5 m) to estimate biomass for individual trees across different species. The proposed method accurately estimated biomass for 53 trees (R² = 0.82, rRMSE = 0.44), with species-specific datasets, showing an average 25.2% increase in R² and a 14.8% reduction in rRMSE. A novel vegetation index combining forest structure parameters with vegetation indices (VIs) was developed using high-resolution multispectral satellite data (3 m) to explore its relationship with individual tree biomass. Combining forest structural parameters with VIs further improved estimation accuracy, achieving an R²of 0.89 and an rRMSE of 0.34. Species-specific datasets show an 11.6% increase in R²compared to methods without VIs, and a 22.2% improvement over methods using only VIs. SHapley Additive exPlanations (SHAP) analysis shows that the volume feature played a key role in model performance and remained stable throughout the training process. Altogether, the proposed approach enhances individual tree biomass and carbon sink estimations, showing great potential for large-scale precise forest carbon monitoring using multi-source remote sensing data.

59 BASIC BIOLOGICAL SCIENCES↗

ITreeForeCast: An integrated modeling software to simulate tree level growth and forest carbon storage

Healthy trees in forest act as a natural carbon sink, capturing carbon. As they grow, they store carbon in their trunks, leaves and roots. Not all trees store carbon at the same rate, or in the same quantities, as it depends on a variety of biophysical and climatic factors. Furthermore, although carbon estimation in trees can be complex, the precision of estimates is tightly linked to trees growth, both in diameter and height. However, the simulation of carbon uptake by forest and forest growth has each been modeled separately, and independently at differing levels of detail and spatial resolution. In this paper, we introduce ITreeForeCast, a simulation model combining the two types of modeling on a unified platform, enabling the investigation of impacts of management strategies on carbon sequestration and wood products. ITreeForeCast is a user-extendable framework that offers new opportunities to model, simulate, and visualize the dynamics of individual trees in a forest, simulate management strategies over time, and carbon uptake.

09 - BIOMASS FUELS↗

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions↗

Distributed Augmentation, Hypersweeps, and Branch Decomposition of Contour Trees for Scientific Exploration

Contour trees describe the topology of level sets in scalar fields and are widely used in topological data analysis and visualization. A main challenge of utilizing contour trees for large-scale scientific data is their computation at scale using highperformance computing. To address this challenge, recent work has introduced distributed hierarchical contour trees for distributed computation and storage of contour trees. However, effective use of these distributed structures in analysis and visualization requires subsequent computation of geometric properties and branch decomposition to support contour extraction and exploration. In this work, we introduce distributed algorithms for augmentation, hypersweeps, and branch decomposition that enable parallel computation of geometric properties, and support the use of distributed contour trees as query structures for scientific exploration. Finally, we evaluate the parallel performance of these algorithms and apply them to identify and extract important contours for scientific visualization.

97 MATHEMATICS AND COMPUTING↗

Tree architectural characteristics and stem and leaf functional traits for 17 individuals in the Central Amazon

Given recent increases in tree mortality rates in the Amazon forest following extreme drought and wind events, we tested if lower wood density and acquisitive plant functional traits were associated with increased growth and mortality for common co-occurring trees in the Central Amazon. Research was conducted at the ZF2 Research Station located north or Manaus, Brazil, managed by the Instituto Nacional de Pesquisas da Amazônia (INPA). Seventeen trees of different species with similar sizes but a range in wood density (WD) and wood traits were felled, then assessed for 27 different individual functional parameters, including whole tree architecture, stem xylem anatomical and hydraulic traits and leaf traits. Wood logs were collected at DBH, 50% stem length and at 100% stem length (at the base of the canopy). For wood anatomy samples, n=3-6 subsamples from each height. For leaf samples, 30 leaves were collected from the upper sunlit canopy. The methodology is detailed in the accompanying manuscript. The trait data are summarized in this file: "Trait_Summary.CSV". Summary Trait code abbreviations and units are described in this file: "Sample_Info_Traits_Summary.CSV". Stem traits measured along the bole from the base of the tree (DBH, diameter breast height), mid-stem, and base of the canopy are described in these files: "Sapwood_Area_height.CSV"; "Species_Info_height.CSV"; "Sample_Info_height.CSV"

54 ENVIRONMENTAL SCIENCES↗

CHESS 2025: Leaf Area Index (LAI) for meadow, shrub, tree, and understory vegetation

This dataset contains Leaf Area Index (LAI) measurements made as part of the Colorado Headwaters Ecological Spectroscopy Study (CHESS) during June and July of 2025. Data were collected in the Upper Gunnison Basin, Colorado, across three study domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). Field observations of LAI were collected within 72 hours of airborne data collection by the National Ecological Observatory Network’s Aerial Observation Platform (NEON AOP). The NEON AOP collected waveform LiDAR (Light Detection and Ranging) and imaging spectrometer data in 426 spectral bands from the visible to shortwave infrared. LAI measurements were collected using the LICOR LAI-2200C Plant Canopy Analyzer following protocols outlined in the instrument manual (LI-COR 2019). Sampling targeted four distinct vegetation types: meadows, shrubs, trees, and aspen forest understory. We have archived data separately by site type because different field methods were used for each. At meadow sites, measurements were made at the four corners of 1m x 1m plots, with the instrument moving inward toward the center of the plot. At shrub sites, we measured the canopies of individual shrubs. At tree sites, we made measurements within a 10m x 10m subplot centered around a focal tree, with 30 observations taken on a regular grid. At aspen understory sites, we measured overstory trees following the tree protocol and understory herbaceous vegetation following the meadow protocol. All measurements included above-canopy (A) and below-canopy (B) readings, with specific protocols for scattering correction measurements in direct-sun conditions. Data were processed using the R package `rlai` (Worsham 2025). This package includes functions to calculate LAI, gap fraction, apparent clumping factor (Ω), scattering correction, and other canopy metrics. Package contents: Full file descriptions appear in ‘flmd.csv’. Files named according to the convention ‘lai_*_summary_data_cleaned.csv’ contain summary values of LAI, apparent clumping factor (Ωapp), and scattering correction factors for each site. These are the analysis-ready products that most data users will work with. Files named ‘lai_*_metadata_cleaned.csv’ contain additional site-level observations made during field collection. We have also archived intermediate and supplementary data for users who wish to check our processing approach or apply alternative methods. ‘raw_lai_2200C.zip’ contains the raw files as read from the LI-COR instrument, with no processing applied, in TXT format. The zip archive contains subdirectories by site type, which are further subdivided by sampling area. Filenames correspond to the sampling site number. ‘intermediate_results.zip’ contains detailed output from the processing routines, in JSON format. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘scattering_correction_logs.zip’ contains logfiles from the implementation of Kobayashi et al.'s (2013) scattering correction algorithm. The logfiles report values of several parameters at each iteration of the algorithm, as the model converges toward a stable solution. They are intended for users who want to verify scattering correction performance. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘spot_checks.csv’ reports LAI and other values for a small number of files processed with LI-COR FV2200 software (LI-COR 2013) using the same control parameters as in our R-based approach. Additional metadata are provided in a data dictionary describing column names and definitions (dd.csv), and in a file-level metadata file (flmd.csv). All zip files can be expanded with common archive utilities. TXT, CSV, and JSON files can be ingested into R or Python computing environments or read in common text editor utilities. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. * Todorov and Worsham are co–first authors.

2018 NEON and 2025 CHESS Campaigns↗

Spring-Summer Temperatures Since AD 1780 Reconstructed from Stable Oxygen Isotope Ratios in White Spruce Tree-Rings from the Mackenzie Delta, Northwestern Canada

High-latitude delta(exp 18)O archives deriving from meteoric water (e.g., tree-rings and ice-cores) can provide valuable information on past temperature variability, but stationarity of temperature signals in these archives depends on the stability of moisture source/trajectory and precipitation seasonality, both of which can be affected by atmospheric circulation changes. A tree-ring delta(exp 18)O record (AD 1780-2003) from the Mackenzie Delta is evaluated as a temperature proxy based on linear regression diagnostics. The primary source of moisture for this region is the North Pacific and, thus, North Pacific atmospheric circulation variability could potentially affect the tree-ring delta(exp 18)O-temperature signal. Over the instrumental period (AD 1892-2003), tree-ring delta(exp 18)O explained 29% of interannual variability in April-July minimum temperatures, and the explained variability increases substantially at lower-frequencies. A split-period calibration/verification analysis found the delta(exp 18)O-temperature relation was time-stable, which supported a temperature reconstruction back to AD 1780. The stability of the delta(exp 18)O-temperature signal indirectly implies the study region is insensitive to North Pacific circulation effects, since North Pacific circulation was not constant over the calibration period. Simulations from the NASA-GISS ModelE isotope-enabled general circulation model confirm that meteoric delta(exp 18)O and precipitation seasonality in the study region are likely insensitive to North Pacific circulation effects, highlighting the paleoclimatic value of tree-ring and possibly other delta(exp 18)O records from this region. Our delta(exp 18)O-based temperature reconstruction is the first of its kind in northwestern North America, and one of few worldwide, and provides a long-term context for evaluating recent climate warming in the Mackenzie Delta region.

Canada↗

Very High Resolution Tree Cover Mapping for Continental United States using Deep Convolutional Neural Networks

Uncertainties in input land cover estimates contribute to a significant bias in modeled above ground biomass (AGB) and carbon estimates from satellite-derived data. The resolution of most currently used passive remote sensing products is not sufficient to capture tree canopy cover of less than ca. 10-20 percent, limiting their utility to estimate canopy cover and AGB for trees outside of forest land. In our study, we created a first of its kind Continental United States (CONUS) tree cover map at a spatial resolution of 1-m for the 2010-2012 epoch using the USDA NAIP imagery to address the present uncertainties in AGB estimates. The process involves different tasks including data acquisition ingestion to pre-processing and running a state-of-art encoder-decoder based deep convolutional neural network (CNN) algorithm for automatically generating a tree non-tree map for almost a quarter million scenes. The entire processing chain including generation of the largest open source existing aerial satellite image training database was performed at the NEX supercomputing and storage facility. We believe the resulting forest cover product will substantially contribute to filling the gaps in ongoing carbon and ecological monitoring research and help quantifying the errors and uncertainties in derived products.

High Resolution↗

Tree-ring cellulose δ(18)O records similar large-scale climate influences as precipitation δ(18)O in the Northwest Territories of Canada

Stable oxygen isotopes measured in tree rings are useful for reconstructing climate variability and explaining changes in physiological processes occurring in forests, complementing other tree-ring parameters such as ring width. Here, we analyzed the relationships between different climate parameters and annually resolved tree-ring δ(18)O records (δ(18)O(TR)) from white spruce (Picea glauca [Moench]Voss) trees located near Tungsten (Northwest Territories, Canada) and used the NASA GISS ModelE2 isotopically-equipped general circulation model (GCM) to better interpret the observed relationships. We found that the δ(18)O(TR) series were primarily related to temperature variations in spring and summer, likely through temperature effects on the precipitation δ(18)O in spring, and evaporative enrichment at leaf level in summer. The GCM simulations showed significant positive relationships between modelled precipitation δ(18)O over the study region and surface temperature and geopotential height over northwestern North America, but of stronger magnitudes during fall-winter than during spring–summer. The modelled precipitation δ(18)O was only significantly associated with moisture transport during the fall-winter season. The δ(18)O(TR) showed similar correlation patterns to modelled precipitation δ(18)O only during spring–summer when water matters more for trees, with significant positive correlations with surface temperature and geopotential height, but no correlations with moisture transport. Overall, the δ(18)O(TR) records for northwestern Canada reflect the same significant large-scale climate patterns as precipitation δ(18)O for spring–summer, and therefore have potential for reconstructing past atmospheric dynamics in addition to temperature variability in the region.

paleoclimate↗

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie↗