Search NASASearch

SEARCH · Search NASA

Results for “Tree Classifiers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An end-to-end deep learning solution for automated LiDAR tree detection in the urban environment

Cataloging and classifying trees in the urban environment is a crucial step in urban and environmental planning; however, manual collection and maintenance of this data is expensive and time-consuming. Although algorithmic approaches that rely on remote sensing data have been developed for tree detection in forests, they generally struggle in the more varied urban environment. This work proposes a novel end-to-end deep learning method for the detection of trees in the urban environment from remote sensing data. Specifically, we develop and train a novel PointNet-based neural network architecture to predict tree locations directly from LiDAR data augmented with multi-spectral imagery. We compare this model to a number of high-performing baselines on a large and varied dataset in the Southern California region, and find that our method outperforms all baselines in terms of tree detection ability (75.5% F-score) and positional accuracy (2.28 meter root mean squared error), while being highly efficient. We then analyze and compare the sources of errors, and how these reveal the strengths and weaknesses of each approach. Our results highlight the importance of fusing spectral and structural information for remote sensing tasks in complex urban environments.

54 ENVIRONMENTAL SCIENCES

Decision-tree structures utilizing a phase-transition material

The rich internal physics due to competing electronic phases present in phase-transition materials such as VO2 offer the potential for compact building block design for emerging non-von Neumann computing technologies. Here, based on the relaxation dynamics of an insulator-metal phase transition, we demonstrate experimentally a decision-tree classifier embedded within a single volatile resistive switching device. The tree is constructed by the combination of the voltage pulse and relaxation time and can adapt to different tasks. We use machine learning to analyze the relaxation process, enabling a predictive voltage-relaxation time phase diagram for the electrical resistance state. Classification of the etiology of the chronic cough is presented as a proof-of-principle use case. Further, our approach can be generalized to broader classes of solid-state and solid-liquid interfacial systems that demonstrate a variety of phase relaxations.

36 MATERIALS SCIENCE

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)

Machine Learning–Guided Boolean Matrix Inference for Real-Time O-RAN Conflict Detection

Open Radio Access Networks (O-RAN) are emerging, software-driven cellular architectures that promote flexibility by enabling components from different vendors to interoperate. Multiple control applications called xApps can independently adjust network parameters in near real time, often without awareness of each other's actions. This creates a system highly prone to unintended conflicts and performance degradation due to the inherent complexity of such openness. To model such systems and ultimately prevent or mitigate xApp conflicts, it is essential to understand the dynamic relationships between xApps (A), the control parameters they adjust (P), and the resulting KPI responses (K). While the mappings from A to P and from K to A can often be derived from xApp specifications, the relationship from P to K is typically hidden within the system’s dynamics and must be inferred from observed data. We propose a novel data-driven Boolean inference framework that uncovers the hidden P?K dependencies using machine learning and interpretable rule induction. Continuous parameters and KPIs are first binarized using decision tree classifiers, and a binary influence matrix L is then inferred by solving Boolean matrix equations over time. This compact representation improves interpretability and enables real-time tracking of dynamically evolving parameter-KPI dependencies. We demonstrate the effectiveness of our method in a realistic mobile handover scenario, where it accurately recovers the underlying logic and enables proactive conflict detection.

42 - ENGINEERING

Searching for Neutrino Tridents in the NOvA Near Detector

This dissertation presents a search for neutrino trident production in the NOvA near detector through the coherent ``dimuon" channel: $\nu_\mu +\hspace{1pt}\text{X} \rightarrow \nu_\mu + \mu^- + \mu^+ +\hspace{1pt}\text{X}$. Trident production is a rare, purely electroweak process with sensitivity to physics beyond the Standard Model. The theoretical background, motivation for studying the process, and previous experimental measurements are reviewed. The analysis uses data collected by the NOvA near detector (ND) from Fermilab's Neutrinos at the Main Injector (NuMI) beam between November 2014 and February 2024, corresponding to an exposure of $25.5\times 10^{20}$ protons on target. The ND is a segmented tracking calorimeter located 800~m from the beam target, receiving neutrinos with a mean energy of 2~GeV. A multi-pass background reduction strategy is implemented, including the development of a novel dimuon-specific tracking technique. Trident candidates are identified using a boost ed decision tree classifier trained on simulated signal and background events. Limited background Monte Carlo statistics necessitate the use of functional fits to sideband data, which are extrapolated to estimate backgrounds in the signal region. The unblinded data contain 9 trident-like events, with an estimated background of 5.66 $\pm$ 5.15 events. This yields a best fit estimate of 3.34 tridents compared to the Standard Model prediction of 4.66. A profiled Feldman-Cousins method is used to determine a 90\% confidence interval of [0,9.1] on the number of signal events, corresponding to an upper limit of 1.95$\times$ the Standard Model prediction. This result represents the lowest energy search for trident events to date, and the first experimental contribution to the process in 27 years.

Bowles, Reed Scott [Indiana U.]

A procedure for rule extraction from a Self-Organising plasma disruption predictor for JET

In a previous paper, a Self-Organizing Map had proven to be able to identify the regions of the plasma operative space characterizing the pre-disruptive phase at JET without relying on any a priori information. One of the strengths of this disruption predictor lies in its inherent self-organization capability. The Self-Organizing Map discovers non-trivial relationships and captures the complicated interplay of device diagnostics on the internal plasma states directly from the experimental data. Moreover, the provided model allows the visualization of high-dimensional plasma parameters and facilitates easy interrogation of the model to understand the reasons behind its correlations. In this paper, an additional step is taken towards the interpretability of models for predicting disruptions by training a Decision Tree to classify the plasma states according to the interpretation provided by the Self-Organizing Map (stable or at high risk of disruptions). The Decision tree provides a set of rules which describe the transition of the plasma towards the pre-disruptive phase as visualized in the Self-Organizing Map. The obtained rules for the database explored in the study identify four regions in the map, two of which are at risk of disruption. These regions correspond to partitions of a 3D space based on the peaking factors of the core and divertor radiation, as well as the Locked Mode. The agreement between the Self-Organizing Map answers and the rules supplied by the Decision Tree is confirmed by the comparison of the performance exhibited by the two models in the prediction of disruptions.

Setzu, Samuele [Univ. of Cagliari, Monserrato, Cag

Reference Shapefiles and Pre-trained Random Forest Classification Models for Detecting Aufeis on the North Slope of Alaska in Landsat Imagery

This dataset provides shapefiles and trained machine learning models used for aufeis detection at four sites on the North Slope of Alaska. It includes reference data for evaluating Landsat-based detection methods, supporting research on remote sensing approaches for identifying aufeis. The ReferenceData folder contains ArcGIS shapefiles of semi-automated land cover classifications for 217 Landsat Collection 2 images, categorizing pixels into six classes: aufeis, snow, ground, none, water, and cloud. The SiteBuffers.zip file includes 10-kilometer buffer shapefiles defining regions of interest around four aufeis fields (Canning21, FH1, Firth, and Kuparuk), used to test three detection techniques. Additionally, the TrainedRFModels folder contains six pre-trained Scikit-Learn Random Forest classifiers (100 trees, max depth = 30) designed to predict aufeis presence in Landsat Collection 2 Surface Reflectance images using Red, Blue, SWIR2, NDVI, and NDWI bands. This dataset supports the development and validation of remote sensing methods for mapping aufeis in Arctic environments.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES

Influence of loblolly pine anatomical fractions and tree age on oil yield and composition during fast pyrolysis

Fast pyrolysis of woody materials is a technology pathway for producing renewable fuels and chemicals. This is a presentation of isolating needles, bark, and stemwood from a single tree as well as isolating stemwood and whole tree samples from the same species of tree with different ages and pyrolyzing each individually as well as in mixtures. This gives insight into the role of tree anatomical fractions on the resulting intermediate oil product as well as into interactions between these components. The highest carbon content oil (45.1 wt% as received) was produced from a one-to-one mixture of stemwood and needles, followed by the pure stemwood (43.4–43.8 wt% as received), while the lowest oil carbon content was from a one-to-one blend of bark and needles (26.7 wt% as received). The pyrolysis oil yield (combining oil and aqueous where separation occurred) varied from 54 wt% as received (needles) to 72.3 wt% as received (stemwood). When comparing trees of different ages, we find the change in the ratio of the anatomical fractions is a dominant factor in the product composition and yields, while the product composition and yields vary slightly with tree age when only the stemwood is pyrolyzed. Here, in this study, we present the bench-scale pyrolysis, yields, and product characterization of loblolly pine feedstocks (13- vs. 23 year-old, residues, air-classified residues, whole tree, needles, bark, and stemwood).

09 BIOMASS FUELS

Measurements of Beam Spin Asymmetries in p+p0 and p´p0 Dihadron Production at CLAS12

Semi-Inclusive Deep Inelastic Scattering (SIDIS) is a powerful experimental tool for studying the internal structure and dynamics of the proton, revealing how quarks and gluons are distributed and interact within it. SIDIS describes a process where an elec tron scatters off one of the constituent quarks within the proton, causing it to undergo hadronization, creating multiple hadrons in the final state. Through factorization, the full process can be split into probabilistic components: one which describes the internal structure of the proton using Parton Distribution Functions (PDFs), and another which describes the hadronization process using Fragmentation Functions (FFs). These functions are non-perturbative quantities of Quantum Chromodynamics (QCD), meaning they cannot be calculated directly from first principles and must instead be extracted from experimental measurements. Acommon approach for accessing PDFs and FFs using SIDIS is to measure asymmetries. In this context, asymmetries correspond to subtle differences in the angular distribution of outgoing particles that arise when the spin orientation of the incoming beam or target is reversed. Because many of these effects only appear when spin is involved, they isolate specific, nuanced properties of the proton’s spin-structure that are otherwise hidden in spin averaged measurements. In practice, they show up as specific azimuthal modulations (e.g., sin ¿R, sin(¿h ´ ¿R)), whose amplitudes isolate convolutions of PDFs and FFs at leading and subleading twist. Non-zero asymmetries of these angular distributions can be traced back to unique combinations of PDFs and FFs, offering a way to probe them directly. In this work, we measure SIDIS by analyzing high energy electron-proton scattering events using the CLAS12 detector at Jefferson Lab. This study focuses on subset of SIDIS referred to as dihadron SIDIS, where pairs of hadrons — here p+p0 and p´p0 — are observed. We analyzed these dihadrons using detector data collected during Fall 2018 and Spring 2019, where longitudinally polarized electrons from the CEBAF accelerator were incident on a liquid hydrogen target. A photon classifier using a Gradient Boosted Trees (GBTs) architecture was trained using Monte Carlo simulations to reduce the amount of iv false combinatorial background p0’s. When deployed on experimental data, the model in creases our dihadron statistics by up to five-fold compared to previous CLAS12 p0 analyses. This work reports the first measurements of beam spin asymmetries for p+p0 and p´p0 dihadron production in SIDIS. The measured asymmetries offer new insights to the spin-dependent structure and dynamics within the proton, as well as the spin-dependent properties of quark fragmentation. Non-zero twist-3 sin¿R amplitudes are observed, pro viding sensitivity to the subleading twist PDF e(x). The PDF e(x) encodes quark-gluon correlations within the proton — a property that is otherwise inaccessible at leading twist. Additionally, this work measured significant twist-2 modulations carried by sin(¿h ´ ¿R) and sin(2¿h ´2¿R), providing experimental access to the helicity dihadron fragmentation function (DiFF) GK 1 . Because there is no equivalent quark helicity-dependent FF in single pion SIDIS, the DiFF GK 1 offers a unique lens into novel spin-dependent fragmentation. For instance, the twist-2 modulations observed in this study are enhanced by vector mesons created during fragmentation — a behavior predicted by phenomenological models. This study broadens our understanding of dihadron fragmentation, revealing new details about the flavor and charge dependence of hadronization.

Matousek, Gregory [Duke Univ., Durham, NC (United

Identification of low-momentum muons in the CMS detector using multivariate techniques in proton-proton collisions at $\sqrt{s}$ = 13.6 TeV

“Soft” muons with a transverse momentum below 10 GeV are featured in many processes studied by the CMS experiment, such as decays of heavy-flavor hadrons or rare tau lepton decays. Maximizing the selection efficiency for these muons, while simultaneously suppressing backgrounds from long-lived light-flavor hadron decays, is therefore important for the success of the CMS physics program. Multivariate techniques have been shown to deliver better muon identification performance than traditional selection techniques. To take full advantage of the large data set currently being collected during Run 3 of the CERN LHC, a new multivariate classifier based on a gradient-boosted decision tree has been developed. It offers a significantly improved separation of signal and background muons compared to a similar classifier used for the analysis of the Run 2 data. The performance of the new classifier is evaluated on a data set collected with the CMS detector in 2022 and 2023, corresponding to an integrated luminosity of 62 fb -1 .

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Sensor Reduction for Diversion Detection in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors are designed as a smaller, cheaper, and safer alternative to traditional nuclear power plants. Their non-traditional characteristics and prospect of mass production and deployment will likely require new approaches to nuclear safeguards. The primary proliferation concern with microreactors is the diversion of fuel material. Such diversion may produce measurable defects in key physical attributes like neutron flux, which may in turn be detectable using machine learning models. Preliminary work has demonstrated this ability for modeled nominal and diversion scenarios using large quantities of energy integrated neutron flux data. In practice, the number of available sensors for such measurements will be limited and energy integrated flux information will not be available. This work explores the ability of tree-based gradient boosted ensemble models to classify a given microreactor core is nominal or diversion, and determine the number of fuel pins diverted in the case of diversion with reduced numbers of sensors and more realistic detector responses. Classification accuracy of greater than 98% and regression errors as low as 5% of the total number of fuel pins were achieved with as few as 15 sensors, compared to 99% and 4.1% with a maximum of 240 sensors.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Machine Learning Correlation of Electron Micrographs and ToF-SIMS for the Analysis of Organic Biomarkers in Mudstone

The spatial distribution of organics in geological samples can be used to determine when and how these organics were incorporated into the host rock. Mass spectrometry (MS) imaging can rapidly collect a large amount of data, but ions produced are mixed without discrimination, resulting in complex mass spectra that can be difficult to interpret. Here, we apply unsupervised and supervised machine learning (ML) to help interpret spectra from time-of-flight-secondary ion mass spectrometry (ToF-SIMS) of an organic-carbon-rich mudstone of the Middle Jurassic of England (UK). It was previously shown that the presence of sterane molecular biomarkers in this sample can be detected via ToF-SIMS (Pasterski, M. J. et al., Astrobiology 2023, 23, 936). We use unsupervised ML on scanning electron microscopy–electron dispersive spectroscopy (SEM-EDS) measurements to define compositional categories based on differences in elemental abundances. We then test the ability of four ML algorithms─k-nearest neighbors (KNN), recursive partitioning and regressive trees (RPART), eXtreme gradient boost (XGBoost), and random forest (RF)─to classify the ToF-SIM spectra using (1) the categories assigned via SEM-EDS, (2) organic and inorganic labels assigned via SEM-EDS, and (3) the presence or absence of detectable steranes in ToF-SIMS spectra. In terms of predictive accuracy and balanced accuracy, KNN was the best performing model and RPART the worst. The feature importance, or the specific features of the ToF-SIM spectra used by the models to make classifications, cannot be determined for KNN, preventing posthoc model interpretation. Nevertheless, the feature importance extracted from the other models was useful for interpreting spectra. In conclusion, we determined that some of the organic ions used to classify biomarker containing spectra may be fragment ions derived from kerogen which is abundant in this mudstone sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa

AmeriFlux FLUXNET-1F BR-Sa1 Santarem-Km67-Primary Forest

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site BR-Sa1 Santarem-Km67-Primary Forest. This is the FLUXNET version of the carbon flux data for the site BR-Sa1 Santarem-Km67-Primary Forest produced by applying the standard ONEFlux (1F) software. Site Description - The LBA Tapajos KM67 Mature Forest site is a closed-canopy terra firme (upland) forest, located in the FLONA Tapajos, or National Forest, a 450,000 ha government conservation unit in the Brazilian Amazon. Bounded by the Tapajos River in the west and highway BR-163 to the east, the tower is located on a flat plateau (or planalto) that extends up to 40 km to the north, south, and east. The forest at the tower site is classified as primary or "old-growth"" predominantly by its uneven age distribution, emergent trees, numerous epiphytes and abundant large logs. In January 2006 and again in November 2023, falling trees hit the tower guy wires destroying the tower and halting measurements. In each case, the tower was restored and measurements resumed, first in August of 2008 (enabled by a Partnership for International Research and Education, or PIRE, grant from the U.S. National Science Foundation) and again in June 2024.

Restrepo-Coupe, Natalia [University of Arizona, Cu

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification

Distinguishing Orbiting and Infalling Dark Matter Particles with Machine Learning

Dark matter halos are typically defined as spheres that enclose some overdensity, but these sharp, somewhat arbitrary boundaries introduce nonphysical artifacts such as backsplash halos, pseudo-volution, and an incomplete accounting of halo mass. A more physically motivated alternative is to define halos as the collection of particles that are physically orbiting within their potential well. However, existing methods to classify particles as orbiting or infalling suffer from trade-offs between accuracy, computational cost, and generalizability across cosmologies. We present an efficient, yet accurate, supervised machine learning approach using decision trees. The classification is based on only the particle radii and velocities at two epochs. Compared to detailed analysis of particle trajectories, we find that our model matches the classification of 97% of particles. Consequently, we are able to quickly and accurately reproduce the density profiles of the orbiting and infalling components out to many virial radii. We demonstrate that our model generalizes to a significantly different cosmology that lies outside the training data set. We make publicly available both our final model and the code to train similar models.

79 ASTRONOMY AND ASTROPHYSICS