Search NASA⌕ Search

SEARCH · Search NASA

Results for “workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Earth Science Data Processing With Nextflow

Earth science data processing tasks present many challenges. These tasks often process large input datasets and require scores of CPU-hours to generate results. All but the simplest tasks will be decomposed into a series of computational or data manipulation steps, also known as a scientific workflow. In order to reduce the burden of orchestrating and running the dependent processing steps, a workflow execution engine is required. This poster describes the lessons learned by the CLARREO Pathfinder (CPF) team while developing multiple scientific workflows and utilizing the open-source Nextflow engine to execute them in a cloud computing environment. The Nextflow engine is designed with the following stated goals: first, the engine does not dictate how individual steps in the task are implemented (i.e. it is language and interface agnostic); second, the engine supports easy configuration and modularity at the workflow level so that others can easily execute our workflows to reproduce results; lastly, the engine eases development by transparently scaling execution from local to remote environments. Nextflow was developed for the bioinformatics domain but is a good fit for other scientific workflows where the overall task is well-described by a dataflow diagram. The CPF team has developed Nextflow pipelines (i.e. scientific workflows) to simulate CLARREO radiance, generate large look-up tables for inter-calibration algorithms, and generate L4 intercalibration data products. These pipelines consume from single-digits to hundreds of thousands of CPU-hours. In the development and evolution of these pipelines we have discovered many design patterns, pitfalls, and solutions to common problems. Our goal is to demonstrate important aspects of how to design, implement, run, and ultimately share Nextflow pipelines in the domain of Earth science.

Aron D Bartle↗

Untargeted Spatial Metabolomics and Spatial Proteomics on the Same Tissue Section

An increasing number of spatial multiomic workflows have been recently developed. Some of these approaches have leveraged initial mass spectrometry imaging (MSI)-based spatial metabolomics to inform region of interest (ROI) selection for downstream spatial proteomics. However, these workflows have been limited by varied substrate requirements between modalities or have required analyzing serial sections (i.e., one section per modality). To mitigate these issues, we present a novel multiomic workflow that uses desorption electrospray ionization (DESI)-MSI to identify representative spatial metabolite patterns on-tissue prior to spatial proteomic analyses on the same tissue section. Further, this workflow is demonstrated here with a model mammalian tissue (coronal rat brain section) mounted on a polyethylene naphthalate-membrane slide. Initial DESI-MSI resulted in 160 annotations (SwissLipids) within to the METASPACE platform (≤20% false discovery rate). A segmentation map from the annotated ion images informed downstream ROI selection for spatial proteomics characterization from the same sample. The unspecific substrate requirements and minimal sample disruption inherent to DESI-MSI allowed for an optimized, downstream spatial proteomics assay, resulting in 3888 ± 240 to 4717 ± 48 proteins being confidently directed per ROI (200 µm x 200 µm). Finally, we demonstrate the integration of multiomic information, where we found ceramide localization to be correlated with SMPD3 abundance (ceramide synthesis protein), and we also utilized protein abundance to resolve metabolite isomeric ambiguity. Overall, the integration of DESI-MSI into the multiomic workflow allows for complementary spatial and molecular-level information to be achieved from optimized implementations of each MS assay inherent to the workflow itself.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scenario Generation for Built Environment Decision Support under Uncertainty: Case Studies of Airflow Modeling and Climate-Resilient Infrastructure System Design

When confronted with unforeseen challenges, practicing informed decision making is crucial for enhancing resilience in the built environment. While scan-to-building information modeling (BIM) is a well-established approach for creating detailed digital representations of physical assets, its application in assessing and improving infrastructure resilience remains underexplored. This study addresses this gap by proposing a novel application of scan-to-BIM, namely, scan-to-BIM-to-digital twin (S-BIM-DT) workflow. By integrating reality capture and digital twin technologies, this workflow creates continuously updated and accurate digital representations of physical assets, enabling the generation of various scenarios. Unlike traditional methods, the S BIM-DT workflow facilitates continuous model refinement, supporting informed resilience strategies. By combining these technologies into a cohesive process, the workflow facilitates decision making under uncertainty, enabling stakeholders to evaluate and respond to various scenarios effectively. We demonstrate the implementation of the S-BIM-DT workflow through two use cases that highlight its capability to enhance resilience at different scales. The first use case involves the Combined Transportation, Emergency, and Communications Center (CTECC) in Austin, Texas. BIM-enriched computational fluid dynamics (CFD) modeling simulates airflow and develops alternative scenarios for optimizing the heating, ventilation, and air conditioning (HVAC) systems. This approach enhances resilience against airborne health threats in a postCOVID context. The second use case focuses on designated areas within Beaumont, Texas, as part of the Southeast Texas Urban Integrated Field Laboratory (SETx-UIFL) research. By developing inundation maps to assess extreme weather events, this modeling aids in preparedness efforts and informs the development of climate-resilient infrastructure in vulnerable neighborhoods. Results indicate that the S-BIM-DT workflow effectively generates scenarios that enhance resilience in the built environment by facilitating informed decision making. Furthermore, this study serves as a bridge between advanced scan-to-BIM methodologies and the practical strategies needed to improve built infrastructure resilience.

Built environment↗

Developing a Prototype Methodology to Rank CO2-EOR Wells and Assess Their Reuse Potential for Geologic Carbon Storage

This paper presents a prototype methodology to assess the possible transition of Class II carbon dioxide-enhanced oil recovery (CO2-EOR) wells to Class VI wells. The focus is on wellbore construction materials—casing, cement, tubing, and the packer—and includes comprehensive workflows to evaluate these materials, with primary emphasis on compliance with Environmental Protection Agency (EPA) Class VI well construction and conversion guidelines. These workflows systematically assess material properties and performance criteria to ensure regulatory compliance and optimize long-term wellbore integrity and functionality. Utilizing Python scripts and JavaScript Object Notation (JSON) representations, the study automates checks on digitized Texas Railroad Commission (TRRC) data to rank wells based on workflow criteria. By emphasizing critical factors such as casing integrity, cementing techniques, tubing compatibility, and packer selection, the methodology helps well owners and operators prioritize wells for potential reuse as CO2 injection wells. Given limitations in digitized data, manual user verification is required in some sections. Future improvements include integrating non-digitized data through web scraping and machine learning techniques. This research serves as a practical guide for stakeholders, supporting environmental compliance and sustainable well operations.

geologic carbon sequestration↗

FY25 Theory and Simulation Performance Target: Development of an integrated modeling framework for fusion reactor design and assessment (Final Report)

This report documents the FY25 Theory and Simulation Performance Target (TSPT) of developing an integrated modeling framework for fusion reactor design and assessment (FREDA). Over Q1-Q4, new capabilities were developed across both plasma and engineering domains and demonstrated on an example representation of a Compact Advanced Tokamak with a Dual Cooled Lead Lithium blanket. This represents a first-of-a-kind demonstration of coupled core-to-wall-to-engineering for a reactor. Self-consistent CESOL workflows were applied to provide core, pedestal, and SOL prediction; new modules were developed for energetic particle stability (FAR3D) and transport (TGLF-EP) analysis; and boundary plasma modeling (SOLPS-ITER, BOUT++/Hermes-3) was expanded to evaluate wall and divertor heat fluxes and interface with engineering thermal analysis. A parameterized CAD tool, TRACER, was expanded to generate medium-fidelity divertor, blanket, and coil geometries; OpenFOAM and Diablo workflows were applied for first-wall and divertor thermal analyses with helium cooling; and reduced-order models were created for high-mass-flux divertor cooling. Magnet multiphysics capabilities were verified between Elmer, Diablo, and a new MFEM-based solver, and workflows enable stress, thermal, and neutron-fluence analysis of TF coils with neutronics-driven heating. Nuclear and blanket analysis workflows were demonstrated, including tritium breeding, transport, and CFD-informed thermo-mechanical assessment. Preliminary multi-fidelity uncertainty quantification workflows were applied to boundary modeling codes and shown to achieve variance reductions with fewer high-fidelity boundary simulations. Key findings highlight the challenges of resolving the ITEP gap to find suitable balance between wall and divertor loads, neutron heating, and practical limits of PFC cooling. Next step priorities are to develop automated workflows to check boundary code convergence and detachment, implement tighter physics-engineering CAD provenance tracking, and inclusion of plasma-material interface models for SLAG and tungsten cracking behavior. Collectively, these developments establish sophisticated capabilities for predictive, multi-fidelity, whole-device modeling that integrates plasma physics, materials, magnets, and nuclear engineering to guide pathways to viable Fusion Pilot Plant design points.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Utilizing Distributed Heterogeneous Computing with PanDA in ATLAS

In recent years, advanced and complex analysis workflows have gained increasing importance in the ATLAS experiment at CERN, one of the large scientific experiments at LHC. Support for such workflows has allowed users to exploit remote computing resources and service providers distributed worldwide, overcoming limitations on local resources and services. The spectrum of computing options keeps increasing across the Worldwide LHC Computing Grid (WLCG), volunteer computing, high-performance computing, commercial clouds, and emerging service levels like Platform-as-a-Service (PaaS), Container-as-a-Service (CaaS) and Function-as-a-Service (FaaS), each one providing new advantages and constraints. Users can significantly benefit from these providers, but at the same time, it is cumbersome to deal with multiple providers, even in a single analysis workflow with fine-grained requirements coming from their applications’ nature and characteristics. In this paper, we will first highlight issues in geographically-distributed heterogeneous computing, such as the insulation of users from the complexities of dealing with remote providers, smart workload routing, complex resource provisioning, seamless execution of advanced workflows, workflow description, pseudointeractive analysis, and integration of PaaS, CaaS, and FaaS providers. We will also outline solutions developed in ATLAS with the Production and Distributed Analysis (PanDA) system and future challenges for LHC Run4.

97 MATHEMATICS AND COMPUTING↗

Streaming Large-Scale Microscopy Data to a Supercomputing Facility

Data management is a critical component of modern experimental workflows. As data generation rates increase, transferring data from acquisition servers to processing servers via conventional file-based methods is becoming increasingly impractical. The 4D Camera at the National Center for Electron Microscopy generates data at a nominal rate of 480 Gbit s -1 (87,000 frames s -1 ⁠), producing a 700 GB dataset in 15 s. To address the challenges associated with storing and processing such quantities of data, we developed a streaming workflow that utilizes a high-speed network to connect the 4D Camera’s data acquisition system to supercomputing nodes at the National Energy Research Scientific Computing Center, bypassing intermediate file storage entirely. In this work, we demonstrate the effectiveness of our streaming pipeline in a production setting through an hour-long experiment that generated over 10 TB of raw data, yielding high-quality datasets suitable for advanced analyses. Additionally, we compare the efficacy of this streaming workflow against the conventional file-transfer workflow by conducting a postmortem analysis on historical data from experiments performed by real users. Our findings show that the streaming workflow significantly improves data turnaround time, enables real-time decision-making, and minimizes the potential for human error by eliminating manual user interactions.

4D-STEM↗

Ensemble Simulations on Leadership Computing Systems

Scientific productivity can be enhanced through workflow management tools, relieving large High Performance Computing (HPC) system users from the tedious tasks of scheduling and designing the complex computational execution of scientific applications. This paper presents a study on the usage of ensemble workflow tools to accelerate science using the Summit and Frontier supercomputing systems. The research aims to connect science domain simulations using Oak Ridge Leadership Computing Facility (OLCF) supercomputing platforms with ensemble workflow methods in order to accelerate HPC-enabled discovery and boost scientific impact. We present the coupling, porting and optimization of Radical-Cybertools on three applications: Chroma, NAMD and LAMMPS. The tools augment traditional HPC monolithic runs with a pilot scheduler. Lessons-learned are discussed for physics, biology and materials science applications. We discuss intrinsic limitations of coupling and porting ensemble workflow tools to applications that run on large HPC systems. The origins of technical challenges and their solutions developed during the implementation process are discussed. Data management strategies, OLCF’s policies for ensembles, and natively supported workflow tools are also summarized.

Georgiadou, Antigoni [ORNL] (ORCID:000000020977631↗

Hydrologic Model Data for the East Fork Poplar Creek Watershed Simulated with the Advanced Terrestrial Simulator (ATS): Streamflow and Network Expansion–Contraction Dynamics

This dataset supports hydrologic modeling and stream network expansion–contraction analysis for the East Fork Poplar Creek (EFPC) Watershed in Tennessee. It includes a Jupyter notebook for model setup, model configuration files, simulation outputs, and derived products used to evaluate model performance and investigate stream dynamics under varying hydrologic conditions. The dataset was generated using the Watershed Workflow Python package and the Advanced Terrestrial Simulator (ATS), enabling integrated surface–subsurface hydrologic simulations using a stream-aligned mesh. Outputs include high-resolution time series of streamflow, active network length, water table depth, and related hydrologic variables. Also included are spatially explicit stream persistency indices and classifications of reaches as perennial or non-perennial. These data facilitate reproducibility and support further research on stream intermittency and variability in network extent.The model data archive is organized in following directories:1) model_setup_inputsContains the Watershed Workflow Jupyter notebooks (accessed through any open source code editor), selected input datasets, and resulting ATS input files, including XML files (access through any open source code editor), computational mesh (.exo files can be viewed using Paraview), and meteorological forcing files (.h5 files can be accessed through h5py python package and HDFView open source software). 2) model_outputsIncludes ATS simulation outputs relevant to this study. Time series of spatially integrated or averaged variables (e.g., streamflow, water table depth) are provided as CSV files. Select spatial fields (e.g., ponded depth and water table depth) are saved as pickled Python objects to reduce file size, and can be accessed through pickle package in Python. Key geometry objects from Watershed Workflow—such as the surface mesh and river tree—are also included to support analysis of streamflow persistency and expansion–contraction dynamics. These files can also be accessed through Watershed Workflow Python package.3) model_evaluationProvides observed streamflow time series and field survey-based flow regime classifications used to evaluate model performance. Jupyter notebooks for processing ATS outputs and comparing model predictions with observations to build confidence in the model prior to scientific analysis are also included.4) Q_L_relationshipsContains workflows for generating time series of discharge, active network length, and related hydrologic variables used in the stream network expansion–contraction analysis. Includes routines for delineating baseflow-dominated periods. For each catchment, notebooks and processed data (as pickled DataFrames accessed through Pandas Python package) are provided. 5) figure_scriptsProvides the Jupyter notebooks used to generate the figures presented in the paper.

54 ENVIRONMENTAL SCIENCES↗

NASA Earth eXchange (NEX) App Store

NASA Earth Exchange (NEX), and her public cloud version OpenNEX, have become platforms supporting scientific collaboration, knowledge sharing and research for the entire Earth science community. To date, a number of custom tools and capabilities have been integrated into the platforms. However, such integration has to undergo a case-by-case manual process thus lacks scalability. This timely project builds an App Store onto OpenNEX as a building block. Climate data analytics tools/programs can be easily uploaded, shared, organized, searched, and recommended like photos and videos on the YouTube. The foundation of our App Store is a provenance server, which not only records metadata but also execution history of climate data analytics apps including the input data and parameters, output data and products, who runs the app for which purpose, and how apps may be chained into workflows. Researchers can thus understand, reproduce, and repurpose existing apps and workflows. Machine learning approaches are applied to mine provenance to provide recommend-as-you-go services for Earth scientists, such as to recommend suitable apps and workflow snippets. A browser-based workflow tool is also provided for researchers to explore the provenance server and design value-added workflows. Scalability, sustainability, extensibility, usability, adaptability, security and privacy are considered in the App Store.

eXchange↗

Enabling the broader adoption of fusion simulation on complex geometry

This project addressed a key barrier to advanced fusion and nuclear simulation: the difficulty of performing high-fidelity Monte Carlo neutronics directly on complex, real-world CAD geometry. Traditional workflows require engineers to rebuild CAD models as simplified constructive solid geometry, a time-consuming and error-prone process that limits design iteration and broader adoption of simulation tools. The goal of this Phase I SBIR was to make CAD-based neutronics practical, accessible, and robust for industrial and research users. During the project, Coreform significantly enhanced the Direct Accelerated Geometry Monte Carlo (DAGMC) workflow and fully integrated it into Coreform Cubit as a first-class capability. Major achievements include optimized material assignment and surface meshing workflows, substantial performance improvements to geometry imprinting and preparation, native export of DAGMC models, and new visualization tools to support OpenMC source definition and lost-particle debugging. Coreform also expanded Cubit’s capabilities as a full OpenMC preprocessor, including the ability to convert OpenMC constructive solid geometry models back into CAD for visualization, multiphysics coupling, and debugging. In collaboration with Argonne National Laboratory, the project delivered comprehensive new DAGMC documentation and training materials, transforming DAGMC from a research-oriented tool into a production-ready workflow. Results were disseminated through tutorials, conference training, and multiple well-attended webinars demonstrating integrated CAD-based neutronics and multiphysics workflows. Overall, this project demonstrated that high-fidelity Monte Carlo simulations can be performed directly on complex CAD geometry, reducing setup time, improving usability, and enabling faster, more informed design decisions for fusion and nuclear energy systems.

42 ENGINEERING↗

GDSA framework, a computational framework for complex modeling problems in radioactive waste management

This paper details a computational framework to produce automated, graphical workflows, and how this framework can be deployed to support complex modeling problems like those in nuclear engineering. Key benefits of the framework include: automating previously manual workflows; intuitive construction and communication of workflows through a graphical interface; and automated file transfer and handling for workflows deployed across heterogeneous computing resources. This paper demonstrates the framework's application to probabilistic post-closure performance assessment of systems for deep geologic disposal of nuclear waste. However, the framework is a general capability that can help users running a variety of computational studies.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao↗

Surrogate models for development of unconventional shale reservoirs by an integrated numerical approach of hydraulic fracturing, flow and geomechanics, and machine learning

We develop well-completion surrogate models by taking an integrated workflow of hydraulic fracturing, flow, geomechanics, and machine learning simulation. There are three steps in the proposed workflow. First, history-matching processes are conducted with the field data including pumping and production data for characterization. Second, full-physics simulation is performed with various parameters of the field development (e.g., cluster spacing, clusters per stage, pumping rates and times, amount of proppant, and well spacing) to generate multiple simulation results by changing the parameters of the completion design with well-known hydraulic fracturing, reservoir, geomechanics simulators to calculate fracture geometry, reservoir depressurization, induced stress changes. The workflow is demonstrated over a field in the Southern Midland Basin. Here, we take two completion scenarios: a single well case followed by a multi-well case. Finally, a Long Short-Term Memory (LSTM) machine learning algorithm is employed to create surrogate models that can replicate the full-physics simulation results. Furthermore, results show that the trained models applied in the single well and multi-well cases for a particular geological system can provide good accuracy close to those provided by full-physics simulations. Specifically, the site-specific surrogate models can predict fracture parameters (length, height, and surface area) and cumulative production accurately with computational efficiency, suggesting our proposed workflow can be used as a pragmatic tool for expediting the well completion optimization process.

Geomechanics↗

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats↗

Machine Learning-Assisted Recovery of Delicate Kinetic Information from Transient Reactor Experiments

Identifying active sites and their roles in chemical reaction steps remains a vital challenge in heterogeneous catalysis. Transient experiments offer a unique way to probe active sites and distinguish subtle kinetic features. Although physics-based analysis methods may be well-developed, they can be highly susceptible to experimental noise, and smoothing methods may erase or even distort important features; a smooth curve is not always the best curve. We demonstrate a new workflow for the direct interpretation of intrinsic kinetic information from exit flux curves measured in transient reactor experiments. This workflow contains three artificial neural networks (ANNs), including a noise reducer, a concentration predictor, and a rate predictor to analyze experimental data, followed by the virtual TAP (VTAP) physics-based reactor model and density functional theory (DFT) calculations of adsorption energies on specific sites. We use this workflow to analyze the data from experiments titrating Pt/Al 2 O 3 and Pt/SiO 2 catalysts with carbon monoxide (CO) in the temporal analysis of products (TAP) reactor. Our workflow separates the time-evolving chemical reaction and mass transfer information contained in the TAP pulse response. The existence of strong- and weak-binding sites on the Pt/Al 2 O 3 catalyst is observed in the catalyst titration experiment in the transient reactor. The structures of the strong- and weak-binding sites are then identified by using DFT calculations. We find that the Pt/SiO 2 catalyst has only strong-binding sites, which aligns with the inactive support effect of SiO 2 . We demonstrate how machine learning methods provide unique insights with high-resolution data analysis that cannot be achieved by using state-of-the-art physics-based methods.

Adsorption↗

Challenges of conventional iterative all-atom and coarse-grained multiscale molecular dynamics

In this work, we evaluate the biomolecular dynamics behaviors when conventionally iterating between all-atom (AA) and coarse-grained (CG) molecular dynamics (MD) simulations over multiple cycles. We implemented the workflow to iterate between AA and CG in OpenMM, namely the iterative multiscale MD (iMMD) simulation workflow. In particular, we aim to identify practical applications for iterating between AA and CG simulations in a conventional manner without any constraints or model modifications. We evaluate the iMMD workflow on four representative systems, spanning folding of two soluble proteins and protein-protein as well as protein-lipid interactions of two membrane proteins. We observe that iteration between AA and CG representations could help the soluble proteins exit undesirable metastable states to fold, resulting from random protein structural distortions due to cycling. Consequently, the most reliable use of iterative AA and CG simulations appears to be to accelerating complex lipid mixing for membrane-bound protein systems rather than sampling protein conformational space. Our work explores the practical usages and limitations for iterative AA and CG simulations using readily available AA and CG force fields. The evaluated iMMD workflow in OpenMM is made available at https://github.com/lanl/iMMD.

59 BASIC BIOLOGICAL SCIENCES↗

Automated ICRF heating surrogate modeling via machine learning

This work introduces automated machine learning workflows that address critical bottlenecks in surrogate model development for Ion Cyclotron Range of Frequencies (ICRF) heating applications. The automated framework includes data analysis tools that transform raw datasets into actionable insights in seconds, replacing weeks of manual exploratory effort and ensuring consistent, reproducible dataset characterization. By integrating advanced hyperparameter optimization (HPO) methods including Bayesian optimization via BoTorch and Tree-structured Parzen Estimators (TPE), the framework significantly reduces model development time from weeks to hours, decreasing computational cost and required expertise, while enabling high-accuracy surrogate models. Compared to traditional hyperparameter scanning (HPS) techniques such as methodical, randomized, and grid searches, HPO methods achieve superior convergence and predictive performance, even when compared to already well-tuned reference models. On NSTX High Harmonic Fast Wave (HHFW) heating datasets, both Random Forest Regressor (RFR) and neural network surrogates demonstrate improved accuracy, achieving R 2 values beyond 0.97 and 0.98, respectively. The results show that while HPO gains are modest for robust architectures like RFR, they become essential for more sensitive models such as neural networks, highlighting the trade-offs across optimization strategies. Through automated workflows that eliminate manual hyperparameter tuning and require minimal ML expertise, this work enables widespread adoption of high-fidelity surrogate models across the fusion community for real-time plasma control, uncertainty quantification, rapid experimental scenario development, and integrated system optimization.

Sanchez-Villar, Alvaro [Princeton Plasma Physics L↗