Search NASA⌕ Search

SEARCH · Search NASA

Results for “Scalable Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Lowering entry barriers to developing custom simulators of distributed applications and platforms with SimGrid

Researchers in parallel and distributed computing (PDC) often resort to simulation because experiments conducted using a simulator can be for arbitrary experimental scenarios, are less resource-, labor-, and time-consuming than their real-world counterparts, and are perfectly repeatable and observable. Many frameworks have been developed to ease the development of PDC simulators, and these frameworks provide different levels of accuracy, scalability, versatility, extensibility, and usability. Further, the SimGrid framework has been used by many PDC researchers to produce a wide range of simulators for over two decades. Its popularity is due to a large emphasis placed on accuracy, scalability, and versatility, and is in spite of shortcomings in terms of extensibility and usability. Although SimGrid provides sensible simulation models for the common case, it was difficult for users to extend these models to meet domain-specific needs. Furthermore, SimGrid only provided relatively low-level simulation abstractions, making the implementation of a simulator of a complex system a labor-intensive undertaking. In this work we describe developments in the last decade that have contributed to vastly improving extensibility and usability, thus lowering or removing entry barriers for users to develop custom SimGrid simulators.

97 MATHEMATICS AND COMPUTING↗

Deep Learning Advances Arctic River Water Temperature Predictions

The accelerated warming in the Arctic poses serious risks to freshwater ecosystems by altering streamflow and river thermal regimes. However, limited research on Arctic River water temperatures exists due to data scarcity and the absence of robust methodologies, which often focus on large, major river basins. To address this, we leveraged the newly released, extensive AKTEMP data set and advanced machine learning techniques to develop a Long Short-Term Memory (LSTM) model. By incorporating ERA5-Land reanalysis data and integrating physical understanding into data-driven processes, our model advanced river water temperature predictions in ungauged, snow- and permafrost-affected basins in Alaska. Our model outperformed existing approaches in high-latitude regions, achieving a median Nash-Sutcliffe Efficiency of 0.95 and root mean squared error of 1.0°C. The LSTM model learned air temperature, soil temperature, solar radiation, and thermal radiation—factors associated with energy balance—were the most important drivers of river temperature dynamics. Soil moisture and snow water equivalent were highlighted as critical factors representing key processes such as thawing, melting, and groundwater contributions. Glaciers and permafrost were also identified as important covariates, particularly in seasonal river water temperature predictions. Our LSTM model successfully captured the complex relationships between hydrometeorological factors and river water temperatures across varying timescales and hydrological conditions. This scalable and transferable approach can be potentially applied across the Arctic, offering valuable insights for future conservation and management efforts.

54 ENVIRONMENTAL SCIENCES↗

Scalable bottom-up synthesis of Co-Ni–doped graphene

Introducing heteroatoms into graphene is a powerful strategy to modulate its catalytic, electronic, and magnetic properties. At variance with the cases of nitrogen (N)– and boron (B)–doped graphene, a scalable method for incorporating transition metal atoms in the carbon (C) mesh is currently lacking, limiting the applicative interest of model system studies. This work presents a during-growth synthesis enabling the incorporation of cobalt (Co) alongside nickel (Ni) atoms in graphene on a Ni(111) substrate. Single atoms are covalently stabilized within graphene double vacancies, with a Co load ranging from 0.07 to 0.22% relative to C atoms, controllable by synthesis parameters. Structural characterization involves variable-temperature scanning tunneling microscopy and ab initio calculations. The Co- and Ni-codoped layer is transferred onto a transmission electron microscopy grid, confirming stability through scanning transmission electron microscopy and electron energy loss spectroscopy. This method holds promise for applications in spintronics, gas sensing, electrochemistry and catalysis, and potential extension to graphene incorporation of similar metals.

Science & Technology - Other Topics↗

Geospatial Data Workflow Orchestration and Architecture

In an era characterized by explosive growth in geospatial data, the selection of appropriate technologies for data storage, processing, and orchestration is critical for organizations aiming to maintain competitive advantages. This white paper provides a comprehensive analysis of how Oak Ridge National Laboratory (ORNL) has effectively employed various cloud technologies, including containerized applications, container orchestrators, and workflow orchestrators, to develop robust geospatial data processing solutions. We explore the fundamental concepts behind these technologies and compare multiple deployment models tailored to diverse use cases. Our findings conclude that while Kubernetes has emerged as the preferred platform for truly scalable and fault-tolerant production workflows, the choice of workflow orchestration tool requires careful consideration of team needs, pipeline complexity, and deployment environments. This paper aims to serve as a strategic guide for organizations leveraging geospatial data, articulating the balance between technology choices and practical implementation to enhance workflow efficacy and scalability.

97 MATHEMATICS AND COMPUTING↗

Machine-learning-based estimates of global natural vegetated wetland methane emissions (2000–2025)

Wetlands are the largest natural source of atmospheric methane (CH 4 ), yet comprehensive global budgets are typically delayed by years, preventing a timely understanding of CH 4 sources, sinks, and trends. To reduce this delay, we present a model emulator-driven framework and accompanying workflow that enable timely, continuous emission updates using a machine-learning emulator to reconstruct spatially explicit monthly emission fields at 1° × 1° resolution. We apply this framework to a global dataset of natural vegetated wetland CH 4 emissions to extend the most recent Global Methane Budget (GMB; Saunois et al., 2025) record that covers the 2000–2020 emissions through 2025. In the test data (∼ 30 % of the total dataset), the emulator achieved a global R 2 of 0.65 ± 0.003 (mean ± 95 % CI, hereafter) and an RMSE of 5.49 ± 0.12×10 -3 Tg CH 4 yr −1 . The emulator is trained on 35 GMB model estimates, including 22 process-based models and 13 atmospheric inversions, paired with 10 ensemble realizations of 11 gridded climate predictor variables from atmospheric reanalyses. Our results show that the global mean predicted wetland CH 4 emissions for 2021–2025 (157.8 ± 2.4 Tg CH 4 yr −1 ) are not significantly higher (∼ 0.05 Tg CH 4 yr −1 ) than the 2000–2020 baseline. However, this stability masks a significant hemispheric redistribution of emissions. We detect an increase in Northern Hemisphere (NH) emissions in 2021–2025, with mid- and high-latitudes increasing by 0.76 ± 0.07 and 0.35 ± 0.03 Tg CH 4 yr −1 , respectively, while the tropics and Southern Hemisphere (SH) extratropics show offsetting negative trends (−0.95 ± 0.19 and -0.11 ± 0.02 Tg CH 4 yr −1 , respectively). The predicted emissions are able to capture the low emissions in 2023 in South America linked to El Niño-related drought, as reported by recent studies (Ciais et al., 2026; Quinn et al., 2025). Furthermore, we identify a distinct seasonal amplification of global emission trends that peaks in late boreal summer. This new modeled dataset and operational framework bridge the gap between the latest updated budgets and low-latency monitoring, providing a scalable capacity to frequently update global emission estimates and critical early warnings of regional wetland feedback loops. The data are publicly available at https://doi.org/10.5281/zenodo.18870108 (Li et al., 2026).

Li, Mengze [National University of Singapore (Sing↗

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection↗

Model-Based Energy and Cost Analysis of Direct Air Capture Using ePTFE-Based Laminate-Structured Gas–Solid Contactors

Carbon dioxide removal (CDR) technologies will play a significant role in limiting global warming if implemented on a large scale. Direct air capture (DAC) is a scalable approach for removing atmospheric carbon, yet the true scope of its scalability remains unclear due to the early stage of technology development and high first plant costs. This study provides groundwork for understanding the technoeconomic trade-offs in developing DAC systems using laminate-structured gas–solid contactors, encompassing the analysis of both contactor and process design spaces. The robust mass transfer and process models outlined in this study provide tools for evaluating DAC processes and designing DAC plants based on cost and energy analysis. First, the key contactor geometrical parameters are identified to understand the CO 2 productivity–energy demand trade-offs, where geometries yielding higher mass transfer rates can achieve higher CO 2 productivities at the expense of energy consumption by fans and steam use. Next, a detailed process parametric study is conducted for DAC systems coupled with steam-assisted temperature-vacuum swing adsorption (S-TVSA) to visualize the trade-offs in the multidimensional design space. The main cost driver dramatically changes over different process conditions, but the operating cost prevailed on the Pareto front, with potential to operate as low as 150 $/tonne-CO 2 (within the cost range of 148–504 $/tonne-CO 2 in this study where the DAC system is coupled with industrial facilities for steam production).

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Numerical simulation of involute-plate research reactor flow behavior using RANS, LES and DNS

This paper investigates the flow behavior of involute-plate research reactors by performing Reynolds-Averaged Navier Stokes simulation (RANS), Large Eddy Simulation (LES) and Direct Numerical Simulation (DNS) of the channel flow between fuel plates. By modeling turbulence with different numerical approaches, this study provides data with three levels of fidelity. For the RANS simulation, three widely used turbulence models, i.e., k-ε, k-ω, Reynolds Stress Turbulence model (RST) are applied by using the commercial CFD code STAR-CCM +. For LES and DNS, the open-source CFD code, Nek5000, is used given its outstanding scalability on High Performance Computer (HPC) and high-order technique. The results from RANS simulations are compared with that from LES and DNS for benchmarking. Both macroscale parameters and turbulence statistics, such as velocity magnitude, lateral velocity and turbulence kinetic energy, are presented and analyzed. The results from RANS simulation achieve good agreement with LES and DNS on velocity and turbulence kinetic energy prediction. The RST turbulence model predicts the most similar flow pattern of lateral velocity as compared to LES and DNS. The Lambda-2 (λ2) criterion with a reasonable threshold is used to demonstrate the instantaneous vortices distribution in the involute channel from both LES and DNS calculation. The DNS simulation captures more detailed turbulence especially near the corner, which explains the discrepancy between LES and DNS results near the corner. The normalized RMS error are defined and calculated to assess the performance of those turbulence models. The RST model captures the anisotropic feature of turbulence, which enable it to outperform other turbulence models for predicting the flow behavior in an involute channel. Although some discrepancies are found between LES and DNS results in the corner, the overall deviations between LES and DNS are found to be small. In conclusion, given that the computational cost of DNS calculation is an order of magnitude higher, using LES data for benchmarking RANS model is a cost-effective approach.

DNS↗

Streamlining Ocean Dynamics Modeling with Fourier Neural Operators: A Multiobjective Hyperparameter and Architecture Optimization Approach

Training an effective deep learning model to learn ocean processes involves careful choices of various hyperparameters. We leverage DeepHyper’s advanced search algorithms for multiobjective optimization, streamlining the development of neural networks tailored for ocean modeling. The focus is on optimizing Fourier neural operators (FNOs), a data-driven model capable of simulating complex ocean behaviors. Selecting the correct model and tuning the hyperparameters are challenging tasks, requiring much effort to ensure model accuracy. DeepHyper allows efficient exploration of hyperparameters associated with data preprocessing, FNO architecture-related hyperparameters, and various model training strategies. We aim to obtain an optimal set of hyperparameters leading to the most performant model. Moreover, on top of the commonly used mean squared error for model training, we propose adopting the negative anomaly correlation coefficient as the additional loss term to improve model performance and investigate the potential trade-off between the two terms. The numerical experiments show that the optimal set of hyperparameters enhanced model performance in single timestepping forecasting and greatly exceeded the baseline configuration in the autoregressive rollout for long-horizon forecasting up to 30 days. Utilizing DeepHyper, we demonstrate an approach to enhance the use of FNO in ocean dynamics forecasting, offering a scalable solution with improved precision.

97 MATHEMATICS AND COMPUTING↗

Reduced-order modeling for efficient cross section library development in high-temperature gas reactor pebble-bed depletion analysis

Accurate modeling of running-in and equilibrium conditions in pebble-bed reactors (PBRs) requires precise microscopic multigroup neutron cross sections. In Griffin, deterministic neutronics calculations rely on multivariate interpolation over large cross section libraries, resulting in significant memory usage and performance bottlenecks. This work, together with a companion paper on Griffin integration, explores reduced-order models (ROMs) to replace interpolation with lightweight surrogates. Several ROM techniques are benchmarked, with deep neural networks (DNNs) demonstrating superior memory efficiency, scalability, and predictive accuracy. A total of 295 DNNs were trained to build a comprehensive isotope library, integrated into Griffin through a custom LibTorch interface for depletion analysis. Initial results demonstrate that DNN-based ROMs drastically reduce memory demands while preserving accuracy, enabling finer tabulations and additional state variables without overhead. In conclusion, the framework also supports online cross section generation and real-time DNN updates through transfer learning, improving fidelity by capturing self-shielding and evolving nuclide compositions during burnup.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Description of FY25 Theory and Simulation Performance Target: Development of an integrated modeling framework for fusion reactor design and assessment

The urgency to deliver fusion power is growing now more than ever, with increasing pressure for both public programs and private companies to meet milestones timelines and overcome significant remaining technical challenges to ensure growth of a nascent fusion industry in time to meet rapidly growing clean energy demands. With incredible advancements in computation and years of investment in fusion model development and validation, integrated modeling is poised to fill a key role in accelerating the timeline to a fusion pilot plant (FPP). Future fusion pilot plants will operate in regimes far beyond current experience, and device design will rely on physics-based prediction and extrapolation. Many concepts will also rely on simulation to assess safety (shielding, tritium management, materials activation and lifetimes), economics and scalability before the decision to build. Importantly, integrated simulation can be used to reveal and solve the complexities of system integration that may otherwise not be apparent in physical components or models developed in isolation. New experimental test facilities that produce relevant conditions to validate and resolve key technical challenges for various subsystems (materials, blankets, fuel cycle, etc.) have been repeatedly called for by the fusion community but are not yet realized. Integrated modeling has an important role in identifying realistic load conditions (thermal, electromagnetic, plasma, neutron and photon loads, etc.) and defining the components and experiments for these test facilities in order to ensure meaningful validation that sufficiently reduces modeling uncertainties and technical risk for the full integrated reactor. The Fusion REactor Design and Assessment (FREDA) SciDAC project is building a component-based integrated modeling framework & data structure to enable self-consistent, multi-fidelity, iterative optimization workflows for the fusion reactor design process. FREDA aims to shorten the time to viable designs by providing a set of flexible workflows to support the various stages of the design process using an integrated model hierarchy, ranging from the simple analytic descriptions to the highest fidelity, theory-based plasma and engineering modeling developed by the fusion and fission communities. These tools are expected to be needed for timely support of FPP design in the milestone program and in the FIRE collaboratives. The plasma simulation backbone of FREDA is IPS-FASTRAN with newly developed coupled Core-Edge Pedestal-SOL (CESOL) workflows, which is being extended to the far-SOL region up to the plasma facing components. FREDA incorporates the FERMI engineering modeling suite and will enable self-consistent evaluation of the thermal shields, limiters, blanket, magnets, and other surrounding structures with predictions of temperatures, erosion, dpa, activation, tritium generation and transport, creep, corrosion, material degradation, etc. Parametric generation of 3D CAD enables rapid iteration of component geometry in response to plasma and loading specifications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Unified Design Theory for Multi-Port Polyphase Transformers Enabling Scalable Power-Multiplexed EV Fleet Charging Systems

This paper presents a unified analytical design theory for multi-port polyphase transformers, targeting scalable and isolated high-power Electric Vehicle (EV) fleet charging systems with power multiplexing capability. As fleet electrification accelerates, conventional one-to-one charger architectures face significant challenges in infrastructure cost, peak power demand, and low utilization of installed power electronics. Power-multiplexed charging architectures, which dynamically distribute power from a shared pool of converter modules across multiple vehicles, have emerged as a promising solution. However, such architectures require scalable, isolated multi-port power interfaces capable of routing energy among multiple inputs and outputs, whose design remains complex and dependent on iterative modeling. To address this gap, the proposed theory provides closed-form expressions for self-inductance, leakage inductance, and mutual coupling terms for arbitrary multi-phase, multi-port transformer structures. The formulation enables direct synthesis of isolated multi-input and multi-output resonant converter systems without reliance on geometry-specific finite-element analysis or extensive parameter extraction. This capability is particularly critical for power-multiplexed systems, where modular converter structures must interface with multiple vehicles while maintaining galvanic isolation and flexible power allocation. The effectiveness of the proposed framework is demonstrated through the design of a 360 kW multi-phase system operating over a 700–900 VDC input and 400–1250 VDC output range. PLECS simulation results confirm accurate prediction of system behavior and validate the applicability of the approach to multi-port, power-multiplexed charging scenarios. The proposed method significantly reduces design complexity while enabling scalable, cost-effective, and fully utilized EV fleet charging infrastructure.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Data for FUN-PROSE: A Deep Learning Approach to Predict Condition-Specific Gene Expression in Fungi

mRNA levels of all genes in a genome is a critical piece of information defining the overall state of the cell in a given environmental condition. Being able to reconstruct such condition-specific expression in fungal genomes is particularly important to metabolically engineer these organisms to produce desired chemicals in industrially scalable conditions. Most previous deep learning approaches focused on predicting the average expression levels of a gene based on its promoter sequence, ignoring its variation across different conditions. Here we present FUN-PROSE—a deep learning model trained to predict differential expression of individual genes across various conditions using their promoter sequences and expression levels of all transcription factors. We train and test our model on three fungal species and get the correlation between predicted and observed condition-specific gene expression as high as 0.85. We then interpret our model to extract promoter sequence motifs responsible for variable expression of individual genes. We also carried out input feature importance analysis to connect individual transcription factors to their gene targets. A sizeable fraction of both sequence motifs and TF-gene interactions learned by our model agree with previously known biological information, while the rest corresponds to either novel biological facts or indirect correlations.

Genomics↗

The global Sahel monsoon ocean-pressure index reconciles its regional and large-scale features

Monitoring Sahelian rainfall variability is increasingly critical as climate extremes intensify across the region. Here, we develop the Sahelian Monsoon Ocean-Pressure Index (SMOPI), a novel global synthetic indicator constructed from five dynamically coherent sea-level pressure regions statistically linked to June-September Sahel monsoon rainfall. SMOPI captures intra-seasonal and interannual variability, and crucially, reflects the influence of both regional processes and large-scale teleconnections on monsoon dynamics. It aligns with the dominant rainfall variability mode in reanalyses and 29 CMIP6 models. Strong/positive SMOPI phases coincide with wet years and are associated with enhanced convergence, favorable jet configurations, and robust Pacific, Atlantic, and Indian Ocean teleconnections. Conversely, weak/negative SMOPI phases correspond to drought conditions and divergent moisture fluxes. SMOPI exposes model failures in reproducing historical droughts and offers new physical insights into rainfall-driving mechanisms. It stands out as a scalable, potentially transferable diagnostic tool for monitoring/forecasting and evaluating Sahelian monsoon rainfall under global warming.

Tamoffo, Alain T. [Helmholtz-Zentrum Hereon GmbH, ↗

Industry-driven Training and Curriculum Development Process

The development of a sustainable, skilled fusion workforce requires coordinated strategy between all sectors of fusion industry. This paper outlines a framework to align training programs with evolving technical and professional demands of fusion, including enhancing existing curricula, the establishment of new programs at educational institutions, and the identification of workforce gaps informed through industry engagement. Effective curriculum development requires input from both educators and employers to ensure that academic content reflects real-world challenges and can prepare students for successful transitions into the field. Collaborative models, such as industry-led training programs, inter-institutional partnerships, and faculty development initiatives, are highlighted as mechanisms for scalable and inclusive workforce development. Continued program success and relevance will be dependent on continuous review processes, including feedback from employers, alumni, and advisory boards. The combination of these programs supports the formation of flexible, industry-informed training pathways. This approach aims to foster a competent workforce capable of advancing fusion energy research and commercialization.

Gehrig, Monica [ORNL] (ORCID:0000000341022612)↗

Privacy-Preserving Federated Learning for Science: Challenges and Research Directions

This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific AI models, in particular, foundation models~(FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy-- an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.

Kim, Kibaek [Argonne National Laboratory (ANL)]↗