Search NASA⌕ Search

SEARCH · Search NASA

Results for “Multi - threaded programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Developing a robust strength model using physically-informed genetic programming

The strength of materials is influenced by a range of external conditions, such as temperature and deformation rate. Consequently, materials that demonstrate substantial variations in their mechanical behavior due to fluctuations in temperature and strain rate require complex strength models to accurately predict material performance in real-world applications. To predict such complex behavior, a robust and flexible strength model is necessary. In this work, we utilize genetic programming-based symbolic regression (GPSR) to develop data-driven strength models that accurately represent the measured stress–strain responses of tin across a wide range of strain, strain rate and temperature regimes. The GPSR models are constrained by physically-informed conditions, which leads to significant improvement in extrapolation. The best model is integrated into a multi-physics code to perform Taylor impact simulations, validating the model’s accuracy and robustness. In conclusion, the model predictions showed excellent agreement with experimental results, particularly when compared to predictions using traditional strength models.

Genetic programming↗

An interactive machine learning platform for analyzing multi-particle coincidence data from cold target recoil ion momentum spectroscopy

We present SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training), a comprehensive software platform for analyzing tabulated high-dimensional multi-particle coincidence data from Cold Target Recoil Ion Momentum Spectroscopy (COLTRIMS) experiments. The software addresses critical challenges in modern momentum spectroscopy by integrating advanced machine learning techniques with physics-informed analysis in an interactive web-based environment. SCULPT implements uniform manifold approximation and projection for non-linear dimensionality reduction to reveal correlations in high-dimensional data. We also discuss potential extensions to deep autoencoders for feature learning and genetic programming for automated discovery of physically meaningful observables. A novel adaptive confidence scoring system provides quantitative reliability assessments by evaluating user-selected clustering quality metrics with predefined weights that reflect each metric’s robustness. The platform features configurable molecular profiles for different experimental systems, interactive visualization with selection tools, and comprehensive data filtering capabilities. Utilizing a subset of SCULPT’s capabilities, we analyze photo-double-ionization data measured using the COLTRIMS method for three-body dissociation of the D 2 O molecule, revealing distinct fragmentation channels and their correlations with physics parameters. The software’s modular architecture and web-based implementation make it accessible to the broader atomic and molecular physics community, significantly reducing the time required for complex multi-dimensional analyses. This opens the door to finding and isolating rare events exhibiting non-linear correlations on the fly during experimental measurements, which can help steer exploration and improve the efficiency of experiments.

Artificial neural networks↗

Low-Mass and High-Efficiency Engine for Medium-Duty Truck Applications

This collaborative multi-year advanced engine technology project summarized in the report is to develop low mass and high efficiency medium-duty truck engine. The project was proposed as a large-scale engine design and demonstration enabled by an advanced materials and manufacturing development program, with a comprehensive plan spanning a period of four years (two phases).

33 ADVANCED PROPULSION SYSTEMS↗

Integrated Molten Salt Reactor Modeling Capabilities in NEAMS Thermal Hydraulics Tools

The DOE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program supports a full range of computational thermal fluids analysis capabilities and code developments for a broad range of advanced reactor concepts. The research and development approach under the thermal fluids technical area synergistically combines three length and time scales in a hierarchical multi-scale approach. To enable multi-scale thermal fluids capability using these codes, a key joint effort has been underway to develop an integrated system- and engineering-scale thermal fluids analysis capability, through integration of SAM and Pronghorn codes, both based on the MOOSE framework.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Control and Optimization of Energy Storage System in Power Distribution System

The widespread adoption of electric vehicles (EVs) and transportation electrification is encumbered by two chief barriers: i) the limited driving range of EVs in the market today and ii) inadequate fast-charging infrastructure for long-distance trips. Extreme fast charging (XFC) technology can recharge EVs in less than 10 minutes for 200 miles range. Firstly, a novel robust optimization-based mixed integer linear programming model is proposed to size a battery energy storage system (BESS) and PV system in an XFCS. In this part, it is assumed that the sizing and location of the XFCS are known. Secondly, the aforesaid assumption is relaxed, and a strategic multi-period coordinated planning model is proposed to optimally site and size BESS-assisted charging stations in a highway transportation network and PV systems in a power distribution network by considering the coupling between both networks. Optimal operation and control of BESS-assisted EV charging stations are vital to alleviate the adverse impact of extreme fast charging of EVs on the host power network. A joint solution is proposed to mitigate the steady state and transient impact of extremefast charging of EVs and ensure grid-friendly integration of XFCSs with the host grid. Lastly, to make the operation of the XFCS cost-effective, a multi-layered energy management framework is proposed for the XFCS by considering forecast uncertainties, monthly demand charges reduction, and BESS degradation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Integrated Molten Salt Reactor Modeling Capabilities in NEAMS Thermal Hydraulics Tools

The DOE neams program supports a full range of computational thermal fluids analysis capabilities and code developments for a broad range of advanced reactor concepts. The research and development approach under the thermal fluids technical area synergistically combines three length and time scales in a hierarchical multi-scale approach. To enable multi-scale thermal fluids capability using these codes, a key joint effort has been underway to develop an integrated system- and engineering-scale thermal fluids analysis capability, through integration of SAM and Pronghorn codes, both based on the MOOSE framework. This report summarizes recent advances in developing an integrated system- and engineering-scale modeling capability for the msr concept, which has gained significant interest in recent years. A consistent framework was established by coupling Pronghorn and SAM through the Saline interface, with thermophysical properties provided by the Molten Salt Thermal Property Database (MSTDB-TP). Further improvements were made to the coupling schemes and domain-overlapping strategies, enhancing the stability and robustness of multi-code simulations. Verification and validation efforts demonstrate the accuracy of this integration across a range of benchmark problems, including one-dimensional heated pipe flows, three-dimensional natural convection loops with evolving isotopic compositions, and \gls{msre} demonstration cases. Within Pronghorn, new capabilities were introduced to model corrosion and noble-metal plating phenomena, supported by an extended thermal-hydraulics framework and refined turbulence treatments. To capture two-phase flow behavior, a multiphase Euler–Euler model was implemented in Pronghorn, including advanced closure relations, high-resolution advection techniques, and capillary force reconstruction. Preliminary verification cases confirm the fidelity of the approach, while planned validation efforts target canonical multiphase benchmarks and application to msr components such as the msre pump bowl. Finally, updates to SAM’s msr mass transfer modeling were extended to consider noble gas migration into porous structures like graphite. The point kinetics model was updated to include reactivity feedback contributions from any defined species, such as xenon. The gas transport model was expanded for applicability to gas mixtures, bubble efflux phenomena, and species transport between liquid and gas phases. A selection of multi-scale Sherwood number correlations from MOSCATO/NekRS and multi-phase correlations from literature have been added for improved accuracy in calculating mass transfer coefficients. A companion effort on developing system-level redox corrosion has also been incorporated into SAM. Collectively, these enhancements strengthen the predictive capability of SAM and Pronghorn for simulating MSR thermal-hydraulics, corrosion, multiphase behavior, and fission-product transport, providing a more complete toolset for design, safety analysis, and licensing support of next-generation \gls{msr}s.

42 - ENGINEERING↗

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Optimal Operation of Residential High Performance Water Heater for Reduction of Electricity Cost and Peak Demand Through Field Validation

Water heating accounts for about 18% of a typical US home’s energy use. Modern water heaters have enabled control options through APIs, offering customers the opportunity to reduce their energy cost and peak demand by dynamically adjusting settings. A water heater’s capacity to store energy using its storage tank makes it an asset for peak demand reduction and energy cost savings. For this reason, a mixed-integer linear programming model is proposed to minimize the energy cost of a high-performance water heater while also reducing the peak demand of the residential household under a time-of-use utility rate by dynamically changing the water heater’s running mode. Specifically, a multi-objective optimization model is formulated to determine the mode settings of the water heater considering hot water use, time-of-use rate, and peak demand limit of the residential household. The mode settings are associated with different dead bands of water temperature for triggering on/off action of the heat pump and heating element. A 66-gal hybrid electric high performance water heater was used for numerical simulation and practical experiments. The simulation results were well aligned with measurements of practical experiments, validating the soundness of the thermodynamic model. In addition, reductions of energy cost, enabling affordability, and reducing peak demand are demonstrated. The research team also developed a software framework with dashboards to automatically and continuously monitor and manage devices.

Liu, Guodong [ORNL] (ORCID:0000000213498608)↗

A multi-objective optimization model for cropland design considering profit, biodiversity, and ecosystem services

More sustainable agricultural methods are needed to alleviate the decreases in biodiversity and ecosystem services that have occurred because of industrial agriculture. One such method is the inclusion of alternative crops into croplands that can support biodiversity, reduce erosion and chemical runoff, and sequester carbon in the soil. However, the question of where such crops should be planted to balance competing economic and environmental objectives remains open. To this end, we develop a mixed-integer quadratically constrained program to optimize the layout of a cropland considering economic, biodiversity, greenhouse gas emissions, and water quality objectives. We include spatially varying fertilization as a decision variable in addition to crop establishment location. We further include the effect of core area and edges between different crops on biodiversity. To demonstrate the applicability of the model, we apply it to an example field, showing how the optimal cropland design changes as a decision-maker prioritizes different objectives and as edges have different impacts on biodiversity.

54 ENVIRONMENTAL SCIENCES↗

On a Simplified Approach to Achieve Parallel Performance and Portability Across CPU and GPU Architectures

This paper presents software advances to easily exploit computer architectures consisting of a multi-core CPU and CPU+GPU to accelerate diverse types of high-performance computing (HPC) applications using a single code implementation. The paper describes and demonstrates the performance of the open-source C++ matrix and array (MATAR) library that uniquely offers: (1) a straightforward syntax for programming productivity, (2) usable data structures for data-oriented programming (DOP) for performance, and (3) a simple interface to the open-source C++ Kokkos library for portability and memory management across CPUs and GPUs. The portability across architectures with a single code implementation is achieved by automatically switching between diverse fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. The MATAR library solves many longstanding challenges associated with easily writing software that can run in parallel on any computer architecture. This work benefits projects seeking to write new C++ codes while also addressing the challenges of quickly making existing Fortran codes performant and portable over modern computer architectures with minimal syntactical changes from Fortran to C++. We demonstrate the feasibility of readily writing new C++ codes and modernizing existing codes with MATAR to be performant, parallel, and portable across diverse computer architectures.

97 MATHEMATICS AND COMPUTING↗

Compiler-Driven FPGA Virtualization with SYNERGY

FPGAs are increasingly common in modern applications, and cloud providers now support on-demand FPGA acceleration in datacenters. Applications in datacenters run on virtual infrastructure, where consolidation, multi-tenancy, and workload migration enable economies of scale that are fundamental to the provider's business. However, a general strategy for virtualizing FPGAs has yet to emerge. While manufacturers struggle with hardware-based approaches, we propose a compiler/runtime-based solution called Synergy. We show a compiler transformation for Verilog programs that produces code able to yield control to software atsub-clock-tickgranularity according to the semantics of the original program. Synergy uses this property to efficiently support core virtualization primitives: suspend and resume, program migration, and spatial/temporal multiplexing, on hardware which is availabletoday.We use Synergy to virtualize FPGA workloads across a cluster of Intel SoCs and Xilinx FPGAs on Amazon F1. The workloads require no modification, run within 3--4x of unvirtualized performance, and incur a modest increase in FPGA fabric usage.

Computer Science↗

IMPACT: Design of Integrated Multiphysics Producible Additive Components for Turbomachinery

The overall objective of the IMPACT program was to enable a dramatic reduction in design maturation time for an additive hot-section turbomachinery component through the following: • A fast crack-risk producibility surrogate model generated from machine learning applied to additive process simulation data generated via exascale computing, • Linking this surrogate model to multi-physics topology optimization (TO) to enable the creation of producible, near-optimal structural/thermal designs for additive hot-section components, • Maturing this toolset to reduce hot-section component design-for-manufacturing iterations by a large fraction, and eventually, • Using these tools to develop more efficient gas turbines in much shorter design cycle times.

42 ENGINEERING↗

Formal Methods for Provably Secure Software and Firmware

This project addresses a gap observed in verifying the programming in embedded devices used in international arms control: namely verifying that embedded programming in an arms control device does exactly what it is supposed to do, no more and no less, every time without fail, and without disclosing unauthorized information accidentally or intentionally. In critical military, aerospace, and industrial safety systems this problem is sometimes addressed using formal methods (FM). This multi-year project seeks to identify formal methods toolsets useable in arms control regimes, with emphasis on applicability, ease of use, long term availability, and support.

formal methods, Arms Control Verification↗

Measurement of $\nu_\mu$ CC Interactions With Two-Proton Final State in MINERvA

This dissertation presents a measurement of charged–current (CC) muon–neutrino interactions with exactly two protons and no pions in the final state (CC~$2p\,0\pi$), using data collected by the MINERvA detector in the NuMI medium–energy beam at Fermilab. Such two–proton topologies are a sensitive probe of nuclear dynamics in the few–GeV regime, including multi–nucleon correlations (npnh, notably $2p2h$) and intranuclear final–state interactions (FSI) such as pion absorption and nucleon rescattering. A precise experimental characterization of these processes is essential both for neutrino–interaction theory and for reducing systematic uncertainties in oscillation experiments that rely on accurate modeling of neutrino–nucleus interactions. Events are selected by requiring a $\nu_\mu$ CC interaction with a reconstructed $\mu^-$ and two proton tracks originating from a common vertex in MINERvA’s finely segmented scintillator tracker, with no reconstructed mesons. Muon charge and momentum are constrained by matching to the MINOS Near Detector, while proton identification exploits energy–loss profiles and stopping–proton features. Backgrounds from pion–producing channels that enter the signal region through FSI or reconstruction effects are constrained with data–driven sidebands (Michel–electron and isolated–cluster “blob” samples) and tuned via a simultaneous fit across signal and sideband regions. To correct detector resolution and acceptance effects, the analysis employs iterative Bayesian unfolding with extensive validation: statistical pseudo–experiments, and robustness checks against generator systematic “universes” and additional strong shape warps. Single–differential cross sections are reported for three observables tailored to the two–proton final state: the opening–angle cosine $\cos\!\left(\theta_{pp}\right)$, the leading–proton momentum, and the subleading–proton momentum. Systematic uncertainties include contributions from neutrino flux, interaction modeling (e.g., npnh and resonance parameters, pion FSI), and detector response (calibration, reconstruction efficiencies). The resulting distributions provide targeted constraints on the interplay of multi–nucleon dynamics and FSI that shape CC~$2p\,0\pi$ final states on hydrocarbon. Comparisons to modern GENIE–based simulations highlight kinematic regions where model components require refinement. These measurements thus inform generator tuning and improve the reliability of neutrino–energy reconstruction strategies for current and future long–baseline oscillation programs.

Syrotenko, Vladyslav S. [Tufts U.]↗

CommBench: Micro-Benchmarking Hierarchical Networks with Multi-GPU, Multi-NIC Nodes

Modern high-performance computing systems have multiple GPUs and network interface cards (NICs) per node. The resulting network architectures have multilevel hierarchies of subnetworks with different interconnect and software technologies. These systems offer multiple vendor-provided communication capabilities and library implementations (IPC, MPI, NCCL, RCCL, OneCCL) with APIs providing varying levels of performance across the different levels. Understanding this performance is currently difficult because of the wide range of architectures and programming models (CUDA, HIP, OneAPI). We present CommBench, a library with cross-system portability and a high-level API that enables developers to easily build microbenchmarks relevant to their use cases and gain insight into the performance (bandwidth & latency) of multiple implementation libraries on different networks. We demonstrate CommBench with three sets of microbenchmarks that profile the performance of six systems. Our experimental results reveal the effect of multiple NICs on optimizing the bandwidth across nodes and also present the performance characteristics of four available communication libraries within and across nodes of NVIDIA, AMD, and Intel GPU networks.

Hidayetoglu, Mert↗

Variability of Eastern North Atlantic Summertime Marine Boundary Layer Clouds and Aerosols Across Different Synoptic Regimes Identified With Multiple Conditions

Abstract This study estimates the meteorological covariations of aerosol and marine boundary layer (MBL) cloud properties in the eastern North Atlantic (ENA) region, characterized by diverse synoptic conditions. Using a deep‐learning‐based clustering model with mid‐level and surface daily meteorological data, we identify seven distinct synoptic regimes during the summer from 2016 to 2021. Our analysis, incorporating reanalysis data and satellite retrievals, shows that surface aerosols and MBL clouds exhibit clear regime‐dependent characteristics, whereas lower tropospheric aerosols do not. This discrepancy likely arises from synoptic regimes determined by daily large‐scale conditions, which may overlook air mass histories that predominantly dictate lower tropospheric aerosol conditions. Focusing on three regimes dominated by northerly winds, we analyze the Atmospheric Radiation Measurement Program (ARM) ENA observations on Graciosa Island in the Azores. In the subtropical anticyclone regime, fewer cumulus clouds and more single‐layer stratocumulus clouds with light drizzle are observed, along with the highest cloud droplet number concentration (Nd), surface cloud condensation nuclei (CCN) and surface aerosol levels. The post‐trough regime features more broken or multi‐layer stratocumulus clouds with slightly higher surface rain rate, and lower Nd and surface CCN levels. The weak trough regime is characterized by the deepest MBL clouds, primarily cumulus and broken stratocumulus clouds, with the strongest surface rain rate and the lowest Nd, surface CCN and surface aerosol levels, indicating strong wet scavenging. These findings highlight the importance of considering the covariation of cloud and aerosol properties driven by large‐scale regimes when assessing aerosol indirect effects using observations.

54 ENVIRONMENTAL SCIENCES↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗