Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Spherical tokamak physics research in preparation for the operation of NSTX-U

The National Spherical Torus Experiment Upgrade (NSTX-U) is preparing to resume operation, representing a crucial step toward realizing compact, cost-effective fusion pilot plants. In advance of this, extensive modeling and data analysis have been conducted to advance the physics basis for low-aspect-ratio, high-performance plasma regimes, focusing on three core objectives: confinement and stability, power and particle handling, and steady-state operation. Significant progress has been made in understanding the electron temperature flattening in high-β plasmas, which is shown to be driven by a complex interplay of magnetohydrodynamic instabilities (e.g. non-resonant infernal modes), fast-ion-driven Alfvén eigenmodes, and electron and ion-scale micro-instabilities, particularly Kinetic Ballooning Modes (KBMs), whose destabilization is strongly dependent on parallel magnetic field fluctuations (δB ∥ ). Furthermore, a new gyrokinetic critical pedestal model was developed, accurately predicting pedestal structure by identifying KBMs as the primary stability limit, offering a critical constraint for future high-confinement scenarios. To address the challenge of high heat flux, novel liquid lithium plasma-facing components were modeled. The analysis confirmed that lithium vapor shielding is a self-regulating mechanism for heat mitigation, while also emphasizing that strong main ion parallel flow is essential to minimize core lithium contamination. Finally, progress toward steady-state operation was anchored by developing the required physics basis and control tools. This includes predictive modeling for reversed magnetic shear sustainment, demonstrating that magnetic island-induced bootstrap current reduction is negligible in STs, and advancing real-time control and disruption avoidance capabilities. The development of high-speed surrogate models (e.g. MMMNet) provides computationally efficient tools vital for non-inductive scenario optimization and integrated, low-disruptivity operations planned for NSTX-U.

NSTX-U↗

Relative molecular orientation can impact the onset of plasticity in molecular crystals

Abstract Creating or moving dislocations is the first step to dissipating mechanical energy via plastic deformation under contact loading. In molecular crystals there is both a lattice that defines crystal orientation and a relative orientation of the basis of the molecules. We define a normalization parameter which relates strain at yield, the hardness of the bulk crystal, and a distance parameter analogous to a Burgers vector that nominally predicts the relative ease of initiating plasticity in this broad class of materials. Analyzing the yield behavior of 10 different molecular crystals of varying space groups shows the inter-molecular orientation predicts the experimentally observed applied stress needed to nucleate dislocations. When molecules are oriented ‘parallel’ relative to one another the normalized maximum shear stress at the onset of plasticity is on the order of 3–5 times lower than when molecules within the crystal are ‘anti-parallel’, and molecules with a more equiaxed shape fall in between these bounds. This provides an initial indication of a structural feature which predicts the relative ease of initiating plasticity during contact loading in molecular crystals.

36 MATERIALS SCIENCE↗

Linear complexity

We present factorization and solution phases for a new linear complexity direct solver designed for concurrent batch operations on fine-grained parallel architectures, for matrices amenable to hierarchical representation. We focus on the strong-admissibility-based $\mathscr{H}^{2}$ format, where strong recursive skeletonization factorization compresses remote interactions. We build upon previous implementations of $\mathscr{H}^{2}$ matrix construction for efficient factorization and solution algorithm design, which are illustrated graphically in stepwise detail. The algorithms are ‘blackbox’ in the sense that the only inputs are the matrix and right-hand side, without analytical or geometrical information about the origin of the system. We demonstrate linear complexity scaling in both time and memory on four representative families of dense matrices up to one million in size. Parallel scaling up to 16 threads is enabled by a multi-level matrix graph coloring and avoidance of dynamic memory allocations thanks to prefix-sum memory management. An experimental backward error analysis is included. We break down the timings of different phases, identify phases that are memory-bandwidth limited, and discuss alternatives for phases that may be sensitive to the trend to employ lower precisions for performance.

Boukaram, Wajih↗

Measuring and unbiasing the BAO shift in the Ly α forest with AbacusSummit

ABSTRACT The Dark Energy Spectroscopic Instrument (DESI) places sub- per cent constraints on measurements of the Baryon Acoustic Oscillation (BAO) scaling parameters from the Ly $\alpha$ forest. However, no systematic error budget stemming from non-linearities in the three-dimensional clustering of the Ly $\alpha$ forest is included in the DESI-Ly $\alpha$ analysis. In this work, we measure the size of the shift of the BAO peak using large Ly $\alpha$ forest mocks produced on the N-body simulation suite AbacusSummit, which adopt the Fluctuating–Gunn–Peterson Approximation (FGPA). Specifically, we measure the Ly $\alpha$ autocorrelation and the Ly $\alpha$-quasar cross-correlation functions. To mitigate the noise, we adopt a linear control variates technique, reducing the error bars by a factor of up to $\sim \sqrt{50}$ on large scales. From the autocorrelation, we detect a small positive shift in radial direction of $\Delta \alpha _{\parallel }= 0.35~{{\ \rm per\ cent}}$ at the 3$\sigma$ level and virtually no shift in the transverse direction, $\alpha _\perp$. From the cross-correlation, we see a similar shift to $\Delta \alpha _\parallel$, albeit with larger error bars, and a small negative shift, $\Delta \alpha _{\perp }=\sim$0.25 per cent, at the 2$\sigma$ level. We also make a connection with the Ly $\alpha$ forest effective field theory (EFT) framework and find that the one-loop EFT power spectrum yields unbiased measurements of the BAO shift parameters in radial and transverse direction for Ly $\alpha$ auto- and the Ly $\alpha$-quasar cross-correlation measurements. When using the one-loop EFT framework, we find that we can recover the BAO parameters without a shift, which has important implications for future Ly $\alpha$ forest analyses based on EFT. This work paves the way for novel full-shape analyses of the currently observing DESI and future surveys such as the PFS, WEAVE-QSO, and 4MOST.

Hadzhiyska, Boryana↗

Towards modelling AR Sco: calibration – reproducing high-energy pulsar emission and testing convergence to Aristotelian electrodynamics

In recent years, kinetic simulations have been crucial to further our understanding of pulsar electrodynamics. Yet, due to the large-scale separation between the gyro-period and the stellar rotation period, resolving the particle gyration has been computationally unfeasible for realistic pulsar parameters. The main aim of this work is comparing our gyro-phase-resolved model with a gyro-centric pulsar model, where our model solves the general equations of motion with included radiation reaction using a higher order numerical solver with adaptive time-steps. Specifically, we aim to (i) reproduce a pulsar’s high-energy emission maps, namely one with 10 per cent of the surface B-field strength of Vela, and the spectra produced by an independent gyro-centric pulsar emission model; and (ii) test convergence of these results to the radiation-reaction limit of Aristotelian electrodynamics. (iii) Additionally, we identify the effect that a large $E_{\parallel }$-field has on the trajectories and radiation calculations. We find that we can reproduce the curvature radiation emission maps and spectra well, using 10 per cent field strengths of the Vela pulsar and injecting our particles at a higher altitude in the magnetosphere. Using sufficiently large $E_{\parallel }$-fields, our numeric results converge to the analytic radiation-reaction limit trajectories. Additionally, we illustrate the importance of accounting for the $\mathbf {E}\times \mathbf {B}$-drift in the particle trajectories and radiation calculations, validating the Harding and collaborators’ model approach. Lastly, we found that our model deals very well with the high-radiation-reaction and high-field regimes present in pulsars.

79 ASTRONOMY AND ASTROPHYSICS↗

New scaling and nuclear structure aspects in heavy-ion fusion reactions

Three new behaviors have been found in comparisons of fusion cross sections for different collision systems. root (1) Replacing the energy E with a scaling one, E scal = (E-V g )/($\sqrt{2}$W g ), is successful for washing out the Coulomb interaction in the spectra of fusion cross sections, where V g and W g are barrier height and width of the single-Gaussian barrier distribution model. (2) In a representation of σE vs the scaling energy, E scal , all data sets display in parallel. Here, the ratio for sigma E from any two fusion systems over the whole range is a constant value. That behavior is also studied in another representation, in which the data sets display as parallel horizontal lines for any heavy-ion fusion system. (3) The constant ratio value is the ratio of parameter products, $R^2_gW_g$, of the two systems; where R g is the barrier radius obtained in the single-Gaussian barrier distribution model. Moreover, when comparing neighboring collision systems at the same E scal , the ratio of sigma is near a constant value within a few percent over the whole range. Thus a quantitative comparison for the fusion enhancement for neighboring systems is developed. The present finding could be beneficial for predicting unmeasured fusion cross sections.

Jiang, C. L. [Argonne National Laboratory (ANL), A↗

Geometric origin of the intrinsic transverse spin transport in a canted-antiferromagnet/heavy-metal heterostructure

We theoretically study the conditions under which an intrinsic spin Nernst effect–a transverse spin current induced by an applied temperature gradient–can occur in a canted-antiferromagnet insulator, such as LaFeO 3 and other materials of the same family. The spin Nernst effect may provide a microscopic mechanism for an experimentally observed anomalous thermovoltage in LaFeO 3 /Pt heterostructures, where spin is transferred across the insulator/metal interface when a temperature gradient is applied to LaFeO 3 parallel to the interface. We find that LaFeO 3 exhibits an intrinsic spin Nernst effect when inversion symmetry is broken on the axes parallel to both the applied temperature gradient and the direction of spin transport, which can result in a spin injection across the insulator/metal interface. Furthermore, our paper provides a general derivation of a symmetry-breaking-induced spin Nernst effect, which may open a path to engineering a finite spin Nernst effect in systems where it would otherwise not arise.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Zonal magnetic fields regulate nonlinear edge-localized-mode dynamics via self-consistent force balance

Edge-localized modes (ELMs) eject intense bursts of heat and particles that threaten plasma-facing components in fusion reactors. Nonlinear full-torus BOUT++ simulations show that turbulence-driven zonal magnetic fields (ZMFs) play an essential role in nonlinear ELM evolution by maintaining self-consistent force balance. Zonal flows mitigate the initial crash through shear but do not prevent continued radial transport. When ZMFs are self-consistently included, turbulence-driven zonal currents modify the parallel current distribution and magnetic tension and are associated with a reduction of the axisymmetric (𝑛 = 0) perturbed radial force imbalance. This coincides with a transition from convective, bursty propagation to more localized, diffusive transport. Similar behavior is observed across the regimes considered, including both resistive-ballooning and peeling-ballooning cases. Finally, associated signatures, including radial electric field shear and parallel current redistribution, provide experimentally accessible diagnostics for present devices and ITER-relevant conditions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Real-Time GPU-Accelerated OFDR With an Integrated Auxiliary Interferometer

A GPU-accelerated optical frequency domain reflectometry (OFDR) system with an improved integrated auxiliary interferometer is proposed. Unlike conventional approaches that require separate auxiliary interferometers and multiple detection channels, the proposed OFDR system embeds this functionality directly into the signal via an intentional beat component. This enables self-calibration of laser nonlinearity while maintaining a cost-effective hardware configuration. Building on this simplified configuration, the system leverages GPU acceleration with an NVIDIA RTX 4070 Ti to achieve real-time performance, delivering high-throughput signal processing for continuous OFDR interrogation. The signal processing pipeline comprises signal capture, resampling for nonlinearity compensation, and frequency shift computation, all optimized for parallel execution. Hardware benchmarking demonstrates substantial acceleration over CPU implementations, achieving up to a 45× speedup for resampling and frequency shift computations and enabling processing latencies below 30 ms. Thermal response validation is conducted under two complementary scenarios: localized heating using a water bath and cryogenic-temperature conditions using liquid nitrogen. Under localized heating, the system achieves an accuracy of 0.249 °C with a thermal sensitivity of 5.971 GHz/°C, while cryogenic-temperature validation demonstrates a frequency shift response with a sensitivity of 2.383 GHz/°C and an accuracy of 2.04 °C. The high acceleration of the proposed GPU-accelerated OFDR system and its accuracy are achieved by exploiting CUDA-based stride indexing, enabling efficient parallel segmentation and processing of large datasets without additional memory copies. The benchmarking results confirm the robustness, accuracy, and deployability of the proposed OFDR system across a wide temperature range, establishing it as a practical platform for real-time distributed fiber sensing in structurally dynamic environments.

Harb, Salah [Lawrence Berkeley National Laboratory↗

On-Board AC Charging Topology Integrated with Electric Vehicle Motor Drive System

On-board AC charging is a convenient and widely adopted method for recharging electric vehicles (EVs) directly from standard alternating current (AC) power sources. This paper presents a novel topology for AC charging of EVs that utilizes EV 3-phase electric machine windings as the input inductors, thus eliminating the requirement for bulky grid interfacing inductors and resulting in a compact and cost-effective integrated motor drive and charger system. The proposed approach leverages the motor windings and parallel operating half-bridge inverter during the charging process by interconnecting the inverter phases with the motor windings in a mechanically interleaved and electrically paralleled manner. The implementation of this unique and innovative idea, achieved through precise control and arrangement of the motor winding as a series inductor, successfully eliminates the possibility of unintended motion of the electric machine during the charging process.

ADVANCED PROPULSION SYSTEMS↗

Design of High-Power Polyphase PCB Coil Systems for Wireless Power Transfer

Printed circuit board (PCB) coils have been proposed prior for implementation as inductive wireless charging coils to minimize size and cost. Utilization of PCBs can allow for a reduced cost, improved manufacturability, and a wide range of geometric customization options. To circumvent material limitations on insulation and thermal performance, parallel paths can be implemented to divide the current per path accordingly. Within this paper, an unconventional high-power PCB coil is designed employing all possible techniques for wiring with axial and radial parallel paths with equivalent transposition to minimize circulating currents. Design studies are simulated in 3D finite element analysis (FEA) to evaluate imbalance between phases with and without transposition. Two experimental prototype coils were fabricated with measurements for self-inductance and mutual inductance between phases. These measurements were validated to be sufficiently consistent with FEA results. Additionally, a method is proposed for a two-step optimization of coupling coefficient and coil losses.

Lewis, Donovin D.↗

Modified Andronov-Hopf Oscillator-Based Grid-Forming Converter with Emulated Virtual Cable for Enhanced Power Sharing Performance

Nonlinear oscillator-based grid-forming converters offer superior dynamic and steady-state performance, making them an attractive solution for interconnecting renewable resources. This paper proposes a novel modified Andronov-Hopf oscillator to enhance the operating spectrum and facilitate the integration of renewable energy sources. An inner loop controller based on the Lyapunov energy function is implemented to achieve robust stability and performance, while a virtual cable emulation strategy enables seamless parallel operation. Comprehensive modeling and simulation studies validate the effectiveness of the proposed system, demonstrating its capabilities in addressing diverse operating scenarios, including grid faults, renewable energy fluctuations, and parallel operation. The proposed solution exhibits fast transient response, robust stability, and flexible operation, making it a valuable contribution to the field of renewable energy integration. The results of this study can be used to inform the design and implementation of next-generation grid-forming converters, enabling a more sustainable and reliable energy future. Additionally, the proposed system's ability to operate in both grid-connected and islanded modes makes it an ideal candidate for remote and off-grid renewable energy applications. The proposed solution's scalability and modularity also make it suitable for large-scale renewable energy integration. The proposed system is verified through MATLAB/Simulink and PLECS simulations, demonstrating its effectiveness in ensuring robust and efficient operation.

Andronov-Hopf Oscillator (AHO)↗

Interpretable Models for Workflow Differentiation in High-Performance Scientific Networks

Scientific workflows in high-performance networks spawn hundreds of interdependent flows that must be managed collectively—yet existing network classifiers treat each flow in isolation, leading to fragmented QoS decisions and missed interflow patterns. We present a novel traffic classification solution that operates at the workflow level, distinguishing entire filetransfer operations from streaming analytics by capturing how concurrent flows interact and burst together. We introduce a workflow identification window (WIW) that ingests raw packet headers from parallel flows into unified tensors, preserving the spatial-temporal patterns that differentiate scientific workflows. This approach achieves 98.7% accuracy using CNN, LSTM, and hybrid architectures, while maintaining 84% accuracy on production traffic collected a week later—demonstrating robustness to temporal drift. By integrating SHAP and GradCAM explainability, we reveal that early-packet timing patterns and cross-flow correlations drive classification decisions, providing operators with interpretable insights. Our system enables coherent workflow-level QoS enforcement and dynamic bandwidth allocation in scientific networks, eliminating manual per-flow configuration while maintaining classification latency at millisecond level.

Giannakou, Anna [LBL, Berkeley]↗

DC Microgrid Reliability Enhancement with Adaptive Converter Thermal Management

Due to the different device selections, aging levels, and thermal dissipation performance, some converters may take additional thermal stress on switching devices than others in paralleled converter systems, which will reduce system reliability. To address this problem, this paper proposes a power-sharing strategy with adaptive thermal management. First, the temperature-based power loss model and electrical-thermal model are established. Based on that, a high-accuracy IGBT junction temperature estimate considering the power loss-temperature coupling can be achieved. Further, the thermal-sharing for all the switching devices in paralleled converters can be achieved with the proposed adaptive thermal management strategy. The proposed strategy can change the power-sharing ratio adaptively according to the system operation conditions, which will contribute to the system reliability enhancement. The effectiveness of the proposed strategy is verified through PLECS thermal simulation and joint real-time simulation with Dspace and RT-box.

DC microgrid↗

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif↗

Virtual Self-Excited Induction Generator-Based Grid-Forming Inverter Control for Robust Voltage Regulation Under Nonideal Loading

This paper presents a generator-inspired control methodology for grid-forming (GFM) inverters that deliberately emulates a self-excited induction generator so that the inverter can hold its voltage and frequency under difficult loading and severe terminal disturbances across wide voltage and frequency ranges. The design integrates a Lyapunov energy function-based inner loop to provide high bandwidth and strong disturbance rejection, and it complements this with a passivity-based argument that furnishes a coherent large-signal stability guarantee beyond small-signal limits. Analytical insights are developed via the Krylov-Bogoliubov-Mitropolsky averaging method, which reveals an intrinsic resistive droop characteristic; these closed-form relations both explain the observed dynamics and yield simple, decentralized tuning rules. The methodology is validated on a controller-hardware-in-the-loop platform and exercised in real time across balanced, unbalanced, and nonlinear loads, as well as during parallel operation. Across these scenarios, the inverter maintains balanced three-phase voltages, limits harmonic content, settles quickly with well-damped transients, and remains resilient when multiple units operate in parallel. The contributions are a self-excited-machine-inspired GFM controller with enhanced dynamic performance and robustness, a single stability rationale grounded in passivity, closed-form expressions that guide tuning, and comprehensive hardware-in-the-loop validations demonstrating effectiveness and superiority under challenging operating conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Efficient Probabilistic Visualization of Local Divergence of 2D Vector Fields with Independent Gaussian Uncertainty

This work focuses on visualizing uncertainty of local divergence of two-dimensional vector fields. Divergence is one of the fundamental attributes of fluid flows, as it can help domain scientists analyze potential positions of sources (positive divergence) and sinks (negative divergence) in the flow. However, uncertainty inherent in vector field data can lead to erroneous divergence computations, adversely impacting downstream analysis. While Monte Carlo (MC) sampling is a classical approach for estimating divergence uncertainty, it suffers from slow convergence and poor scalability with increasing data size and sample counts. Thus, we present a two-fold contribution that tackles the challenges of slow convergence and limited scalability of the MC approach. (1) We derive a closed-form approach for highly efficient and accurate uncertainty visualization of local divergence, assuming independently Gaussian-distributed vector uncertainties. (2) We further integrate our approach into Viskores, a platform-portable parallel library, to accelerate uncertainty visualization. In our results, we demonstrate significantly enhanced efficiency and accuracy of our serial analytical (speed-up up to 1946×) and parallel Viskores (speed-up up to 19698×) algorithms over the classical serial MC approach. We also demonstrate qualitative improvements of our probabilistic divergence visualizations over traditional mean-field visualization, which disregards uncertainty. We validate the accuracy and efficiency of our methods on wind forecast and ocean simulation datasets.

Ouermi, Timbwaoga [University of Utah]↗

Bias-Modulated ALD of ZnO: Insights into Precursor-Surface Interactions for ZnO Films

Atomic layer deposition (ALD) is widely used to deposit conformal thin films but is often limited in the tunability of the resulting material’s properties. Substrate bias and electric fields alter precursor-surface interactions and provide means to tune material properties. To explore this, we performed zinc oxide (ZnO) ALD using diethylzinc (DEZ) and water on silicon native oxide substrates at 150 °C in a sample holder designed to create a static electrical field by biasing one plate of a parallel plate capacitor-style sample holder during deposition. ZnO films prepared in an electric field/on a biased sample holder were thinner, changed relative crystalline composition, and contained more carbon compared to samples grown in identical sample holders without bias. The thickness was independent of the magnitude of the eletric field between plates, indicating that the primary driver for the change was substrate biasing not the electric field between plates of the parallel plate capacitor-style sample holder. Density functional theory calculations showed enhanced electron migration between dissociatively adsorbed DEZ molecules and the ZnO (002) facet with increasing force from an electric field at the substrate surface, which strengthens the electronic interactions between the surface and the adsorbate. These models offer a compelling explanation for the inhibited growth, changes in the crystallinity, and increase in carbon content of films grown in an electric field/on biased plates.

Jones, Jessica C. (ORCID:0000000174754620)↗