Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49

Direct Measurement of Diffusion Coefficients: Evidence for Diffusive Stochastic Heating in Collisionless Plasmas

Open questions in collisionless plasma dissipation can be addressed using space-based observations in different astrophysical environments, with implications for both astrophysical and laboratory plasma systems. We study a low-𝛽, highly imbalanced, sub-Alfvénic stream observed by Parker Solar Probe (PSP) to identify and distinguish between signatures of stochastic heating (SH) and resonant heating (RH) by parallel ion cyclotron waves (∥-ICWs). Prior work studying this stream [Trevor A. Bowen et al., Stochastic heating in the sub-Alfvénic solar wind, Phys. Rev. Lett. 135, 255201 (2025)] showed that the SH rate, accounting for intermittency, matched the amplitude of the local energy transfer (LET) rate, while the RH rate did not. This comparison relied on a number of assumptions regarding the nature of the diffusive process and the calculation of the LET rate. We introduce a novel technique of inverting the proton guiding center equation to empirically measure velocity-space diffusion coefficients using three-dimensional proton velocity distribution functions, from the ion electrostatic analyzer (the Solar Probe Analyzer for Ions) on PSP. Measured diffusion coefficients are used to determine phase-space heating rates, leading to a calculation of a fully kinetic heating rate independent of assumptions made in prior work. We show that scale-dependent analytic expressions for SH via noncoherent fluctuations match the empirical measurements from PSP data, provided that we account for intermittency in the heating calculation. In contrast, the derived heating rates for SH that accounts for the effects of the helicity barrier and heating rates for RH via ∥-ICWs do not peak in the same region of velocity space as the empirical measurements, nor do they reach the required magnitude. Our approach provides novel methodology to uniquely identify and constrain heating processes in collisionless plasmas and shows evidence of a Fokker-Planck-like diffusive process in the near-Sun solar wind.

Plasma kinetic theory↗

Unconventional magnetotransport and antiferromagnetic order in Eu 2 ⁢InTe 5

Here, we report on the structural characterization and physical properties of Eu 2 ⁢InTe 5 , a distorted square-net motif antiferromagnetic semiconductor. Single crystal x-ray diffraction measurements reveal a distortion of the Te square-net layers such that an orthorhombic supercell is necessary to accurately describe the structure. Magnetization, resistivity, and specific heat measurements confirm antiferromagnetic ordering at 𝑇 𝑁 = 7.3K. Anisotropy in magnetization data suggests the magnetic moments are oriented parallel to the 𝑐 axis. We observe an unconventional transverse magnetoresistance and Hall resistivity in the paramagnetic state (𝑇 ≤ 80K), indicating strong coupling between localized Eu 2+ magnetic moments and conduction electrons. Resistivity and Hall measurements indicate the material is a heavily doped semiconductor, with impurity states related to disorder. Our findings highlight a relatively unexplored Eu-based antiferromagnetic semiconductor with unconventional magnetotransport behavior.

Cook, Matthew S. [Oak Ridge National Laboratory (O↗

Pion and kaon PDFs from lattice QCD via large momentum effective theory and short-distance factorization

In this paper, we present a first-principles lattice-QCD calculation of the unpolarized quark PDF for the pion and the kaon. The lattice data rely on matrix elements calculated for boosted mesons coupled to non-local operators containing a Wilson line. The calculations on this lattice ensemble correspond to two degenerate light, a strange, and a charm quark (𝑁 𝑓 = 2 + 1 + 1), using maximally twisted mass fermions with a clover term. The lattice volume is 32 3 × 64, with a lattice spacing of 0.0934 fm, and a pion mass of 260 MeV. Matrix elements are calculated for hadron boosts of |𝑃 3 | = 0, 0.41, 0.83, 1.25, 1.66, and 2.07 GeV. To match lattice QCD results to their light-cone counterparts, we employ two complementary frameworks: the large-momentum effective theory (LaMET) and the short-distance factorization (SDF). Using these approaches in parallel, we also test the lattice data to identify methodology-driven systematics. Results are presented for the standard quark PDFs, as well as the valence sector. Beyond obtaining the PDFs, we also explore the possibility of extracting information on SU(3) flavor-symmetry-breaking effects. For LaMET, we also parametrize the momentum dependence to obtain the infinite-momentum PDFs. Since the present calculation is performed on a single ensemble at a pion mass of 260 MeV and fixed lattice spacing, the uncertainties reported are statistical only, and systematic uncertainties remain to be addressed in future multi-ensemble studies.

Miller, Joshua [Temple University, Philadelphia, P↗

Positive Neutrino Masses with DESI DR2 via Matter Conversion to Dark Energy

The Dark Energy Spectroscopic Instrument (DESI) is a massively parallel spectroscopic survey on the Mayall telescope at Kitt Peak, which has released measurements of baryon acoustic oscillations determined from over 14 million extragalactic targets. We combine DESI Data Release 2 with CMB datasets to search for evidence of matter conversion to dark energy (DE), focusing on a scenario mediated by stellar collapse to cosmologically coupled black holes (CCBHs). In this physical model, which has the same number of free parameters as Λ⁢CDM, DE production is determined by the cosmic star formation rate density (SFRD), allowing for distinct early- and late-time cosmologies. Using two SFRDs to bracket current observations, we find that the CCBH model: accurately recovers the cosmological expansion history, agrees with early-time baryon abundance measured by BBN, reduces tension with the local distance ladder, and relaxes constraints on the summed neutrino mass ∑𝑚 𝜈 . For these SFRDs, we find a peaked positive ∑𝑚 𝜈 < 0.149 eV (95% confidence) and ∑𝑚 𝜈 = 0.106$^{+0.050}_{−0.069}$ eV, respectively, in good agreement with lower limits from neutrino oscillation experiments. A peak in ∑𝑚 𝜈 > 0 results from late-time baryon consumption in the CCBH scenario and is expected to be a general feature of any model that converts sufficient matter to dark energy during and after reionization.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Combined Influence of Rotation and Scrape-Off Layer Drifts on Recycling Asymmetries in Tokamak Plasmas

Coupled 2D fluid-kinetic simulations of a DIII-D high confinement tokamak plasma show that plasma rotation coupled with drift effects near the plasma edge play a significant role in the creation of the observed poloidal distribution of neutrals. It is observed that including either drift or rotation effects enhances particle flux at the inner target in the case of ion 𝐵×∇𝐵 drift toward the 𝑋-point. However, the particle flux asymmetry is significantly higher with the combination of drifts and rotation than either effect alone. The heightened particle flux asymmetry allows for improved simulation of the strong in-out asymmetry of the Lyman-𝛼 brightness profiles measured in the experiment. Enhancement of radial transport of parallel momentum changes the upstream scrape-off layer flow pattern, increasing the fraction of deuterium flux that reaches the inboard divertor entrance while lowering that which arrives at the outboard. In conclusion, this Letter indicates that by combining drifts, rotation, and viscous coupling, existing boundary plasma models can achieve a satisfactory agreement with experimentally measured neutral asymmetries.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Determination of α lamellae orientation in a β-Ti alloy using electron backscatter diffraction

The spatial orientation of α lamellae in a metastable β-Ti matrix of Timetal LCB (Ti–6.8 Mo–4.5 Fe–1.5 Al in wt%) was examined and the orientation of the hexagonal close-packed α lattice in the α lamella was determined. For this purpose, a combination of methods of small-angle X-ray scattering, scanning electron microscopy and electron backscatter diffraction was used. The habit planes of α laths are close to {111} β , which corresponds to (1$\bar3$20) α in the hexagonal coordinate system of the α phase. The longest α lamella direction lies approximately along one of the $\langle$110$\rangle$ β directions which are parallel to the specific habit plane. Taking into account the average lattice parameters of the β and α phases in aged conditions in Timetal LCB, it was possible to index all main axes and faces of an α lath not only in the cubic coordinate system of the parent β phase but also in the hexagonal system of the α phase.

36 MATERIALS SCIENCE↗

Integrating machine learning interatomic potentials with hybrid reverse Monte Carlo structure refinements in RMCProfile

Structure refinement with reverse Monte Carlo (RMC) is a powerful tool for interpreting experimental diffraction data. To ensure that the under-constrained RMC algorithm yields reasonable results, the hybrid RMC approach applies interatomic potentials to obtain solutions that are both physically sensible and in agreement with experiment. To expand the range of materials that can be studied with hybrid RMC, we have implemented a new interatomic potential constraint in RMCProfile that grants flexibility to apply potentials supported by the Large-scale Atomic/Molecular Massively Parallel Simulator ( LAMMPS ) molecular dynamics code. This includes machine learning interatomic potentials, which provide a pathway to applying hybrid RMC to materials without currently available interatomic potentials. To this end, we present a methodology to use RMC to train machine learning interatomic potentials for hybrid RMC applications.

Cuillier, Paul↗

Solid structure of Li 2 BeF 4 (FLiBe) from room temperature to melting studied by neutron and X-ray diffraction

Molten fluoride salts such as Li 2 BeF 4 (FLiBe) are used in molten salt reactors, fluoride-salt-cooled high-temperature reactors and fusion reactors as a fuel solvent, coolant and/or tritium breeding medium. In engineered systems that use molten salt, solid-state material will be present during melting and freezing scenarios, and therefore the temperature-dependent properties of the solid and solid/liquid phase transition merit investigation. To observe the behavior of the solid state of Li 2 BeF 4 from room temperature to melting, this work used neutron and X-ray diffraction to measure the changes in the lattice parameters and volume of the crystalline unit cell and compared the results with prior low-temperature data for solid Li 2 BeF 4 . From neutron diffraction data it is also possible to identify anisotropy: centimetre-scaled crystals align preferentially with the a axes parallel to the direction of freezing front propagation, and the c axes expand 54% more than the a axes. This work provides the lattice constants as a function of temperature, quantifies the thermal expansion, and determines the equation describing the change in density for solid Li 2 BeF 4 from room temperature to 459°C to be ρ solid (kg m −3 ) = 2182 (3) − 0.115 (2) T (°C) and the volume expansion upon melting to be less than 5%. This density changes depending on molecular weight and enrichment.

36 MATERIALS SCIENCE↗

Discorpy : algorithms and software for camera calibration and correction

Camera or lens-based detector calibration is essential for spatial accuracy in applications like dimensional tomography, optical metrology, and computer vision. Many methods and software exist yet there is still a lack of approaches that achieve both high accuracy and robustness while being easy to use and capable of handling a wide range of distortions. Radial lens distortion is common in high-resolution X-ray detector optics used in parallel-beam tomography at synchrotrons. Achieving sub-pixel accuracy requires calibrating with an optical target image. Although methods for characterizing radial distortion are well established, acquired images often also include perspective distortion and optical center offset. Here, we present our approaches to individually characterize and correct both types of distortion using a single calibration image, implemented in the Discorpy software.

36 MATERIALS SCIENCE↗

Integrating Marine Hydrokinetic and Offshore Wind Energy: A Review of Technologies, Deployment, and Challenges

Together, offshore wind (OSW) and marine hydrokinetic (MHK) technologies have vast potential to expand the world’s access to abundant energy resource. With more than 60 GW of offshore wind energy capacity and 527 MW of ocean energy deployed globally by 2023, there is a significant amount of available resources; however, technical and non-technical challenges prevent the combined large-scale deployment of these technologies. There is still a lack of research that provides a parallel review of both MHK and OSW technologies in order to better understand their synergistic working principles. This paper aimed to address that research gap by presenting a comprehensive side-by-side review of the worldwide technological landscape, global deployment trends, integration strategies, and modeling approaches for MHK and OSW. A particular focus has been given on analyzing existing modeling and simulation techniques, assessing integration and control strategies, and comparing technologies based on water depth. Furthermore, this study provides important insights into the readiness levels of both technologies by highlighting ongoing international projects. By addressing these issues, this review will give researchers and industry stakeholders an outline for assessing the maturity of OSW and MHK systems and facilitating their transition to large-scale, sustainable deployment.

16 - TIDAL AND WAVE POWER↗

Capacitor Design for Self-Resonant Coils for Long-Distance Wireless Power Transfer System

In this paper, an integrated capacitor design is proposed for higher-order resonant tank topologies for self-resonant coils, such as series, parallel, LCC, LLC, etc. The capacitor is one of the large, lossy, and thermally vulnerable components of a high-frequency resonant tank, and designing a high-voltage, thermally stable resonant capacitor can be highly challenging. Designing the extremely high-voltage capacitor as an integral part of the coil reduces the size and complexity of the coil assembly. This paper proposes a low-loss PCB-based high-voltage capacitor design to achieve that target, which can be implemented as an integral part of the coil. The proposed capacitor designs are simulated using Multiphysics FEA and tested experimentally. A 23 kV, 133 pF capacitor prototype was built and tested as part of a 1 kW long-distance wireless charging system. The test results verify the capacitor’s voltage, current, and thermal resiliency performance.

Mohammad, Mostak [ORNL] (ORCID:0000000256388783)↗

Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems

The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globallyconsistent meta-data. In this paper, we propose a novel datastructure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.

Castellana, Vito G. [BATTELLE (PACIFIC NW LAB)] (O↗

Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off Analysis

The scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tool and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the- art methods.

Guo, Luanzheng [BATTELLE (PACIFIC NW LAB)]↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Scaling Ultrahigh-Resolution E3SM Land Model for Leadership-Class Supercomputers

This paper presents advancements in scaling the ultrahigh-resolution E3SM Land Model (uELM) for deployment on leadership-class supercomputers, addressing the increased demand for km-scale Earth system modeling. By focusing on km-scale ELM simulations, we enhance predictive capabilities for climate interactions, facilitating improved responses to climate change impacts on energy systems, agriculture, and water resources. Our approach leverages innovative software architecture optimizations, sophisticated data handling techniques, and advanced parallel processing, achieving strong scalability on two leadership supercomputers (2400 nodes (105,600 cores) on Summit, and 1200 nodes (76,800 cores) on Frontier). Results from extensive scalability assessments on the Summit and Frontier also demonstrate outstanding I/O performance (close to 400 GB/s write throughput) and the model's ability to efficiently handle increasing computational demands. This study not only establishes uELM's capability for high-resolution simulations over vast geographical domains, but also sets a foundation for future Earth system modeling breakthroughs.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

A Versatile Simulated Data Transport Layer for in Situ Workflows Performance Evaluation

In situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes.

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

Game Theoretic Orchestration for Cooperation among Power Distribution System Applications

The evolving transformation with the proliferation of distributed energy resources and advanced metering, necessitates advanced distribution systems to integrate and orchestrate a large number of grid-edge devices while also serving multiple system-level objectives such as resilience, decarbonization, equity and other system mandates. The parallel deployment and control of resources towards achieving diverse objectives may lead to conflicts between applications that want to control overlapping sets of device setpoints, potentially leading to oscillatory behavior and suboptimal performance. This work aims at leveraging game theoretic framework to drive cooperative behavior among competitive applications. The work proposes a weighted-consensus based game design to facilitate conflict resolution through consensus-building iterations for modular platform. Simulation-based evaluation on a sample test system demonstrates the performance the proposed deconfliction strategy in resolving operational conflicts and achieving close-to-optimal trade off among the applications. Results also compare the proposed strategy with a distribution optimization approach and illustrate it effectiveness in diverse apps regardless of their design while also incentivizing apps with flexible design.

Advanced distribution operations, cooperation, app↗

Characterization and Optimization of the Fitting of Quantum Correlation Functions

This case study presents a characterization and optimization of an application code for extracting parton distribution functions from high energy electron-proton scattering data. Profiling this application code reveals that the phase-space density computation accounts for 93% of the overall execution time for a single iteration on a single core. When executing multiple iterations in parallel on a multicore system, the application spends 78% of its overall execution time idling due to load imbalance. We address these issues by first transforming the application code from Python to C++ and then tackling the application load imbalance via a hybrid scheduling strategy that combines dynamic and static scheduling. These techniques result in a 62% reduction in CPU idle time and a 2.46x speedup in overall execution time per node. In addition, the typically enabled power-management mechanisms in supercomputers (e.g., AMD Turbo Core, Intel Turbo Boost, and RAPL) can significantly impact intra-node scalability when more than 50% of the CPU cores are used. This finding underscores the importance of understanding system interactions with power management, as they can adversely impact application performance, and highlights the necessity of intra-node scaling tests to identify performance degradation that inter-node scaling tests might otherwise overlook.

Chuang, Pi-Yueh [Virginia Tech,Dept. of Computer S↗