Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Grand challenges of wind energy science – meeting the needs and services of the power system

The share of wind power in power systems is increasing dramatically, and this is happening in parallel with increased penetration of solar photovoltaics, storage, other inverter-based technologies, and electrification of other sectors. Recognising the fundamental objective of power systems, maintaining supply–demand balance reliably at the lowest cost, and integrating all these technologies are significant research challenges that are driving radical changes to planning and operations of power systems globally. In this changing environment, wind power can maximise its long-term value to the power system by balancing the needs it imposes on the power system with its contribution to addressing these needs with services. A needs and services paradigm is adopted here to highlight these research challenges, which should also be guided by a balanced approach, concentrating on its advantages over competitors. The research challenges within the wind technology itself are many and varied, with control and coordination internally being a focal point in parallel with a strong recommendation for a holistic approach targeted at where wind has an advantage over its competitors and in coordination with research into other technologies such as storage, power electronics, and power systems.

17 WIND ENERGY↗

rustpix

rustpix is a high-performance, open-source Rust library with first-class Python bindings (via PyO3) for processing pixel-detector data in neutron imaging. It targets time-stamping detectors such as Timepix3 (TPX3) at ORNL's Spallation Neutron Source (VENUS beamline), where each detected neutron deposits charge across a cluster of pixels within a very high-rate event stream (96M+ hits/sec). rustpix parses TPX3 event data in parallel using memory-mapped I/O, offers four interchangeable clustering algorithms (ABS adjacency-based search, DBSCAN, graph/union-find connected components, and a parallel grid method), and extracts weighted, super-resolved centroids to produce neutron-event lists. A streaming architecture lets it process files larger than available memory. rustpix is distributed as a pip-installable Python package (with NumPy integration), Rust crates, a command-line tool, and an interactive GUI; it writes HDF5, Apache Arrow, and CSV; and it is designed to extend to TPX4 and other detector types. Released as open-source under the MIT License.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

Performance Evaluation of Multi-Vendor Grid-Forming Inverters for Grid-Connected Operation Through Hardware Experimentation

This paper presents the functional performance evaluation tests of multiple (three) commercial grid-forming (GFM) inverters when they operate in parallel with the grid through hardware experiments. The goal of these tests is to explore and benchmark the GFM inverters' functionalities and dynamic response when they are operated in parallel with power grids. Both steady-state (changing the inverter's frequency and voltage droop) and transient (adding step changes to grid's frequency/voltage) tests are performed for each GFM inverter with the same testing circuit and testing protocol. The key findings are summarized as follows: 1) The GFM inverters can be dispatched through frequency and voltage droop intercepts to output the target power; and 1) the GFM inverters automatically respond to system frequency and voltage events to output the needed power; 3) all the GFM inviters shows stability issues when absorbing reactive power from the grid.

frequency droop↗

Importance of Higher Fidelity Model Geometries during Optimization of Critical Experiments

PARADIGM, PARallel Approach of Differential and InteGral Measurements, is a cross-collaborative effort at Los Alamos National Laboratory between nuclear data theorists, differential and integral experimenters, as well as machine learning statisticians to tackle uncertainties in the intermediate region of 239 Pu. In essence, the idea behind PARADIGM is to remove the linear conceptualization of the nuclear data pipeline, shown in Figure 1, and replace it with a far more parallelized approach. The novel approach leverages machine learning to guide which differential measurements and integral experiments will result in the largest decrease in uncertain ties for a nuclide reaction pair in a given energy range. The concept builds off earlier work, EUCLID, which focused on the fast region of 239 Pu. The practical benefit of having evaluation, differential measurement, and integral experiment personnel in collaboration with machine learning is to represent the entire nuclear data in one snapshot. This enable large reduction in the time to deliver improved nuclear data, which using the PARADIGM approach could be done in 3 years. A general outline of PARADIGM and specific topics are available in other papers. The discussion here will pertain directly to the integral experiment design. More specifically, the process of taking a rough design and transforming it into a finalized neutronic model will be discussed.

97 MATHEMATICS AND COMPUTING↗

Investigation of Benchmark $k$ eff Sensitivity and Uncertainty for 239 Pu fission in Specific Energy Ranges

Nuclear data at intermediate energies (from 1 to 100s of keV) are evaluated based on scarce differential data and theory unable to capture physics’ expected structure. There is also a lack of integral data. This is a known deficiency and is challenging to address. Calculated effective multiplication factor, k eff , values for intermediate energy experiments are ~25× further from experiment than for fast energies and are often well outside the experimental uncertainties. The goal of the PARADIGM (PARallel Approach of Differential and InteGral Measurements) project is to significantly re duce the uncertainties of intermediate energy nuclear data for 239 Pu. To this end, PARADIGM simultaneously optimizes experiments at both the Los Alamos Neutron Science Center (LANSCE) and National Criticality Experiments Research Center (NCERC). The combined set of data will inform new intermediate-energy nuclear data. By execution of differential and integral experiments, establishment of new theory, and undertaking nuclear data evaluation in parallel, the timeline to deliver improved nuclear data to users will be reduced significantly that is to three years. For the PARADIGM project, it was decided to optimize an integral experiment for two neutron energy ranges, within the full intermediate energy range. The low energy range goes from 1 to 30 keV, while the higher energy range goes from 30 to 600 keV. This work focuses on nuclear data sensitivities and uncertainties for 239 Pu fission for existing experiments in the International Criticality Safety Benchmark Evaluation Project (ICSBEP). When designing new experiments, it is important to understand what benchmarks currently exist. For a more traditional experiment design (in which a specific application model(s) exists), comparisons would be made between the application model(s) and existing benchmarks. For PARADIGM, there is no specific application model, but instead the specific nuclear data reaction and energy ranges of interest can be explored for existing benchmarks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Streaming Matching and Edge Cover in Practice

Graph algorithms with polynomial space and time requirements often become infeasible for massive graphs with billions of edges or more. State-of-the-art approaches therefore employ approximate serial, parallel, and distributed algorithms to tackle these challenges. However, such approaches require storing the entire graph in memory and thus need access to costly computing resources such as clusters and supercomputers. In this paper, we present practical streaming approaches for solving massive graph problems using limited memory for two prototypical graph problems: maximum weighted matching and minimum weighted edge cover. For matching, we conduct a thorough computational study on two of the semi-streaming algorithms including a recent breakthrough result that achieves a $1/(2+\varepsilon)$-approximation of the weight while using $O( n \log W /\epsilon)$ memory (here $n$ is the number of vertices and $W$ is the maximum edge weight), designed by Paz and Schwartzman [SODA, 2017]. Empirically, we show that the semi-streaming algorithms produce matchings whose weight is close to the best $1/2$-approximate offline algorithm while requiring less time and an order-of-magnitude less memory. For minimum weighted edge cover, we develop three novel semi-streaming algorithms. Two of these algorithms require a single pass through the input graph, require $O(n \log n)$ memory, and provide a 2-approximation guarantee on the objective. We also leverage a relationship between approximate maximum weighted matching and approximate minimum weighted edge cover to develop a two-pass $3/2+\epsilon$-approximate algorithm with the memory requirement of Paz and Schwartzman's semi-streaming matching algorithm. These streaming approaches are compared against the state-of-the-art 3/2-approximate offline algorithm. The semi-streaming matching and the novel edge cover algorithms proposed in this paper can process graphs with several billions of edges in under 30 minutes using 6 GB of memory, which is at least an order of magnitude improvement from the offline (non-streaming) algorithms. For the largest graph, the best alternative offline parallel approximation algorithm (GPA+ROMA) could not finish in three hours even while employing hundreds of processors and 1 TB of memory. We also demonstrate an application of the semi-streaming algorithm by computing a matching using linearly bounded memory on item intersection graphs derived from three machine learning datasets, whereas the existing offline algorithms could not complete on one of these datasets since their memory requirements exceeded 1TB.

Ferdous, S M.↗

Packaging a 650V/400A GaN Half-bridge Power Module with Ultra-low Parasitics for Electric Vehicle Drive Applications

This paper proposes a compact and efficient half-bridge power module with three 650 V / 150 A GaN dies in parallel. The power module incorporates a main power printed circuit board (PCB), an interface PCB, and a flex PCB to achieve low parasitics in both power loop and gate-side connection, resolving the issue of high parasitics typically encountered with wire bonding in high-current applications. Additionally, the interface PCB decouples the design constraints between the power loop and the gate loops. The proposed design is optimized with a vertical loop configuration to reduce power loop inductance through magnetic flux cancellation. Finite element analysis indicates that the power loop inductance is 0.58 nH at 100 MHz, while the maximum die junction temperature reaches 131 °C under an ambient temperature of 65 °C and a load current of 385 A. The proposed multi-piece PCB structure reduces the inductance of the drive circuit to minimize EMI and to mitigate false triggering. At the same time, it reduces impedance mismatches across different driver circuits, thereby achieving dynamic current sharing in multi-chip parallel configurations. Under simulation conditions of 400 V / 385 A, the current imbalance among chips was limited to 5 A. A 400 V / 385 A double-pulse test was conducted to experimentally validate the performance of the proposed power module.

30 DIRECT ENERGY CONVERSION↗

Multi-resonant switched capacitor power converter architecture

A switched-capacitor (SC) network in an SC converter is controlled to operate at varying resonant modes to achieve high conversion ratio efficiency, at a low circuit component count. These power converters are suited to numerous application areas including improving energy efficiency of data centers. A family of resonant switched capacitor (SC) converters with multiple operating phases are presented “Multi-Resonant SC Converters”. Described in detail are an 8-to-1 Multi-Resonant-Doubler (MRD) converter and a 6-to-1 Cascaded Series-Parallel (CaSP). The topology of these converters make them amenable to combining like units in parallel toward reaching higher power levels.

Ye, Zichao↗

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip↗

How Cloud is Accelerating Research at NREL

This presentation coincides with AWS's announcement of their new Parallel Computing Service (PCS) which allows for easy creation of HPC-style clusters in their AWS cloud computing platform. I helped them beta test this service before it was made generally available in August. AWS asked if we would be interested in discussing our experience with the PCS service, and our experience with HPC workloads in the cloud in general, so this slideshow discusses a brief history of scientific computing at NREL and shares a bit of our experiences and approach to utilizing cloud services for HPC-style workloads.

97 MATHEMATICS AND COMPUTING↗

HARD: A performance portable radiation hydrodynamics code based on FleCSI framework

Hydrodynamics And Radiation Diffusion (HARD) is an open-source application for high-performance simulations of compressible hydrodynamics with radiation-diffusion coupling. Built on the FleCSI (Bergen et al., 2021 [1]) (Flexible Computational Science Infrastructure) framework, HARD expresses its computational units as tasks whose execution can be orchestrated by multiple back-end runtimes, including Legion (Bauer et al., 2012 [2]), MPI (Forum, 1994 [3]), and HPX (Kaiser et al., 2020 [4]). Node-level parallelism is handled through Kokkos (Edwards et al., 2014 [5]), providing a single-source, portable code base that runs efficiently on laptops, small homogeneous clusters, and the largest heterogeneous supercomputers currently available. To ensure scientific reliability, HARD includes a regression test suite that automatically reproduces canonical verification problems such as the Sod and LeBlanc shock tubes, and the Sedov blast wave, comparing numerical solutions against known analytical results. The project is distributed under an OSI-approved license, hosted on GitHub, and accompanied by reproducible build scripts and continuous integration workflows. This combination of performance portability, verification infrastructure, and community-focused development makes HARD a sustainable platform for advancing radiation hydrodynamics research across multiple domains.

97 MATHEMATICS AND COMPUTING↗

DL_POLY Quantum 2.0: A modular general-purpose software for advanced path integral simulations

DL_POLY Quantum 2.0, a vastly expanded software based on DL_POLY Classic 1.10, is a highly parallelized computational suite written in FORTRAN77 with a modular structure for incorporating nuclear quantum effects into large-scale/long-time molecular dynamics simulations. This is achieved by presenting users with a wide selection of state-of-the-art dynamics methods that utilize the isomorphism between a classical ring polymer and Feynman’s path integral formalism of quantum mechanics. Here, the flexible and user-friendly input/output handling system allows the control of methodology, integration schemes, and thermostatting. DL_POLY Quantum is equipped with a module specifically assigned for calculating correlation functions and printing out the values for sought-after quantities, such as dipole moments and center-of-mass velocities, with packaged tools for calculating infrared absorption spectra and diffusion coefficients.

36 MATERIALS SCIENCE↗

Performance on HPC Platforms Is Possible Without C++

Computing at large scales has become extremely challenging due to increasing heterogeneity in both hardware and software. More and more scientific workflows must tackle a range of scales and use machine learning and AI intertwined with more traditional numerical modeling methods, placing more demands on computational platforms. These constraints indicate a need to fundamentally rethink the way computational science is done and the tools that are needed to enable these complex workflows. The current set of C++-based solutions may not suffice, and relying exclusively upon C++ may not be the best option, especially because several newer languages and boutique solutions offer more robust design features to tackle the challenges of heterogeneity. In June 2023, we held a mini symposium that explored the use of newer languages and heterogeneity solutions that are not tied to C++ and that offer options beyond template metaprogramming and Parallel. For for performance and portability. In conclusion, we describe some of the presentations and discussion from the mini symposium in this article.

97 MATHEMATICS AND COMPUTING↗

A VSC-HVDC-Assisted Black-Start Strategy in Bulk Power Systems a Case Study in San Diego

With the worldwide growth in deploying high-voltage direct current (HVDC) transmission systems, their ability to facilitate black-start (BS) restoration has been a research topic of interest. In this context, voltage source converter (VSC)-HVDC is regarded as a BS resource, and this paper proposes a VSC-HVDC-assisted parallel BS restoration strategy in bulk power systems. The proposed strategy consists of two stages: 1) determination of the VSC and generator startup sequence and 2) load restoration simulation. In the first stage, the entire blackout system is sectionalized into multiple subsystems. Each subsystem includes a VSC-HVDC station or traditional BS unit, it independently determines its generator startup timeline and the energization timelines for buses and lines. The second stage involves load restoration, conceptualized as a modified unit commitment problem, with the timelines established in the first stage work as critical inputs. The proposed BS restoration strategy is tested on the San Diego power system to simulate the 2011 Southwest blackout. The simulation results validate the effectiveness of using VSC-HVDC links as a BS resource which not only speeds up the restoration process but also reduces both energy and economic losses.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Optimal Control Strategy With Efficiency and Reliability Improvement for Offshore DC Microgrids

Offshore microgrids, due to their remote location and lack of external energy support, face significant challenges in wide-range load operation and maintenance. Consequently, efficiency and reliability are critical concerns for converters in offshore dc microgrids. This article presents an optimal control strategy aimed at enhancing both efficiency and reliability. A normalized nonlinear relationship between power loss and thermal stress of a paralleled converter is first established. Based on this, a dual-objective optimization function with an active weight function as well as a system overall performance index is established. The active weight function dynamically adjusts the control priority based on converter efficiency and switching device thermal stress. Then, the optimal power-sharing strategy is derived by the Lagrange multiplier method with the proposed optimal function. Additionally, to accommodate a wide load range, an optimal selection strategy for operating converter combinations is proposed, requiring only low-bandwidth communication. Experiment verification is given to validate the effectiveness of the proposed control strategy. The experiment results demonstrate that the proposed control strategy can improve the overall performance of offshore microgrids by optimizing efficiency and reliability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Virtual Time III, Part 3: Throttling and Message Cancellation

This is Part 3 of a trio of papers that unify in a natural way the two historically distinct parallel discrete event synchronization paradigms, optimistic and conservative, combining the best properties of both into a single framework called Unified Virtual Time (UVT). In this part, we survey the synchronization effects that can be achieved by restricting to corner cases the relationships permitted among the control variables, GVT, CVT, TVT, and LVT, which were defined in Part 1. Here we also survey various throttling policies from the literature and describe how they can be implemented in UVT by controlling the value of TVT, including policies that can take advantage of rollback in addition to LP blocking. A significant result is a new category of efficient and higher precision throttling algorithms for optimistic execution that are based on optimistic lookahead, defined in a way that is symmetric to what we now call the conservative lookahead information that is traditionally used for conservative synchronization. Finally, we present a novel algorithm allowing the choice between lazy and aggressive cancellation to be made on a message-by-message basis using either external logic expressed in the model code, or policy code internal to the simulator, or a mixture of both.

throttling↗

Direct numerical simulations for hybrid rocket boundary layers: Performance modeling and scaling

This paper presents a comprehensive performance and scaling analysis of direct numerical simulations for reacting boundary layers, focusing on slab burner configurations. Using a PETSc-based finite volume CFD framework, the study evaluates the scalability and computational cost of flow, chemistry, and radiation evaluations across 2D and 3D simulations. Polymethyl methacrylate (PMMA) is the fuel with pure O 2 as the oxidizer, modeled using a detailed chemical kinetics mechanism with 113 species and 660 reactions. A ray-tracing-based radiation solver, designed for distributed memory applications, is implemented to model radiation heat transfer. Parallel scalability is analyzed for the coupled flow, chemistry, and radiation heat transfer processes. Weak and strong scaling studies are conducted on up to 15,000 computational ranks, revealing robust performance when flow cells exceed 200 per rank. Chemistry evaluations dominate the computational cost in large 3D simulations, accounting for approximately 40% of the total runtime, while flow processes contribute around 35%, and radiation solver contributions remain below 10% due to reduced evaluation frequencies. GPU accelerated chemistry evaluation, implemented with Zero-RK, demonstrates significant promise, achieving up to a 4x speedup for workloads exceeding 30,000 cells per GPU. However, diminishing returns are observed for smaller workloads due to CPU-GPU communication overhead. This study identifies key challenges, including memory bottlenecks and the effects of domain partitioning on flow scalability, while highlighting the potential of GPU-accelerated chemistry to reduce computational costs. In conclusion, these findings provide realizable run configurations for 2D, 3D, and GPU-accelerated cases, offering insights for optimizing reactive flow solvers.

CFD Scalability↗

CROCUS Dual Polarization Ceilometer Data at Northeastern Illinois University Rooftop

This dataset is from the Department of Energy Office of Science funded project, CROCUS Urban Integrated Field Laboratory (https://crocus-urban.org/). The dual-polarization ceilometer (Vaisala CL61) is an autonomous lidar system operating at 910 nm wavelength, providing valuable measurements for understanding atmospheric boundary layer evolution, air quality, and cloud-aerosol interactions. The CL61 measures the backscattered signal intensity alternating between parallel- and cross-polarization signals. The unique depolarization measurement capability improves discrimination between different particle types, such as liquid droplets, ice crystals, and aerosols. The depolarization is highly dependent on the scatterer shape and orientation (spherical vs non-spherical particles), with the linear depolarization ratio providing a measure of dominant backscatter signal component from atmospheric particles at various heights, essentially allowing discrimination between liquid and solid particles. With its efficient optical system, CL61’s improved signal-to-noise ratio compared to traditional ceilometers allows studying detailed vertical profiles of aerosols and clouds up to 15 km height.Datasets are stored in a netCDF data format, and we we encourage users to make use of the associated toolkits available from Unidata (https://www.unidata.ucar.edu/software/netcdf/), Project Pythia (https://foundations.projectpythia.org/core/data-formats/netcdf-cf.html), and our “Instrument Cookbooks” (https://crocus-urban.github.io/instrument-cookbooks) for more information on how to process the metadata-rich datasets.

54 ENVIRONMENTAL SCIENCES↗