Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Investigation of Benchmark $k$ eff Sensitivity and Uncertainty for 239 Pu fission in Specific Energy Ranges

Nuclear data at intermediate energies (from 1 to 100s of keV) are evaluated based on scarce differential data and theory unable to capture physics’ expected structure. There is also a lack of integral data. This is a known deficiency and is challenging to address. Calculated effective multiplication factor, k eff , values for intermediate energy experiments are ~25× further from experiment than for fast energies and are often well outside the experimental uncertainties. The goal of the PARADIGM (PARallel Approach of Differential and InteGral Measurements) project is to significantly re duce the uncertainties of intermediate energy nuclear data for 239 Pu. To this end, PARADIGM simultaneously optimizes experiments at both the Los Alamos Neutron Science Center (LANSCE) and National Criticality Experiments Research Center (NCERC). The combined set of data will inform new intermediate-energy nuclear data. By execution of differential and integral experiments, establishment of new theory, and undertaking nuclear data evaluation in parallel, the timeline to deliver improved nuclear data to users will be reduced significantly that is to three years. For the PARADIGM project, it was decided to optimize an integral experiment for two neutron energy ranges, within the full intermediate energy range. The low energy range goes from 1 to 30 keV, while the higher energy range goes from 30 to 600 keV. This work focuses on nuclear data sensitivities and uncertainties for 239 Pu fission for existing experiments in the International Criticality Safety Benchmark Evaluation Project (ICSBEP). When designing new experiments, it is important to understand what benchmarks currently exist. For a more traditional experiment design (in which a specific application model(s) exists), comparisons would be made between the application model(s) and existing benchmarks. For PARADIGM, there is no specific application model, but instead the specific nuclear data reaction and energy ranges of interest can be explored for existing benchmarks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Streaming Matching and Edge Cover in Practice

Graph algorithms with polynomial space and time requirements often become infeasible for massive graphs with billions of edges or more. State-of-the-art approaches therefore employ approximate serial, parallel, and distributed algorithms to tackle these challenges. However, such approaches require storing the entire graph in memory and thus need access to costly computing resources such as clusters and supercomputers. In this paper, we present practical streaming approaches for solving massive graph problems using limited memory for two prototypical graph problems: maximum weighted matching and minimum weighted edge cover. For matching, we conduct a thorough computational study on two of the semi-streaming algorithms including a recent breakthrough result that achieves a $1/(2+\varepsilon)$-approximation of the weight while using $O( n \log W /\epsilon)$ memory (here $n$ is the number of vertices and $W$ is the maximum edge weight), designed by Paz and Schwartzman [SODA, 2017]. Empirically, we show that the semi-streaming algorithms produce matchings whose weight is close to the best $1/2$-approximate offline algorithm while requiring less time and an order-of-magnitude less memory. For minimum weighted edge cover, we develop three novel semi-streaming algorithms. Two of these algorithms require a single pass through the input graph, require $O(n \log n)$ memory, and provide a 2-approximation guarantee on the objective. We also leverage a relationship between approximate maximum weighted matching and approximate minimum weighted edge cover to develop a two-pass $3/2+\epsilon$-approximate algorithm with the memory requirement of Paz and Schwartzman's semi-streaming matching algorithm. These streaming approaches are compared against the state-of-the-art 3/2-approximate offline algorithm. The semi-streaming matching and the novel edge cover algorithms proposed in this paper can process graphs with several billions of edges in under 30 minutes using 6 GB of memory, which is at least an order of magnitude improvement from the offline (non-streaming) algorithms. For the largest graph, the best alternative offline parallel approximation algorithm (GPA+ROMA) could not finish in three hours even while employing hundreds of processors and 1 TB of memory. We also demonstrate an application of the semi-streaming algorithm by computing a matching using linearly bounded memory on item intersection graphs derived from three machine learning datasets, whereas the existing offline algorithms could not complete on one of these datasets since their memory requirements exceeded 1TB.

Ferdous, S M.↗

Packaging a 650V/400A GaN Half-bridge Power Module with Ultra-low Parasitics for Electric Vehicle Drive Applications

This paper proposes a compact and efficient half-bridge power module with three 650 V / 150 A GaN dies in parallel. The power module incorporates a main power printed circuit board (PCB), an interface PCB, and a flex PCB to achieve low parasitics in both power loop and gate-side connection, resolving the issue of high parasitics typically encountered with wire bonding in high-current applications. Additionally, the interface PCB decouples the design constraints between the power loop and the gate loops. The proposed design is optimized with a vertical loop configuration to reduce power loop inductance through magnetic flux cancellation. Finite element analysis indicates that the power loop inductance is 0.58 nH at 100 MHz, while the maximum die junction temperature reaches 131 °C under an ambient temperature of 65 °C and a load current of 385 A. The proposed multi-piece PCB structure reduces the inductance of the drive circuit to minimize EMI and to mitigate false triggering. At the same time, it reduces impedance mismatches across different driver circuits, thereby achieving dynamic current sharing in multi-chip parallel configurations. Under simulation conditions of 400 V / 385 A, the current imbalance among chips was limited to 5 A. A 400 V / 385 A double-pulse test was conducted to experimentally validate the performance of the proposed power module.

30 DIRECT ENERGY CONVERSION↗

Multi-resonant switched capacitor power converter architecture

A switched-capacitor (SC) network in an SC converter is controlled to operate at varying resonant modes to achieve high conversion ratio efficiency, at a low circuit component count. These power converters are suited to numerous application areas including improving energy efficiency of data centers. A family of resonant switched capacitor (SC) converters with multiple operating phases are presented “Multi-Resonant SC Converters”. Described in detail are an 8-to-1 Multi-Resonant-Doubler (MRD) converter and a 6-to-1 Cascaded Series-Parallel (CaSP). The topology of these converters make them amenable to combining like units in parallel toward reaching higher power levels.

Ye, Zichao↗

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip↗

How Cloud is Accelerating Research at NREL

This presentation coincides with AWS's announcement of their new Parallel Computing Service (PCS) which allows for easy creation of HPC-style clusters in their AWS cloud computing platform. I helped them beta test this service before it was made generally available in August. AWS asked if we would be interested in discussing our experience with the PCS service, and our experience with HPC workloads in the cloud in general, so this slideshow discusses a brief history of scientific computing at NREL and shares a bit of our experiences and approach to utilizing cloud services for HPC-style workloads.

97 MATHEMATICS AND COMPUTING↗

HARD: A performance portable radiation hydrodynamics code based on FleCSI framework

Hydrodynamics And Radiation Diffusion (HARD) is an open-source application for high-performance simulations of compressible hydrodynamics with radiation-diffusion coupling. Built on the FleCSI (Bergen et al., 2021 [1]) (Flexible Computational Science Infrastructure) framework, HARD expresses its computational units as tasks whose execution can be orchestrated by multiple back-end runtimes, including Legion (Bauer et al., 2012 [2]), MPI (Forum, 1994 [3]), and HPX (Kaiser et al., 2020 [4]). Node-level parallelism is handled through Kokkos (Edwards et al., 2014 [5]), providing a single-source, portable code base that runs efficiently on laptops, small homogeneous clusters, and the largest heterogeneous supercomputers currently available. To ensure scientific reliability, HARD includes a regression test suite that automatically reproduces canonical verification problems such as the Sod and LeBlanc shock tubes, and the Sedov blast wave, comparing numerical solutions against known analytical results. The project is distributed under an OSI-approved license, hosted on GitHub, and accompanied by reproducible build scripts and continuous integration workflows. This combination of performance portability, verification infrastructure, and community-focused development makes HARD a sustainable platform for advancing radiation hydrodynamics research across multiple domains.

97 MATHEMATICS AND COMPUTING↗

A VSC-HVDC-Assisted Black-Start Strategy in Bulk Power Systems a Case Study in San Diego

With the worldwide growth in deploying high-voltage direct current (HVDC) transmission systems, their ability to facilitate black-start (BS) restoration has been a research topic of interest. In this context, voltage source converter (VSC)-HVDC is regarded as a BS resource, and this paper proposes a VSC-HVDC-assisted parallel BS restoration strategy in bulk power systems. The proposed strategy consists of two stages: 1) determination of the VSC and generator startup sequence and 2) load restoration simulation. In the first stage, the entire blackout system is sectionalized into multiple subsystems. Each subsystem includes a VSC-HVDC station or traditional BS unit, it independently determines its generator startup timeline and the energization timelines for buses and lines. The second stage involves load restoration, conceptualized as a modified unit commitment problem, with the timelines established in the first stage work as critical inputs. The proposed BS restoration strategy is tested on the San Diego power system to simulate the 2011 Southwest blackout. The simulation results validate the effectiveness of using VSC-HVDC links as a BS resource which not only speeds up the restoration process but also reduces both energy and economic losses.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Optimal Control Strategy With Efficiency and Reliability Improvement for Offshore DC Microgrids

Offshore microgrids, due to their remote location and lack of external energy support, face significant challenges in wide-range load operation and maintenance. Consequently, efficiency and reliability are critical concerns for converters in offshore dc microgrids. This article presents an optimal control strategy aimed at enhancing both efficiency and reliability. A normalized nonlinear relationship between power loss and thermal stress of a paralleled converter is first established. Based on this, a dual-objective optimization function with an active weight function as well as a system overall performance index is established. The active weight function dynamically adjusts the control priority based on converter efficiency and switching device thermal stress. Then, the optimal power-sharing strategy is derived by the Lagrange multiplier method with the proposed optimal function. Additionally, to accommodate a wide load range, an optimal selection strategy for operating converter combinations is proposed, requiring only low-bandwidth communication. Experiment verification is given to validate the effectiveness of the proposed control strategy. The experiment results demonstrate that the proposed control strategy can improve the overall performance of offshore microgrids by optimizing efficiency and reliability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Virtual Time III, Part 3: Throttling and Message Cancellation

This is Part 3 of a trio of papers that unify in a natural way the two historically distinct parallel discrete event synchronization paradigms, optimistic and conservative, combining the best properties of both into a single framework called Unified Virtual Time (UVT). In this part, we survey the synchronization effects that can be achieved by restricting to corner cases the relationships permitted among the control variables, GVT, CVT, TVT, and LVT, which were defined in Part 1. Here we also survey various throttling policies from the literature and describe how they can be implemented in UVT by controlling the value of TVT, including policies that can take advantage of rollback in addition to LP blocking. A significant result is a new category of efficient and higher precision throttling algorithms for optimistic execution that are based on optimistic lookahead, defined in a way that is symmetric to what we now call the conservative lookahead information that is traditionally used for conservative synchronization. Finally, we present a novel algorithm allowing the choice between lazy and aggressive cancellation to be made on a message-by-message basis using either external logic expressed in the model code, or policy code internal to the simulator, or a mixture of both.

throttling↗

Direct numerical simulations for hybrid rocket boundary layers: Performance modeling and scaling

This paper presents a comprehensive performance and scaling analysis of direct numerical simulations for reacting boundary layers, focusing on slab burner configurations. Using a PETSc-based finite volume CFD framework, the study evaluates the scalability and computational cost of flow, chemistry, and radiation evaluations across 2D and 3D simulations. Polymethyl methacrylate (PMMA) is the fuel with pure O 2 as the oxidizer, modeled using a detailed chemical kinetics mechanism with 113 species and 660 reactions. A ray-tracing-based radiation solver, designed for distributed memory applications, is implemented to model radiation heat transfer. Parallel scalability is analyzed for the coupled flow, chemistry, and radiation heat transfer processes. Weak and strong scaling studies are conducted on up to 15,000 computational ranks, revealing robust performance when flow cells exceed 200 per rank. Chemistry evaluations dominate the computational cost in large 3D simulations, accounting for approximately 40% of the total runtime, while flow processes contribute around 35%, and radiation solver contributions remain below 10% due to reduced evaluation frequencies. GPU accelerated chemistry evaluation, implemented with Zero-RK, demonstrates significant promise, achieving up to a 4x speedup for workloads exceeding 30,000 cells per GPU. However, diminishing returns are observed for smaller workloads due to CPU-GPU communication overhead. This study identifies key challenges, including memory bottlenecks and the effects of domain partitioning on flow scalability, while highlighting the potential of GPU-accelerated chemistry to reduce computational costs. In conclusion, these findings provide realizable run configurations for 2D, 3D, and GPU-accelerated cases, offering insights for optimizing reactive flow solvers.

CFD Scalability↗

CROCUS Dual Polarization Ceilometer Data at Northeastern Illinois University Rooftop

This dataset is from the Department of Energy Office of Science funded project, CROCUS Urban Integrated Field Laboratory (https://crocus-urban.org/). The dual-polarization ceilometer (Vaisala CL61) is an autonomous lidar system operating at 910 nm wavelength, providing valuable measurements for understanding atmospheric boundary layer evolution, air quality, and cloud-aerosol interactions. The CL61 measures the backscattered signal intensity alternating between parallel- and cross-polarization signals. The unique depolarization measurement capability improves discrimination between different particle types, such as liquid droplets, ice crystals, and aerosols. The depolarization is highly dependent on the scatterer shape and orientation (spherical vs non-spherical particles), with the linear depolarization ratio providing a measure of dominant backscatter signal component from atmospheric particles at various heights, essentially allowing discrimination between liquid and solid particles. With its efficient optical system, CL61’s improved signal-to-noise ratio compared to traditional ceilometers allows studying detailed vertical profiles of aerosols and clouds up to 15 km height.Datasets are stored in a netCDF data format, and we we encourage users to make use of the associated toolkits available from Unidata (https://www.unidata.ucar.edu/software/netcdf/), Project Pythia (https://foundations.projectpythia.org/core/data-formats/netcdf-cf.html), and our “Instrument Cookbooks” (https://crocus-urban.github.io/instrument-cookbooks) for more information on how to process the metadata-rich datasets.

54 ENVIRONMENTAL SCIENCES↗

HPC for Optimizing Process Parameters to Control Material Evolution in Seamless Induction Hardening of Wind Turbine Main Shaft Bearings

Work proposed in this project focused on understanding the effect of martensitic transformation in the steel on the potential for cracking during seamless induction hardening (SIH) as a function of process conditions to allow the process to optimally scale up. Large-scale, three-dimensional phase-field simulations of martensitic transformation were performed using MEUMAPPS-SS (Microstructure Evolution Using Massively Parallel Phase-field Simulations – Solid State) code developed at Oak Ridge National Laboratory. The simulations were guided by location-specific thermal history generated by experimental measurements of time-temperature history generated at The Timken Company. The simulations were able to capture the morphological evolution of the martensite variants in an Fe-1.0C-1.5Cr steel based on the Nishiyama-Wasserman (NW) orientation relationship. The simulations were also able to quantify the stress-state at the interface between impinging martensite variants. The simulations indicated that the magnitude of the various stress and strain components were dependent on the sizes of the impinging plates with a reduction in these quantities with reduced plate size in agreement with experimental findings. The results obtained from the simulations will be used to guide the optimization of the alloy thermal conditions to eliminate quench cracking during SIH of bearing steels.

99 GENERAL AND MISCELLANEOUS↗

Elevating SolTrace's Capabilities for the Next Generation of Concentrating Solar Analysis

SolTrace is an open-source Monte Carlo ray tracing software developed at NREL. SolTrace can characterize concentrating solar thermal (CST) collector optical performance and is CST technology agnostic. Shown in Fig. 1, SolTrace is a foundational tool in NREL's CST system and component modeling suite. SolTrace's generic surface elements can flexibly model novel collector and receiver designs to predict spatial and temporal flux distributions - critical to understand for CST component design, performance prediction, and system integration. Since its initial development, SolTrace has over 1,650 references on Google Scholar, over 9,800 downloads since 2017, and has served the CST research and development community as a benchmark of 3rd party verification. SolTrace provides users with many options for defining surface shape and boundaries. However, SolTrace provides limited documentation which can result in a steep learning curve for new users. Additionally, SolTrace lacks the computational performance required to evaluate optical performance of a CST system over the course of a year and/or iteratively over design parameters in a timely manner. To address this, we are working towards a new release of SolTrace that enables increased computational throughput by implementing ray tracing acceleration structures and enabling GPU parallelization. Additionally, we are working to improve SolTrace's usability, accessibility, and maintainability by (1) automating solar position time-dependent simulation processes, (2) creating general CST collector templates of grouped elements, (3) updating the user interface to better visualize model inputs and outputs, and (4) creating a user support network through forums, "how to" videos, and documentation.

14 SOLAR ENERGY↗

Smart Droplets Stabilized by Designer Surfactants: From Biomimicry to Active Motion to Materials Healing

The science and technologies of emulsion droplets have been a long‐term focus of extensive research endeavors for their practical utility across a breadth of industries, including pharmaceutical products, oil recovery processes, and the food sciences. However, with advances in materials chemistry and characterization tools, new emerging areas are arising with a focus on “smart droplets”. The versatility of emulsion droplets across is based on their ability to partition and create isolated systems with properties defined by the liquid–liquid interface, while preparative routes allow manipulation of droplet size, stability, and encapsulated contents. As described in this article, significant efforts are being devoted to creating new types of droplets by “activating” this interface through the incorporation of reactive structures that trigger droplet response to applied or environmental stimuli (e.g., pH, temperature, salt, or external fields). Moreover, parallels between droplets and live cells inspire efforts to conceive systems that resemble biological motifs or that can produce cellular behaviors that imitate biology (e.g., swarming, communication, or motion). Here, the authors highlight recent advances in smart droplets, with emphasis on organic, polymer, and/or particle surfactants that give rise to inter‐droplet communication (via aggregation, fusion, division, or mass transfer), droplet vehicles for controlled delivery, autonomous droplet motion, and tunable emulsion inversion. Especially emphasized is the macromolecular design to produce reactive and functional surfactants, which are crucial to responsive droplet behavior and their underlying mechanisms. More generally, the exquisite interplay between materials science and biology inspires the review of this research area that provides unique opportunities for insight and inspiration into the capabilities of new droplet designs.

36 MATERIALS SCIENCE↗

Formation and Shape Changing of Conductive Helical Ribbons via Deposition of Highly Stressed Films on Mechanically Responsive Substrates

Abstract This work demonstrates that the electrodeposition of highly stressed films on compliant ribbons is a robust process to obtain helical structures with excellent mechanical stability and potentially high thermal and electrical conductance. Electrodeposition on end‐tethered ribbons alters their axial and bending stiffness while imparting mechanical stress to drive the formation of a helix with a microscale diameter and pitch in a controlled and scalable manner. The process generates helices with diameters and pitches between 80 and 200 µm and lengths as large as several millimeters. The approach is amenable to parallel processing a large number of 3D structures on any substrate, including large‐area semiconductor wafers. This phenomenon is explained in terms of the change of stress gradients as material is added. Applications of the fabricated helices include antennas, metamaterials, and slow‐wave structures in frequency ranges not previously attainable.

Chemistry↗

Repairable and Reconfigurable Structured Liquid Circuits

The advance of printed electronics is significantly bolstered by the development of liquid-state electronics that overcome the inherent limitations in flexibility and reconfigurability of solid-state electronics. By integrating the biocompatibility and conductivity of sulfonated polyaniline (S-PANI) and phytic acid (PA) with the reconfigurability of structured liquids, highly conductive all-liquid threads are developed. The dense packing and overlap of PA/S-PANI complexes at an oil/water interface promotes in-plane electron transport, and standard four-point probe measurements of PA/S-PANI interfacial assemblies demonstrate enhanced electrical properties. Notably, the rapid jetting of the ink phase into the matrix phase allows for liquid threads to be printed, enabling the fabrication of large-scale, conductive pathways between two electrodes and liquid circuits. Upon mechanical cleavage of the liquid wires, circuits can be broken, but will easily self-repair using an electric field, making this motif useful in the design of switches as well as restoring conductive pathways in series or in parallel. In conclusion, the demonstrated flexibility and reconfigurability these PA/S-PANI wires possess hold significant promise for their practical use in the design of flexible and adaptive bioelectronics that can be repaired on demand, signifying a transformative step in the evolution of liquid electronic materials.

36 MATERIALS SCIENCE↗

Fractional Skyrmion Tubes in Chiral‐Interfaced 3D Magnetic Nanowires

Magnetic skyrmions are chiral spin textures with rich physics and great potential for unconventional computing. Typically, skyrmions form in bulk crystals with reduced symmetry or ultrathin film multilayers involving heavy metals. Here, the formation of fractional Bloch skyrmion tubes at room temperature is demonstrated by 3D printing ferromagnetic double‐helix nanowires with two regions of opposite chirality. Using X‐ray microscopy and micromagnetic simulations, it is shown that the coexistence of vortex and anti‐parallel spin states induces the formation of fractional skyrmion tubes at zero magnetic fields, minimizing the energy cost of breaking the coupling between geometric and magnetic chirality. Control over zero‐field states is also demonstrated, including pure vortex, or mixed skyrmion‐vortex states, highlighting the magnetic reconfigurability of these 3D nanowires. This work shows how interfacing chiral geometries at the nanoscale can enable advanced forms of topological spintronics.

X-ray microscopy↗