Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

A worldwide climatology of extreme air masses

Extreme temperature events are among the most damaging weather phenomena. In a warming world, more heat extremes and fewer cold extremes are expected in most regions in the future, a trade-off that warrants further understanding of such events. Here, we track and analyze large, persistent areas of hot and cold extreme temperatures in parallel, relative to the location and time of year, to quantify overall regional exposure to extreme temperatures. To accomplish this, we compare the frequencies, movements, trends, and sources/sinks in each type of extreme air mass, calling them extreme cold or extreme hot air masses (ECAMs/EHAMs). For most land regions, ECAMs occur more often than EHAMs, and ECAMs are more common in each hemisphere’s winter, when cold-air advection is strongest and most widespread. Average movement of ECAMs has a stronger equatorward component in winter than in summer, while movement of EHAMs is eastward all year, with less meridional movement than ECAMs. EHAMs have become more common almost everywhere on land, and the reverse is true for ECAMs, with the strongest trends in the northern hemisphere occurring in autumn, especially in the Arctic. The number of EHAMs around the world is increasing at a higher rate (+ 2.32 events/year) than ECAM numbers are decreasing (-1.27 events/year), causing a net increase in extremes, especially in some midlatitude regions. These results help to document the different processes driving hot and cold extremes, and thus the asymmetric and regionally varying trends in frequency for extreme air masses.

Ryan, James M↗

A rheological model for loose sands with insights from DEM

A rheological model for loose granular media is developed to capture both solid-like and fluid-like responses during shearing. The proposed model is built by following the mathematical structure of an extended Kelvin–Voigt model, where an elastic spring and plastic slider act in parallel to a viscous damper. This arrangement requires the partition of the total stress into rate-independent and rate-dependent stress components. To model the solid-like behavior, a simple frictional plasticity model is adopted without modifications, thus contributing to the rate-independent stress. Instead, the fluid-like or rate-dependent stress is further decomposed into deviatoric and volumetric parts, by proposing a new formulation based on a combination of the μ(I) relation, originally developed under pressure-controlled shear, with a pressure-shear rate relation derived under volume-controlled shear. The proposed formulation allows the model to capture both the increase in the friction coefficient and the enhanced dilation at high shear rates. High-fidelity simulation data, obtained from discrete element method and multiscale modelling, are used to evaluate the performance of the proposed constitutive model. The model provides accurate results under both drained and undrained simple shear paths across a wide range of shear rates. Furthermore, it successfully reproduces at much lower computational cost the flowslide mobility computed through multiscale simulations, which is primarily regulated by the shear rate dependence of the material properties during the dynamic runout stage.

Elasticity↗

Stage-local partitioned two-step runge-kutta methods for large systems of ordinary differential equations

We introduce stage-local partitioned two-step Runge-Kutta methods are an extension of standard two-step Runge-Kutta methods, which are an alternative to the standard additive two-step Runge-Kutta methods currently existing in the literature. Furthermore, these new schemes are designed with an eye towards truly N-partitioned systems and leverage local stage approximations to make several computationally interesting approximations viable. Specifically, the focus on local stage approximations makes possible the construction of truly asynchronous schemes, in the parallel sense, possible. In addition, we show that an implicit-explicit approach to these schemes can lead to methods that require the inversion of only local nonlinear systems.

Applied Dynamical Systems↗

Impact of Forest Canopy Structure on Buoyant Plume Dynamics During Wildland Fires

Heterogeneous forest canopies can generate complex turbulent structures, but in the presence of a fire plume, these interactions are not fully understood. This study investigates the influence of forest canopy heterogeneity on buoyant plume dynamics resulting from surface thermal anomalies representing wildland fires, utilizing Large Eddy Simulation (LES). The Parallelized Large-Eddy Simulation Model (PALM) was employed to simulate six canopy configurations: no canopy, homogeneous canopy, external plume-edge canopy, internal plume-edge canopy, 100 m gap canopy, and 200 m gap canopy. Each configuration was analyzed with and without a static surface heat flux patch of 5000 W ∙ m -2 , resulting in a resting buoyant plume. Simulations were conducted under three crosswind speeds: 0, 5, and 10 m ∙ s -1 . Results show that canopy structure significantly modifies plume behavior, mean flow, and turbulent kinetic energy (TKE) budgets. Plume updraft speed and tilt varied with canopy configuration and crosswind speed. Horizontal pressure gradients associated with plume-atmosphere interaction were modified based on the canopy configuration, resulting in varying crosswind speed reductions at the plume region. Strong momentum absorption was observed above the canopy for the crosswind cases, with the greatest enhancement in the gap canopies. Momentum injection from below the canopy due to the heat source was also observed, resulting in plume structure modulation based on canopy configuration. TKE was found to be the largest in the gap canopy configurations. TKE budget analysis revealed that buoyant production dominated over shear production. At the center of the heat patch, the gap canopy configurations showed enhanced buoyancy within the gap. These results improve our knowledge of fire-canopy-atmosphere interactions that can inform fire models on the impacts of canopy heterogeneity on plume dynamics and ember ejections.

54 ENVIRONMENTAL SCIENCES↗

Hierarchical Phased-Array Antennas Coupled to Al KIDs: A Scalable Architecture for Multi-band Millimeter/Submillimeter Focal Planes

We present the optical characterization of two-scale hierarchical phased-array antenna kinetic inductance detectors (KIDs) for millimeter/submillimeter wavelengths. Our KIDs have a lumped-element architecture with parallel plate capacitors and aluminum inductors. The incoming light is received with a hierarchical phased array of slot dipole antennas, split into 4 frequency bands (between 125 GHz and 365 GHz) with on-chip lumped-element band-pass filters, and routed to different KIDs using microstriplines. Individual pixels detect light for the 3 higher-frequency bands (190–365 GHz), and the signals from four individual pixels are coherently summed to create a larger pixel detecting light for the lowest frequency band (125–175 GHz). The spectral response of the band-pass filters was measured using Fourier transform spectroscopy (FTS), the far-field beam pattern of the phased-array antennas was obtained using an infrared source mounted on a 2-axis translating stage, and the optical efficiency of the KIDs was characterized by observing loads at 294 K and 77 K. We report on the results of these three measurements.

47 OTHER INSTRUMENTATION↗

Randomized Preconditioned Solvers for Strong Constraint 4D-Var Data Assimilation

The Strong Constraint 4D Variational (SC-4DVAR) data assimilation method is widely used in climate and weather applications. SC-4DVAR involves solving a minimization problem to compute the maximum a posteriori estimate, which we tackle using the Gauss-Newton method. The computation of the descent direction is expensive since it involves the solution of a large-scale and potentially ill-conditioned linear system, solved using the preconditioned conjugate gradient (PCG) method. Here, to address this cost, we efficiently construct scalable preconditioners using three different randomization techniques, which all rely on a certain low-rank structure involving the Gauss-Newton Hessian. The proposed techniques come with theoretical guarantees on the condition number, and at the same time, are amenable to parallelization. We also develop an adaptive approach to estimate the sketch size and choose between the reuse or recomputation of the preconditioner. We demonstrate the performance and effectiveness of our methodology on two representative model problems—the Burgers and barotropic vorticity equation—showing a drastic reduction in both the number of PCG iterations and the number of Gauss-Newton Hessian products after including the preconditioner construction cost.

Gauss-Newton↗

Exploratory Investigation of Coal in Nonequilibrium Plasma

Coal is an abundant natural resource and there is motivation to find new uses for it that do not intrinsically involve combustion. One approach is to explore new ways of processing coal, and in this work, we focus on the transformation of coal in a nonequilibrium plasma generated from an equimolar mixture of nitrogen and hydrogen. The outcome of the nonequilibrium plasma reaction is fundamentally different than a thermal control reaction carried out using the same gas composition, pressure, and temperature range. The nonequilibrium plasma produces a gas mixture that is enriched in acetylene and its derivatives. Furthermore, when compared to the thermal control experiment, the solid char byproduct of the nonequilibrium plasma has a very reactive surface and is spontaneously combustible at ambient temperature. Experiments performed to characterize the reaction kinetics of coal in the plasma suggest that the mechanism proceeds through a sequential process by which the coal particle temperature rises to a point where devolatilization can occur, the devolatilization reaction happens, followed by parallel reactions of released organic vapors in the plasma phase and surface activation. In conclusion, the reaction rate appears to be limited by the time it takes for the coal particle temperature to rise, consistent with previous results reported for reactions of coal in thermal plasma.

Acetylene↗

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao↗

Variation of the Passive Film on Compositionally Concentrated Dual-Phase Al 0.3 Cr 0.5 Fe 2 Mn 0.25 Mo 0.15 Ni 1.5 Ti 0.3 and Implications for Corrosion

The passive film on a dual-phase Al 0.3 Cr 0.5 Fe 2 Mn 0.25 Mo 0.15 Ni 1.5 Ti 0.3 FCC + Heusler (L2 1 ) compositionally concentrated alloy formed during extended exposure to an applied potential in the passive range in dilute chloride solution was characterized. Each phase, with its own distinct composition of passivating elements, formed unique passive films separated by a heterophase interface. High-resolution, surface sensitive characterization enabled chemical analysis of the passive film formed over individual phases. The film formed over the L2 1 phase had a higher concentration of Al, Ni, and Ti, while the film formed over FCC phase was of similar thickness but contained comparatively higher Cr, Fe, and Mo concentrations, consistent with the differences in bulk microstructure composition. The passive film was continuous across phase boundaries and the distribution of passivating elements (Al, Cr, and Ti) indicated both phases were independently passivated. Spatially resolved analysis of the surface chemistry of the dual-phase CCA revealed that the cation with the highest composition in passive film formed on the FCC phase was Cr (52.4 at. pct) and for the L2 1 phase was Ti (53.1 at. pct) despite the bulk concentration of each element being below 20 at. pct in their respective phases. Al, Cr, and Ti were enriched in both phases within the passive film relative to their respective bulk compositions. In parallel studies, single-phase alloys with compositions representative of the FCC and L2 1 phases were synthesized to evaluate the corrosion behavior of each phase in isolation. The corrosion behavior of the dual-phase alloy showed passivity evidenced by a pitting potential of 0.615 V SCE in 0.01 M NaCl. The pitting potential and other electrochemical parameters suggested a combination of behaviors of both single-phase samples, suggesting that the global corrosion behavior may be represented by a composite theory applied to phases, their area fractions, and interphase length. However, the interphase in the dual-phase CCA was a local corrosion initiation site and may limit localized corrosion protectiveness. The alloy design implications for optimization of second phase structure and morphology are discussed.

36 MATERIALS SCIENCE↗

The Influence of Residual Stress on Fatigue Crack Growth Rates in Stainless Steel Processed by Different Additive Manufacturing Methods

The properties and microstructure of Type 304L stainless steel produced by two additive manufacturing (AM) methods—directed energy deposition (DED) and powder bed fusion (PBF)—are evaluated and compared. Localized heating and steep temperature gradients of AM processes lead to significant residual stress and distinctive microstructures, which may be process-specific and influence mechanical behavior. Test data show that materials produced by DED and PDF have small differences in tensile strengths but clear differences in residual stress and microstructural features. Measured fatigue crack growth rates (FCGRs) for cracks propagating parallel to and perpendicular to the build directions differ between the two AM materials. To separate the influences of residual stress and microstructure, K-control test procedures with decreasing and constant stress intensity factor ranges are used to measure FCGRs in the near-threshold regime (crack growth rates ≤ 1 × 10 −8 m/cycle). Residual stress is quantified by the residual stress intensity factor, K res , measured by the online crack compliance method. Correcting the FCGR data for differences in K res brings results for specimens of the two AM materials into agreement with each other and with results for wrought specimens, when the latter are corrected for crack closure. Differences in microstructure and tensile strength have an insignificant influence on FCGRs in these tests.

additive manufacturing↗

Expert and operator perspectives on barriers to energy efficiency in data centers

Abstract It was last estimated in 2016 that data centers (DCs) comprise approximately 2% of total US electricity consumption. However, this estimate is currently being updated to account for the massive increase in computing needs due to streaming, cryptocurrency, and artificial intelligence (AI). To prevent energy consumption that tracks with increasing computing needs, it is imperative we identify energy efficiency strategies and investments beyond the low-hanging fruit solutions. In a two-phased research approach, we ask: What non-technical barriers still impede energy efficiency (EE) practices and investments in the data center sector, and what can be done to overcome these barriers? In particular, we are focused on social and organizational barriers to EE. In Phase I, we performed a literature review and found that technical solutions are abundant in the literature, but fail to address the top-down cultural shifts that need to take place in order to adapt new energy efficiency strategies. In Phase II, reported here, we interviewed 16 data center operators/experts to ground-truth our literature findings. Our interview protocols focus on three aspects of DC decision-making: procurement practices, metrics and monitoring, and perceived barriers to energy efficiency. We find that vendors are the key drivers of procurement decisions, advanced efficiency metrics are facility-specific, and there is convergence in the design of advanced facilities due to the heat density of parallelized infrastructure. Our ultimate goals for our research are to design DC decarbonization policies that target organizational structure, empower individual staff, and foster a supportive external market.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗

Optimizing High-Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server

Abstract With machine learning applications now spanning a variety of computational tasks, multi-user shared computing facilities are devoting a rapidly increasing proportion of their resources to such algorithms. Graph neural networks (GNNs), for example, have provided astounding improvements in extracting complex signatures from data and are now widely used in a variety of applications, such as particle jet classification in high energy physics (HEP). However, GNNs also come with an enormous computational penalty that requires the use of GPUs to maintain reasonable throughput. At shared computing facilities, such as those used by physicists at Fermi National Accelerator Laboratory (Fermilab), methodical resource allocation and high throughput at the many-user scale are key to ensuring that resources are being used as efficiently as possible. These facilities, however, primarily provide CPU-only nodes, which proves detrimental to time-to-insight and computational throughput for workflows that include machine learning inference. In this work, we describe how a shared computing facility can use the NVIDIA Triton Inference Server to optimize its resource allocation and computing structure, recovering high throughput while scaling out to multiple users by massively parallelizing their machine learning inference. To demonstrate the effectiveness of this system in a realistic multi-user environment, we use the Fermilab Elastic Analysis Facility augmented with the Triton Inference Server to provide scalable and high-throughput access to a HEP-specific GNN and report on the outcome.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Single-atom materials boosting wearable orthogonal uric acid detection

Abstract Uric acid (UA) is a vital biomarker for the diagnosis and management of various health conditions, including cardiovascular diseases, gout, kidney disorders, metabolic syndrome, and wound healing. Despite significant advances in wearable sensor technology, challenges persist in developing wearable sensors that are capable of maintaining high sensitivity, selectivity, and stability. In this study, we present an epidermal sensing platform enhanced with single-atom materials (SAMs) designed for flexible and orthogonal electrochemical detection of UA. We designed and synthesized an SAM with Fe-N 5 active sites to boost the electrochemical sensing signals, integrating it with laser-engraved graphene (LEG) to fabricate a wearable SAM-based UA patch sensor. This design provides superior UA detection performance compared to sensors based on conventional nanomaterials. In addition, we enhanced the detection accuracy and range by using an orthogonal approach that combines direct oxidation through differential pulse voltammetry (DPV) along with parallel biocatalytic amperometric detection. The resulting SAM-based UA orthogonal sensor patch demonstrated exceptional performance in wearable applications through tests measuring sweat UA levels in subjects before and after consuming a purine-rich diet. Graphical Abstract

Ding, Shichao↗

Enhanced accuracy through ensembling of randomly initialized auto-regressive models for dynamical systems

Computational mechanics simulations using traditional finite element methods (FEM) require prohibitively expensive computational resources for real-time engineering applications, design optimization, and digital twin implementations. While machine learning (ML) surrogate models offer significant computational speedups, autoregressive ML models for time-dependent mechanical systems suffer from error accumulation that compromises long-term prediction reliability - a critical concern for engineering applications where accuracy over extended time horizons is essential for safety and performance assessments. Here, we propose a deep ensemble framework specifically designed to address this challenge in computational mechanics applications, where multiple ML surrogate models with random weight initializations are trained in parallel and their predictions aggregated during inference. This approach leverages statistical diversity to maximize information gain from a fixed set of training data and to mitigate error propagation, while maintaining the computational efficiency that makes ML surrogates attractive for engineering practice. We validate the framework on three representative problems spanning critical areas of computational mechanics: stress field evolution in heterogeneous microstructures under complex loading (relevant to advanced materials design and composite analysis), planetary-scale shallow water dynamics (applicable to environmental and geotechnical engineering), and Gray-Scott reaction-diffusion systems (relevant to mass transport and chemical process engineering). Across all test cases, the ensemble approach demonstrates consistent error reduction of 15-33% compared to individual models. The codes for this work are available on GitHub (https://github.com/Graham-Brady-Research-Group/AutoregressiveEnsemble_SpatioTemporal_Evolution).

autoregressive prediction↗

Progressive Hedging Decomposition for Solutions of Large-Scale Process Family Design Problems

Rapid, wide-scale deployment of green process systems, such as carbon capture or water desalination systems, is essential for combatting climate change. Methods relying on traditional design or modularity fail to capture the benefits of both economies of numbers and economies of scale. We have proposed process family design, which designs a family of processes simultaneously exploiting opportunities for common elements. In previous work, we explored different optimization formulations to solve this problem. In this work, we develop a decomposition approach to tackle larger problems efficiently. We solve a water desalination case study, which is too large to solve within a reasonable timeframe with the discretization formulation. We exploit the block angular structure of the discretization problem to decompose and solve using Progressive Hedging (PH). We use the open-source Python package mpi-sppy to execute PH which allows us to leverage parallelization and a HPC cluster to further improve solution time.

Stinchfield, Georgia↗

Tailoring microstructures with mild magnetic-field processing: A case study of CuNiFe alloys

Combined experimental and computational investigations of the CuNiFe spinodal system confirm that application of a mild magnetic field during thermal treatment alters elemental redistribution and the resulting microstructure, relative to that obtained from zero-field annealing. Spinodal decomposition of a Cu 40 Ni 42 Fe 18 alloy was initiated during thermal treatment at 773 K, conducted either under zero field or modest (60 mT) magnetic f ield conditions for up to 200 h. Periodic (~10 nm) chemical modulations into Cu-rich and NiFe-rich regions were observed under both conditions, with the amplitude and wavelength of the segregated regions increasing with treatment time. However, magnetic field annealing resulted in a more than twofold increase in the amplitude of elemental modulations relative to zero-field conditions – consistent with enhanced diffusional f luxes during spinodal decomposition – while the modulation wavelength remained largely unaffected. These microstructural differences are reflected in various extrinsic magnetic properties. In parallel, first-principles DFT calculations indicate that long-range ferromagnetic order, as induced by an applied magnetic field, substantially alters the strength and nature of atomic interactions, enhancing the thermodynamic instability of the CuNiFe solid solution. Collectively, these results suggest that incorporating a mild (millitesla-level) magnetic field – distinct from the strong (tesla-level) fields commonly used in prior studies – during thermal processing has the potential to deliver enhanced control of microstructures for targeted engineering outcomes.

36 MATERIALS SCIENCE↗

Effect of magneto-mechanical synergism in the process-structure correlation in Fe–C alloys: A phase-field modeling approach

Applied magnetic fields can alter phase equilibria and kinetics in steels; however, quantitatively resolving how magnetic, chemical, and elastic driving forces jointly influence the microstructure remains challenging. We develop a quantitative magneto-mechanically coupled phase-field model for the Fe–C system that couples a CALPHAD-based chemical free energy with demagnetization-field magnetostatics and microelasticity. Here, the model reproduces single- and multi-particle evolution during the α → γ inverse transformation at 1023 K under external fields up to 20 T, including ellipsoidal morphologies observed experimentally at 8 T. Chemically driven growth is isotropic; a magnetic interaction introduces an anisotropic driving force that elongates γ precipitates along the field into ellipsoids, while elastic coherency promotes faceting, yielding elongated cuboidal or “brick-like” particles under combined magneto-elastic coupling. Growth kinetics increase with C content, and decrease with field strength and misfit strain. Multi-particle simulations reveal dipolar interaction-mediated coalescence for field-parallel neighbors and ripening for field-perpendicular neighbors. Incorporating field-dependent diffusivity from experiment slows kinetics as expected; a first-principles-motivated anisotropic diffusivity correction is estimated to be small (<2%). These results establish a process-structure link for magnetically assisted heat treatments of Fe–C alloys and provide guidance for microstructure control via chemo-magneto-mechanical synergism.

Magnetic field↗