Search NASA⌕ Search

SEARCH · Search NASA

Results for “APPROXIMATION METHOD”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Adaptive Patching for High-resolution Image Segmentation with Transformers

Attention-based models are proliferating in the space of image analytics, including segmentation. The standard method of feeding images to transformer encoders is to divide the images into patches and then feed the patches to the model as a linear sequence of tokens. For high-resolution images, e.g. microscopic pathology images, the quadratic compute and memory cost prohibits the use of an attention-based model, if we are to use smaller patch sizes that are favorable in segmentation. The solution is to either use custom complex multi-resolution models or approximate attention schemes. We take inspiration from Adapative Mesh Refinement (AMR) methods in HPC by adaptively patching the images, as a pre-processing step, based on the image details to reduce the number of patches being fed to the model, by orders of magnitude. This method has a negligible overhead, and works seamlessly with any attention-based model, i.e. it is a pre-processing step that can be adopted by any attention-based model without friction. We demonstrate superior segmentation quality over SoTA segmentation models for realworld pathology datasets while gaining a geomean speedup of 6.9× for resolutions up to 64K2, on up to 2, 048 GPUs.

Zhang, Enzhi↗

DROP DURABILITY ASSESSMENT OF ELECTRONIC ASSEMBLIES UNDER OFF-AXIS LOADING WITH SKEWED FIXTURES

This thesis studies drop durability of electronic assemblies when the acceleration vector is oriented at 45° to the out-of-plane direction of the circuit card. The off-axis drop tests are accomplished with a skewed fixture and are conducted as a proxy for multiaxial drop testing. Advanced shock testing and vibration test methods have been developed over the last few decades to better represent real-world field environments during ground-based laboratory testing. However, many of these test methods require expensive and specialized equipment not available in most laboratories. An alternative approach for approximating simultaneous loading along multiple axes on conventional equipment utilizes skewed fixtures which have seen use in off-axis random vibration and drop impact testing. These methods generally rely on the conversion of a uniaxial input load from the test equipment (using a uniaxial drop tower or shaker) into a multiaxial load when resolved in the reference frame of the test article (mounted on a skewed fixture). Skewed fixture design is presented and recommendations for conducting skewed angle drop testing are introduced based on local measurements along the skewed face of the fixture to accurately monitor the impact event. Characterization tests were performed with a skewed fixture, at simultaneous acceleration loads from 500 to 3,000 g in two (in-plane and out-of-plane) directions, while meeting standard time domain tolerances. Upon experimental characterization, drop shock durability tests were conducted on a printed circuit assembly (PCA). Mean drops-to-failure were measured and quantified with Weibull statistics. Dominant solder joint failure modes were identified via failure analysis. Prior work on inclined angle impact testing is limited, and the majority of solder joint interconnect level fatigue studies are conducted considering perpendicular loading normal the circuit card. Low-cycle fatigue curves are generated based on plastic strain and plastic work density within the solder joint. A multiscale nonlinear finite element model is used to relate board-level flexure to solder joint interconnect level plastic strain. A high strain rate solder constitutive model allows for accurate modeling of solder plasticity resulting from high-impact drop shock. Fatigue parameters are computed from the Coffin-Manson relation and Palmgren-Miner damage accumulation. This work serves to apply established low-cycle fatigue methods for conventional drop shock loading (impact normal to circuit card) to non-perpendicular loading with a skewed fixture.

Hower, Jonathan [Kansas City National Security Cam↗

Enhancing photoionization rate calculations in low-temperature plasmas using spectral methods

Photoionization plays a central role in the development of streamer discharges and other non-equilibrium plasma phenomena. It creates seed electrons, which are essential for positive streamer propagation, allowing the ionization front to move forward. Because of this, accurate modeling of photoionization is very important for predicting streamer behavior and plasma evolution. The photoionization process in air (N 2 – O 2 mixture) is often described by the Zheleznyak model (1982). This model is usually solved through Helmholtz-type equations that approximate the Zheleznyak photoionization model (Zheleznyak et al. 1982) as Partial Differential Equations (PDEs). Conventional numerical methods, such as the Finite Difference Method (FDM) or Finite Volume Method (FVM), are widely used to solve these equations. Although they are prevalent, the computational cost of these methods due to their need for matrix operations and iterative solver is demanding. To address this challenge, this work develops a spectral solver based on the Fast Fourier Transform (FFT) combined with Discrete Cosine Transform (DCT) and Discrete Sine Transform (DST) to calculate the photoionization rate efficiently in an axisymmetric cylindrical domain. This method naturally satisfies the boundary conditions used in the model and converts the PDE into algebraic ones in spectral space. Thus, avoids the need for iterative matrix solvers. When compared with FDM results, it is demonstrated that the new solver not only maintains accuracy, but also reduces the computational cost, showing a performance increase of approximately 100 compared to FDM over a wide range of problem sizes. The method is parallelized using Message Passing Interface (MPI) and has been integrated into a fluid plasma model for streamer simulation. Here, this FFT-based approach provides a fast and reliable alternative for calculating photoionization in fluid models, helping large-scale plasma simulations run faster and efficiently, and allows higher-resolution simulation without extra computational cost.

Axisymmetric system↗

Toward improved property prediction of 2D materials using many-body quantum Monte Carlo methods

The field of 2D materials has grown dramatically in the past two decades. 2D materials can be utilized for a variety of next-generation optoelectronic, spintronic, clean energy, and quantum computing applications. These 2D structures, which are often exfoliated from layered van der Waals materials, possess highly inhomogeneous electron densities and can possess short- and long-range electron correlations. The complexities of 2D materials make them challenging to study with standard mean-field electronic structure methods such as density functional theory (DFT), which relies on approximations for the unknown exchange-correlation functional. To overcome the limitations of DFT, highly accurate many-body electronic structure approaches such as diffusion Monte Carlo (DMC) can be utilized. In the past decade, DMC has been used to calculate accurate magnetic, electronic, excitonic, and topological properties in addition to accurately capturing interlayer interactions and cohesion and adsorption energetics of 2D materials. Here, this approach has been applied to 2D systems of wide interest, including graphene, phosphorene, MoS 2 , CrI 3 , VSe 2 , GaSe, GeSe, borophene, and several others. In this review article, we highlight some successful recent applications of DMC to 2D systems for improved property predictions beyond standard DFT.

2D materials↗

Offline Maximizing Minimally Invasive Proper Orthogonal Decomposition for Reduced-Order Modeling of S n Radiation Transport

Deterministic solutions to the Sn radiation transport equation can be computationally expensive to calculate. Reduced-order modeling enables efficient approximation of the full-order model (FOM) solution. We propose a novel method for constructing reduced-order models (ROMs) of the S n radiation transport equation, offline maximizing minimally invasive (OMMI) proper orthogonal decomposition (POD). POD uses the method of snapshots to create a reduced-order basis for constructing an ROM. Minimally invasive POD leverages the sweep infrastructure existing in deterministic transport codes to create a POD-based ROM, even when infeasible by traditional methods. Offline maximizing minimally invasive proper orthogonal decomposition (OMMI-POD) extends minimally invasive POD by performing sweeps offline, therefore maximizing the potential speedup. OMMI-POD does so by creating a library of reduced systems from a training set. This library of reduced systems is then interpolated to provide a rapid approximate solution of the S n radiation transport equation. The model is evaluated on a set of test problems, achieving a low error with a 466 times speedup over the FOM. Also presented is a study of the effect of sampling method on the performance of OMMI-POD, specifically comparing naive uniform sampling to the more accurate and computationally expensive greedy sampling.

97 MATHEMATICS AND COMPUTING↗

Enhanced Collisional Losses from a Magnetic Mirror Using the Lenard-Bernstein Collision Operator

Collisions are crucial in governing particle and energy transport in plasmas confined in a magnetic mirror trap. Modern gyrokinetic codes model transport in magnetic mirrors, but some utilize approximate model collision operators. This study focuses on a Pastukhov-style method of images calculation of particle and energy confinement times using a Lenard-Bernstein model collision operator. Prior work on parallel particle and energy balances used a different Fokker-Planck plasma collision operator. The method must be extended in non-trivial ways to study the Lenard-Bernstein operator. To assess the effectiveness of our approach, we compare our results with a modern finite element solver. Our findings reveal that the particle confinement time scales like a exp( a 2 ) using the Lenard-Bernstein operator, in contrast to the more accurate scaling that the Coulomb collision operator would yield a 2 exp( a 2 ), where a 2 is approximately proportional to the ambipolar potential. We propose that codes solving for collisional losses in magnetic mirrors utilizing the Lenard-Bernstein or Dougherty collision operator scale their collision frequency of any electrostatically confined species. This study illuminates the collision operator’s intricate role in the Pastukhov-style method of images calculation of collisional confinement.

fusion plasma↗

Scalable Implementation of Mean-Field and Correlation Methods Based on Lie-Algebraic Similarity Transformation of Spin Hamiltonians in the Jordan–Wigner Representation

Recent work has highlighted that the strong correlation inherent in spin Hamiltonians can be effectively reduced by mapping spins to Fermions via the Jordan−Wigner transformation (JW). The Hartree−Fock method is straightforward in the Fermionic domain and may provide a reasonable approximation to the ground state. Correlation with respect to the Fermionic mean field can be recovered based on Lie-algebraic similarity transformation (LAST) with two-body correlators. Specifically, a unitary LAST variant eliminates the dependence on site ordering, while a nonunitary LAST yields size-extensive correlation energies. Whereas the first recent demonstration of such methods was restricted to small spin systems, we present efficient implementations using analytical gradients for the optimization with respect to the mean-field reference and the LAST parameters, thereby enabling the treatment of larger clusters, including systems with local spins s > $\frac{1}{2}$.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting U 3 O 8 powder processing conditions: An AI/ML approach analyzing deep learning embeddings of SEM micrographs

High-resolution SEM images of uranium-oxide powders encode micro- and nanoscale clues to their synthesis route and calcination temperature. We trained a ResNet-50 model on 11 commercial-scale U₃O₈ classes, ammonium diuranate (ADU) or uranyl peroxide (H₂O₂) precursors calcined at temperatures ranging from 400 to 750 °C and added a 256-D projection head before the classifier to analyze the learned representation. The best of eight seeds reached 92.4 % accuracy on reserved testing data, but our focus is the structure of the embedding space rather than the accuracy and labels. We quantify class relatedness in the original 256-D space using centroid similarity and distributional distances, and we use Uniform Manifold Approximation Projection (UMAP) for visualization. ‘Unknown’ images from different preparation methods, SEM operators, and from the literature localized near the expected classes under a nearest-centroid analysis without retraining, as well as clustered in similar UMAP space. In conclusion, this embedding-centered workflow complements black-box classification by providing quantitative, similarity-based comparisons of U₃O₈ morphologies and reduces storage space by up to 98 % for image data used in millisecond vector search comparisons.

36 MATERIALS SCIENCE↗

Tests of the DFT Ladder for the Fulminic Acid Challenge

Properties of the historically pivotal fulminic acid (HCNO) molecule have been computed with a panoply of 473 density functionals of all varieties, providing a snapshot of the performance of contemporary density functional theory (DFT) for a challenging chemical system. Exhaustive tabulations and statistical analyses have been carried out for geometric parameters, vibrational frequencies, barriers to linearity, and the HCN–O dissociation energy. As the DFT ladder is climbed, confusion rather than consensus ensues regarding the details of the distinctive, extremely flat H–C–N bending potential of fulminic acid and whether the equilibrium structure is linear or bent. While high-ranking DFT functionals produce the smallest errors for the HCN + O( 3 P) → HCNO reaction energy, lower rungs emerge as the best performers for many of the bond distances and harmonic vibrational frequencies. This research shows that the current DFT zoo of approximations does not constitute a transparent ladder of increasingly accurate methods that consistently converges on definitive predictions for various properties of HCNO. Additional analyses are performed on the side effects of popular dispersion corrections on the covalently bonded properties and thermochemistry of HCNO.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Two- and three-meson scattering amplitudes with physical quark masses from lattice QCD

We study systems of two and three mesons composed of pions and kaons at maximal isospin using four CLS ensembles with 𝑎 ≈ 0.063 fm, including one with approximately physical quark masses. Using the stochastic Laplacian-Heaviside method, we determine the energy spectrum of these systems including many levels in different momentum frames and irreducible representations. Using the relativistic two- and three-body finite-volume formalism, we constrain the two- and three-meson K matrices, including not only the leading 𝑠 wave, but also 𝑝 and 𝑑 waves. By solving the three-body integral equations, we determine, for the first time, the physical-point scattering amplitudes for 3⁢𝜋 + , 3⁢𝐾 + , 𝜋 + ⁢𝜋 + ⁢𝐾 + , and 𝐾 + ⁢𝐾 + ⁢𝜋 + systems. These are determined for total angular momentum 𝐽 𝑃 = 0 − , 1 + , and 2 − . We also obtain accurate results for 2⁢𝜋 + , 𝜋 + ⁢𝐾 + , and 2⁢𝐾 + phase shifts. We compare our results to chiral perturbation theory and to phenomenological fits.

FOS: Physical sciences↗

Lossy Compression: An Online Multi-Stage Technology for High-Fidelity Synchro- Waveform Measurements

Effective real-time monitoring and analysis of distributed grids necessitate the use of synchro-waveform measurements, which capture almost all high-frequency disturbances and transient phenomena. However, due to limitations in high-speed measurements and network bandwidth, it is challenging to transfer all high-fidelity synchro-waveforms losslessly and successfully. To cope with these challenges, a hybrid-based online multi-stage compression algorithm is proposed to significantly improve the compression efficiency for synchro-waveform measurements. Initially, the multiple discrete Wavelet transformation is deployed to deconstruct the waveform components. The delta encoding is further developed to decrease the magnitude. In conjunction with the Lempel-Ziv-Markov chain, the hybrid compression algorithm is implemented to achieve real-time compression for the synchro-waveform measurements. Moreover, an innovative error index that synergizes the time and frequency domain error and correlation is formulated to evaluate the waveform distortion. By integrating compression ratio, suitable parameters can be optimally selected. Finally, the simulation, laboratory experiments, as well as field tests across a spectrum of sampling frequencies and time intervals are conducted to substantiate the efficacy of the proposed method. Here, the outcomes demonstrated that a compression ratio of approximately 15.5 and 17.83 can be reached for 0.5 s and 1 s data under both offline and online scenarios, which equates to a substantial 93.5% to 94.39% reduction in data storage requirements.

High-fidelity synchro-waveform measurements↗

Polynomial Chaos Surrogate Construction for Random Fields with Parametric Uncertainty

Engineering and applied science rely on computational experiments to rigorously study physical systems. The mathematical models used to probe these systems are highly complex, and sampling-intensive studies often require prohibitively many simulations for acceptable accuracy. Surrogate models provide a means of circumventing the high computational expense of sampling such complex models. In particular, polynomial chaos expansions (PCEs) have been successfully used for uncertainty quantification studies of deterministic models where the dominant source of uncertainty is parametric. We discuss an extension to conventional PCE surrogate modeling to enable surrogate construction for stochastic computational models that have intrinsic noise in addition to parametric uncertainty. We develop a PCE surrogate on a joint space of intrinsic and parametric uncertainty, enabled by Rosenblatt transformations, which are evaluated via kernel density estimation of the associated conditional cumulative distributions. Furthermore, we extend the construction to random field data via the Karhunen–Loève expansion. We then take advantage of closed-form solutions for computing PCE Sobol indices to perform a global sensitivity analysis of the model which quantifies the intrinsic noise contribution to the overall model output variance. Additionally, the resulting joint PCE is generative in the sense that it allows generating random realizations at any input parameter setting that are statistically approximately equivalent to realizations from the underlying stochastic model. The method is demonstrated on a chemical catalysis example model and a synthetic example controlled by a parameter that enables a switch from unimodal to bimodal response distributions.

97 MATHEMATICS AND COMPUTING↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Estimating IDP Origins Using ACLED Political Violence Events: Validation with Lebanon IDP Flow Data

A key input in modeling population distribution and flows following conflict events is the inclusion of Internally Displaced Person (IDP) flows between administrative units within a country. These flows are critical for capturing population movement and redistribution driven by current events, particularly conflict. In some cases, IDP destination data are available while origin data is incomplete or unavailable. This creates a gap in understanding where displacement is occurring, limiting the ability to model population redistribution accurately. Without origin data, it is not possible to reallocate population flows or accurately represent where displacement is occurring within the country. This report evaluates whether Armed Conflict Location Event Data (ACLED) political violence event data can be used to estimate IDP origin distributions when direct origin data are unavailable. The approach is validated using historical IDP flow data from Lebanon, where both origins and destinations are observed. Results show that ACLED event distributions strongly correspond to observed IDP-origin patterns, particularly when using cumulative 60-day event windows. The method is most reliable for identifying major origin districts and approximating proportional origin shares. However, it is not intended to reconstruct exact individual displacement flows, but rather to provide a probabilistic spatial allocation of displacement origins.

99 GENERAL AND MISCELLANEOUS↗

Ab initio calculation of atomic solid hydrogen phases based on Gutzwiller many-body wave functions

We apply two ab initio many-body methods based on Gutzwiller wave functions, i.e., correlation matrix renormalization theory (CMRT) and Gutzwiller conjugate gradient minimization (GCGM), to the study of crystalline phases of atomic hydrogen. Both methods avoid empirical Hubbard U parameters and are free from double-counting issues. CMRT employs a Gutzwiller-type approximation that enables efficient calculations, while GCGM goes beyond this approximation to achieve higher accuracy at higher computational cost. By benchmarking against available quantum Monte Carlo (QMC) results, we demonstrate that while both methods are more accurate than the widely used density-functional theory, GCGM systematically captures additional correlation energy missing in CMRT, leading to significantly improved total energy predictions. We also show that by including the correlation energy Ec from local density approximation in the CMRT calculation, CMRT + E c produces energy in better agreement with the QMC results in these hydrogen lattice systems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Physics-based stabilized finite element approximations of the Poisson–Nernst–Planck equations

We present and analyze two stabilized finite element methods for solving numerically the Poisson–Nernst–Planck equations. The stabilization we consider is carried out by using a shock detector and a discrete graph Laplacian operator for the ion equations, whereas the discrete equation for the electric potential need not be stabilized. Discrete solutions stemmed from the first algorithm preserve both maximum and minimum discrete principles. For the second algorithm, its discrete solutions are conceived so that they hold discrete principles and obey an entropy law provided that an acuteness condition is imposed for meshes. Remarkably the latter is found to be unconditionally stable. We validate our methodology through transient numerical experiments that show convergence toward steady-state solutions.

97 MATHEMATICS AND COMPUTING↗

Spatially Accelerated Winding Numbers for Curved Geometry

The generalized winding number (GWN) is a scalar field that supports robust containment queries on curved geometry, including non-watertight, overlapping, and nested boundary representations. While queries can be easily parallelized over samples, direct evaluation on parametric curves and surfaces remains costly for large and complex models. Fast, state-of-the-art GWN approaches leverage a spatial index to approximate the GWN, typically coupled with a Taylor expansion which approximates the GWN contribution for far clusters of geometric primitives. However, such methods operate only on discrete inputs such as triangle meshes and point clouds, and would introduce containment errors near boundaries if applied to curved input. We extend support for fast GWN evaluation over arbitrary collections of NURBS curves in 2D and trimmed NURBS patches in 3D via a Bounding Volume Hierarchy that stores efficiently precomputed moment data in the hierarchy nodes. When querying the hierarchy, approximations for far clusters are used alongside direct evaluation for nearby NURBS primitives, achieving sub-linear complexity while preserving the geometric features in the vicinity of the query point. Central to our performance improvements is an adaptive subdivision strategy for NURBS primitives during a preprocessing phase, creating better spatial partitions while retaining the same accuracy for containment decisions as a direct evaluation. We demonstrate the performance and accuracy of our approach across a large collection of 2D and 3D datasets.

Computer science↗

New particle pusher with hadronic interactions for modeling multimessenger emission from compact objects

We propose novel numerical schemes based on the Boris method in curved spacetime, incorporating both hadronic and radiative interactions for the first time. Once the proton has lost significant energy due to radiative and hadronic losses, and its gyroradius has decreased below typical scales on which the electromagnetic field varies, we apply a guiding center approximation (GCA). We fundamentally simulate collision processes either with a Monte-Carlo method or, where applicable, as a continuous energy loss, contingent on the local optical depth. To test our algorithm for the first time combining the effects of electromagnetic, gravitational, and radiation fields including hadronic interactions, we simulate highly relativistic protons traveling through various electromagnetic fields and proton backgrounds. We provide unit tests in various spatially dependent electromagnetic and gravitational fields and background photon and proton distributions, comparing the trajectory against analytic results. We propose that our method can be used to analyze hadronic interactions in black hole accretion disks, jets, and coronae to study the neutrino abundance from active galactic nuclei.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗