Search NASA⌕ Search

SEARCH · Search NASA

Results for “Iterative”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Realizability-preserving discontinuous Galerkin method for spectral two-moment radiation transport in special relativity

Here we present a realizability-preserving numerical method for solving a spectral two-moment model to simulate the transport of massless, neutral particles interacting with a steady background material moving with relativistic velocities. The model is obtained as the special relativistic limit of a four-momentum-conservative general relativistic two-moment model. Using a maximum-entropy closure, we solve for the Eulerian-frame energy and momentum. The proposed numerical method is designed to preserve moment realizability, which corresponds to moments defined by a nonnegative phase-space density. The realizability-preserving method is achieved with the following key components: (i) a discontinuous Galerkin phase-space discretization with specially constructed numerical fluxes in the spatial and energy dimensions; (ii) a strong stability-preserving implicit-explicit time-integration method; (iii) a realizability-preserving conserved to primitive moment solver; (iv) a realizability-preserving implicit collision solver; and (v) a realizability-enforcing limiter. Component (iii) is necessitated by the closure procedure, which closes higher order moments nonlinearly in terms of primitive moments. The nonlinear conserved to primitive and the implicit collision solves are formulated as fixed-point problems, which are solved with custom iterative solvers designed to preserve the realizability of each iterate. With a series of numerical tests, we demonstrate the accuracy and robustness of this discontinuous-Galerkin-implicit-explicit method.

79 ASTRONOMY AND ASTROPHYSICS↗

Analytical desmearing of Bonse–Hart ultra-small-angle neutron scattering data via truncated Abel inversion

A non-iterative analytical framework based on the truncated Abel inversion is developed for desmearing Bonse–Hart ultra-small-angle neutron scattering (USANS) data. The method directly inverts the slit-averaged intensity without empirical extrapolation or iterative regularization, establishing a closed-form relationship between the measured and intrinsic scattering profiles. Numerical benchmarks on representative models, including a rigid-line form factor, a Lorentzian function and a fractal structural model, demonstrate quantitative recovery of the ground-truth intensity across the full Q range. Application to a deuterated polystyrene/poly(2-vinylpyridine) blend further confirms that the approach yields smooth continuous profiles consistent with companion small-angle neutron scattering data. The truncated Abel inversion thus provides a stable, model-independent and physically transparent route for accurate desmearing of Bonse–Hart USANS measurements.

Huang, Guan-Rong [National Tsing Hua University, T↗

Characterization and Optimization of the Fitting of Quantum Correlation Functions

This case study presents a characterization and optimization of an application code for extracting parton distribution functions from high energy electron-proton scattering data. Profiling this application code reveals that the phase-space density computation accounts for 93% of the overall execution time for a single iteration on a single core. When executing multiple iterations in parallel on a multicore system, the application spends 78% of its overall execution time idling due to load imbalance. We address these issues by first transforming the application code from Python to C++ and then tackling the application load imbalance via a hybrid scheduling strategy that combines dynamic and static scheduling. These techniques result in a 62% reduction in CPU idle time and a 2.46x speedup in overall execution time per node. In addition, the typically enabled power-management mechanisms in supercomputers (e.g., AMD Turbo Core, Intel Turbo Boost, and RAPL) can significantly impact intra-node scalability when more than 50% of the CPU cores are used. This finding underscores the importance of understanding system interactions with power management, as they can adversely impact application performance, and highlights the necessity of intra-node scaling tests to identify performance degradation that inter-node scaling tests might otherwise overlook.

Chuang, Pi-Yueh [Virginia Tech,Dept. of Computer S↗

Distributed Stochastic Optimization of a Neural Representation Network for Time-Space Tomography Reconstruction

4D time-space reconstruction of dynamic events or deforming objects using X-ray computed tomography (CT) is an important inverse problem in non-destructive evaluation. Conventional back-projection based reconstruction methods assume that the object remains static for the duration of several tens or hundreds of X-ray projection measurement images (reconstruction of consecutive limited-angle CT scans). However, this is an unrealistic assumption for many in-situ experiments that causes spurious artifacts and inaccurate morphological reconstructions of the object. To solve this problem, we propose to perform a 4D time-space reconstruction using a distributed implicit neural representation (DINR) network that is trained using a novel distributed stochastic training algorithm. Our DINR network learns to reconstruct the object at its output by iterative optimization of its network parameters such that the measured projection images best match the output of the CT forward measurement model. Here, we use a forward measurement model that is a function of the DINR outputs at a sparsely sampled set of continuous valued 4D object coordinates. Unlike previous neural representation architectures that forward and back propagate through dense voxel grids that sample the object's entire time-space coordinates, we only propagate through the DINR at a small subset of object coordinates in each iteration resulting in an order-of-magnitude reduction in memory and compute for training. DINR leverages distributed computation across several compute nodes and GPUs to produce high-fidelity 4D time-space reconstructions. We use both simulated parallel-beam and experimental cone-beam X-ray CT datasets to demonstrate the superior performance of our approach.

36 MATERIALS SCIENCE↗

A Design for Remanufacturing Framework Incorporating Identification, Evaluation, and Validation: A Case Study of Hydraulic Manifold

In recent years, academic researchers and engineers in the industry have widely recognized the necessity of integrating remanufacturing considerations into product design iterations to advance sustainability objectives. Acknowledging the importance of design for remanufacturing (DfRem), efforts were made to develop tools and guidelines that could be implemented in practice. However, such methods largely rely upon experiential insights and qualitative assessments, leaving a gap in the ability to quantitatively assess the economic and environmental impacts of design choices. To bridge this gap, we investigate existing efforts and present a framework for DfRem that integrates established design and remanufacturing practices into a cohesive workflow with quantitative assessments. To demonstrate its efficacy for making practical design changes for remanufacturing, we apply the framework to a hydraulic manifold in a transmission system for heavy-duty tractors. Through this industry-relevant case study, we focus on showcasing the practical utility of our framework. Based on the identified design modifications from remanufacturability analysis, we estimate the reductions in life cycle costs, energy consumption, and emissions. Afterward, the modifications are tested using physical experiments with plans for integration into future iterations of the hydraulic manifold design and production. Here, we anticipate this framework can illustrate the process of remanufacturing that ensures improvements in sustainability while maintaining performance and reliability standards.

design for X↗

High-throughput homogenization of a quasi-Gaussian ultrafast laser beam using a combined refractive beam shaper and spatial light modulator

Efficiently shaping femtosecond, transverse Gaussian laser beams to flat-top beams with flat wavefronts is critical for large-scale material processing and manufacturing. Existing beam shaping devices fall short either in final beam homogeneity or efficiency. Here, we present an approach that uses refractive optics to perform the majority of the beam shaping and then uses a fine-tune device (spatial light modulator) to refine the intensity profile. For the beam that we selected, circularly asymmetric with intensity fluctuations, our method achieved a uniformity of 0.055 within 90% of the beam area at 92% efficiency. The optimization involved an iterative beam shaping process that converged to optimum within 10 iterations.

43 PARTICLE ACCELERATORS↗

Scalable Multiphysics Block Preconditioning for Low Mach Number Compressible Resistive MHD with Application to Magnetic Confinement Fusion

This study investigates multiphysics block preconditioners that are critical in devising scalable Newton–Krylov iterative solvers for longer time-scale fully implicit fluid plasma models. The specific model of interest is the visco-resistive, low Mach number, compressible magnetohydrodynamics (MHD) model. This model describes the dynamics of conducting fluids in the presence of electromagnetic fields and can be used to study aspects of astrophysical phenomena, important science and technology applications, and basic plasma physics. The specific application of interest that motivates this study is the macroscopic simulation of longer time-scale stability and disruptions of magnetic confinement fusion devices, specifically the ITER Tokamak. The computational solution of the governing balance equations for mass, momentum, heat transfer, and magnetic induction for resistive MHD systems can be extremely challenging. These difficulties arise from both the strong nonlinear, nonsymmetric coupling of fluid and electromagnetic phenomena as well as the significant range of time and length scales that the interactions of these physical mechanisms produce. To handle the range of time and spatial scales of interest, a fully implicit unstructured variational multiscale finite element formulation is employed. For the scalable solution of the Newton linearized systems, fully coupled block preconditioners are designed to leverage algebraic multigrid subsolves. In conclusion, results are presented for the strong and weak scaling of the method as well as the robustness of these techniques for a large range of Lundquist numbers.

97 MATHEMATICS AND COMPUTING↗

G-Mapper: Learning a Cover in the Mapper Construction

The Mapper algorithm is a visualization technique in topological data analysis (TDA) that outputs a graph reflecting the structure of a given dataset. However, the Mapper algorithm requires tuning several parameters in order to generate a “nice” Mapper graph. This paper focuses on selecting the cover parameter. We present an algorithm that optimizes the cover of a Mapper graph by splitting a cover repeatedly according to a statistical test for normality. Our algorithm is based on G-means clustering, which searches for the optimal number of clusters in 𝑘-means by iteratively applying the Anderson–Darling test. Our splitting procedure employs a Gaussian mixture model to carefully choose the cover according to the distribution of the given data. In conclusion, experiments for synthetic and real-world datasets demonstrate that our algorithm generates covers so that the Mapper graphs retain the essence of the datasets, while also running significantly faster than a previous iterative method.

G-means clustering↗

Parallel-in-Time Solution of Hyperbolic PDE Systems via Characteristic-Variable Block Preconditioning

We consider the parallel-in-time solution of both linear and nonlinear hyperbolic partial differential equation (PDE) systems in one spatial dimension. In the nonlinear setting, the discretized equations are solved with a preconditioned residual iteration based on a global linearization. The linear(ized) equation systems are approximately solved parallel-in-time using a block preconditioner applied in the characteristic variables of the underlying linear(ized) hyperbolic PDE. This change of variables is motivated by the observation that intervariable coupling between characteristic variables is weak, at least locally where spatio-temporal variations in the eigenvectors of the associated flux Jacobian are sufficiently small, while that between the original variables is not. For an ℓ-dimensional system of PDEs, applying the preconditioner consists of solving a sequence of ℓ scalar linear(ized)-advection-like problems, each associated with a different characteristic wave-speed in the underlying linear(ized) PDE. Furthermore, we approximately solve these linear advection problems using multigrid reduction-in-time (MGRIT); however, any other suitable parallel-in-time method could be used. Numerical examples are shown for the (linear) acoustics equations in heterogeneous media and for the (nonlinear) shallow water equations and Euler equations of gas dynamics with shocks and rarefactions. For many test problems, the solver converges in just a handful of iterations and with mesh-independent convergence rates.

97 MATHEMATICS AND COMPUTING↗

On High-Order/Low-Order and Micro-Macro Methods for Implicit Time-Stepping of the BGK Model

In this paper, a high-order/low-order (HOLO) method is combined with a micro-macro (MM) decomposition to accelerate iterative solvers in fully implicit time-stepping of the Bhatnagar–Gross–Krook (BGK) equation for gas dynamics. The MM formulation represents a kinetic distribution as the sum of a local Maxwellian and a perturbation. In highly collisional regimes, the perturbation away from initial and boundary layers is small and can be compressed to reduce the overall storage cost of the distribution. The convergence behavior of the MM methods, the usual HOLO method, and the standard source iteration method is analyzed on a linear BGK model. Both the HOLO and MM methods are implemented using a discontinuous Galerkin (DG) discretization in phase space, which naturally preserves the consistency between high- and low-order models required by the HOLO approach. Furthermore, the accuracy and performance of these methods are compared on the Sod shock tube problem and a sudden wall heating boundary layer problem. Overall, the results demonstrate the robustness of the MM and HOLO approaches and illustrate the compression benefits enabled by the MM formulation when the kinetic distribution is near equilibrium.

BGK model↗

Algebraic Multigrid with Filtering: An Efficient Preconditioner for Interior Point Methods in Large-Scale Contact Mechanics Optimization

Large-scale contact mechanics simulations are crucial in many engineering fields such as structural design and manufacturing. In the frictionless case, contact can be modeled by minimizing an energy functional; however, these problems are often nonlinear, nonconvex, and increasingly difficult to solve as mesh resolution increases. In this work, we employ a Newton-based interior-point (IP) filter line-search method, an effective approach for large-scale constrained optimization. While this method converges rapidly, each iteration requires solving a large saddle-point linear system that becomes ill-conditioned as the optimization process converges, largely due to IP treatment of the contact constraints. Such ill-conditioning can hinder solver scalability and increase iteration counts with mesh refinement. Here, to address this, we introduce a novel preconditioner, algebraic multigrid with filtering (AMGF), tailored to the Schur complement of the saddle-point system. Building on the classical AMG solver, commonly used for elasticity, we augment it with a specialized subspace correction that filters near null space components introduced by contact interface constraints. Through theoretical analysis and numerical experiments on a range of linear and nonlinear contact problems, we demonstrate that the proposed solver achieves mesh independent convergence and maintains robustness against the ill-conditioning that notoriously plagues IP methods. These results indicate that AMGF makes contact mechanics simulations more tractable and broadens the applicability of Newton-based IP methods in challenging engineering scenarios. More broadly, AMGF is well suited for problems, optimization or otherwise, where solver performance is limited by a low-dimensional subspace, such as those arising from localized constraints, interface conditions, or model heterogeneities. This makes the method widely applicable beyond contact mechanics and constrained optimization.

Mathematics and Computing↗

On two nucleons near unitarity with perturbative pions

We explore the impact of perturbative pions on the Unitarity Expansion in the two-nucleon S-waves of Chiral Effective Field Theory at next-to-next-to leading order (N 2 LO). Pion exchange explicitly breaks the nontrivial fixed point’s universality, i.e. invariance of S waves under both conformal and Wigner’s combined SU(4) spin–isospin transformations. On the other hand, Unitarity explicitly breaks chiral symmetry. The two seem incompatible in their respective exact-symmetry limits. χEFT with Perturbative Pions in the Unitarity Expansion resolves the apparent conflict in the Unitarity Window (phase shifts 45° ≲ δ(k) ≲ 135°), i.e. around momenta most relevant for low-energy nuclear systems. Its only LO scale is the scattering momentum; NLO adds only scattering length, effective range and non-iterated one-pion exchange (OPE); and N 2 LO only once-iterated OPE. Agreement in the 1 S 0 channel is very good.

Chiral symmetry↗

Future Circular Collider Feasibility Study Report

Volume 3 of the FCC Feasibility Report presents studies related to civil engineering, the development of a project implementation scenario, and environmental and sustainability aspects. The report details the iterative improvements made to the civil engineering concepts since 2018, taking into account subsurface conditions, accelerator and experiment requirements, and territorial considerations. It outlines a technically feasible and economically viable civil engineering configuration that serves as the baseline for detailed subsurface investigations, construction design, cost estimation, and project implementation planning. Additionally, the report highlights ongoing subsurface investigations in key areas to support the development of an improved 3D subsurface model of the region. The report describes the development of the project scenario based on the ‘avoid-reduce-compensate’ iterative optimisation approach. The reference scenario balances optimal physics performance with territorial compatibility, implementation risks, and costs. Environmental field investigations covering almost 600 hectares of terrain—including numerous urban, economic, social, and technical aspects—confirmed the project’s technical feasibility and contributed to the preparation of essential input documents for the formal project authorisation phase. The summary also highlights the initiation of public dialogue as part of the authorisation process. The results of a comprehensive socio-economic impact assessment, which included significant environmental effects, are presented. Even under the most conservative and stringent conditions, a positive benefit-cost ratio for the FCC-ee is obtained. Finally, the report provides a summary of the studies conducted to document the current state of the environment.

Benedikt, M. [European Organization for Nuclear Re↗

Progressive Dynamics for Cloth and Shell Animation

We propose Progressive Dynamics, a coarse-to-fine, level-of-detail simulation method for the physics-based animation of complex frictionally contacting thin shell and cloth dynamics. Progressive Dynamics provides tight-matching consistency and progressive improvement across levels, with comparable quality and realism to high-fidelity, IPC-based shell simulations [Li et al. 2021] at finest resolutions. Together these features enable an efficient animation-design pipeline with predictive coarse-resolution previews providing rapid design iterations for a final, to-be-generated, high-resolution animation. In contrast, previously, to design such scenes with comparable dynamics would require prohibitively slow design iterations via repeated direct simulations on high-resolution meshes. We evaluate and demonstrate Progressive Dynamics's features over a wide range of challenging stress-tests, benchmarks, and animation design tasks. Here Progressive Dynamics efficiently computes consistent previews at costs comparable to coarsest-level direct simulations. Its matching progressive refinements across levels then generate rich, high-resolution animations with high-speed dynamics, impacts, and the complex detailing of the dynamic wrinkling, folding, and sliding of frictionally contacting thin shells and fabrics.

Computer Science↗

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin [Fermilab] (ORCID:0000000157000288↗

Field To Farm Aggregation For Agricultural Systems

The Fields to Farms methodology illustrates the generation of farm parcels from the Crop Data Layer (CDL), a raster dataset containing 133 categories representing various crop types and land uses. This methodology involves two primary steps: Field Delineation and Farm Aggregation. The code specifically addresses the aggregation of pre-delineated fields within a county to form farms, adhering to predefined criteria for farm size categories. It is assumed that the field delineation process, which involves creating vector polygons from CDL raster, has been completed beforehand, possibly through external tools or methods. Upon initialization, the script processes county-level fields, preparing them for farm aggregation. In the Farm Aggregation phase, the code iteratively combines delineated fields into farms based on specified criteria, continuing until the aggregated farm size meets predefined thresholds derived from data from the 2017 National Agricultural Statistics Service (NASS) census. Throughout this iterative process, the script dynamically adjusts the aggregation to ensure alignment with the desired distribution reported by NASS. The resulting output of the script is a GeoDataFrame containing classified farms, which are subsequently saved as GeoPackage files. These files enable further analysis and visualization, facilitating comprehensive exploration of the farm landscape generated through the methodology.

Paudel, Rajiv [Idaho National Laboratory (INL), Id↗

Generic Discretization Library

The GenDiL library is a collection of C++ software abstractions designed to discretize and solve partial differential equations (PDEs) for high-performance computing (HPC) applications. Its primary focus is on modern C++ generic programming, which helps ensure portability across various hardware architectures. The central idea behind the library is to provide building blocks for numerical algorithms-such as discretization methods and iteration patterns-so that domain experts can focus on the math, rather than the low-level details of hardware or implementation. By defining abstractions for data types, iteration over computational grids, and scheduling of operations, the library isolates the high-level PDE algorithms from the platform-specific optimizations needed to achieve efficient performance.

Dudouit, Yohann [Lawrence Livermore National Labor↗

powersqueeze

powersqueeze (psqz) is a truncated power iteration library intended for high-performance computing platforms. psqz efficiently produces low-dimensional, linear measurements of graph matrix spectra by combining classical power iteration with sparse Johnson-Lindenstrauss transforms. psqz is intended to produce high-quality, fast, data-oblivious low-dimensional representations of high-dimensional sparse data such as graphs and term-document matrices. psqz is intended to replace similar workflows that depend on directly approximating a truncated eigendecomposition (e.g., the first step of spectral clustering), which is a much more expensive operation.

Priest, BenjaminW [Lawrence Livermore National Lab↗