Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer implementation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Enriching the physics program of the CMS experiment via data scouting and data parking

Specialized data-taking and data-processing techniques were introduced by the CMS experiment in Run 1 of the CERN LHC to enhance the sensitivity of searches for new physics and the precision of standard model measurements. These techniques, termed data scouting and data parking, extend the data-taking capabilities of CMS beyond the original design specifications. The novel data-scouting strategy trades complete event information for higher event rates, while keeping the data bandwidth within limits. Data parking involves storing a large amount of raw detector data collected by algorithms with low trigger thresholds to be processed when sufficient computational power is available to handle such data. The research program of the CMS Collaboration is greatly expanded with these techniques. The implementation, performance, and physics results obtained with data scouting and data parking in CMS over the last decade are discussed in this Report, along with new developments aimed at further improving low-mass physics sensitivity over the next years of data taking.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Random Phase Approximation Correlation Energy Using Real-Space Density Functional Perturbation Theory

We present a real-space method for computing the random phase approximation (RPA) correlation energy within Kohn–Sham density functional theory, leveraging the low-rank nature of the frequency-dependent density response operator. In particular, we employ a cubic-scaling formalism based on density functional perturbation theory that circumvents the calculation of the response function matrix, instead relying on the ability to compute its product with a vector through the solution of the associated Sternheimer linear systems. We develop a large-scale parallel implementation of this formalism using the subspace iteration method in conjunction with the spectral quadrature method while employing the Kronecker product-based method for the application of the Coulomb operator and the conjugate orthogonal conjugate gradient method for the solution of the linear systems. We demonstrate convergence with respect to key parameters and verify the method’s accuracy by comparing with plane-wave results. We show that the framework achieves good strong scaling to many thousands of processors, reducing the time to solution for a lithium hydride system with 128 electrons to around 150 s on 4608 processors.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Characterization and thermometry of dissipatively stabilized steady states

In this work we study the properties of dissipatively stabilized steady states of noisy quantum algorithms, exploring the extent to which they can be well approximated as thermal distributions, and proposing methods to extract the effective temperature T. We study an algorithm called the relaxational quantum eigensolver (RQE), which is one of a family of algorithms that attempt to find ground states and balance error in noisy quantum devices. In RQE, we weakly couple a second register of auxiliary ‘shadow’ qubits to the primary system in Trotterized evolution, thus engineering an approximate zero-temperature bath by periodically resetting the auxiliary qubits during the algorithm’s runtime. Balancing the infinite temperature bath of random gate error, RQE returns states with an average energy equal to a constant fraction of the ground state. We probe the steady states of this algorithm for a range of base error rates, using several methods for estimating both T and deviations from thermal behavior. In particular, we both confirm that the steady states of these systems are often well-approximated by thermal distributions, and show that the same resources used for cooling can be adopted for thermometry, yielding a fairly reliable measure of the temperature. These methods could be readily implemented in near-term quantum hardware, and for stabilizing and probing Hamiltonians where simulating approximate thermal states is hard for classical computers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Scale-Bridging Optimization Framework for Desalination Integrated Produced Water Networks

In this work, we develop a Pyomo-based non-linear optimization strategy that includes rigorous MVR models. The detailed desalination unit is integrated into the multiperiod produced water network problem using the trust region filter (TRF) method. TRF decomposes the integrated problem into a master problem consisting of the network variables and a simplified surrogate model for the detailed desalination unit. The surrogate is updated using zero and first-order corrections from the optimal solution of the detailed models at every iteration. This framework allows us to co-optimize the design of the desalination units and operating policy for the multiperiod network. A common design is ensured across all periods using global capacity constraints. We validate the solution obtained using the TRF method by solving the full integrated problem for small network instances and show our results on real case studies on produced water networks from the Permian and Appalachian basins. In this work, we describe our TRF formulation, give details on our implementation in Pyomo, and analyze the results obtained by solving the optimization problem using IPOPT. We also present a discussion on the computational efficiency and scaling using the TRF approach against a full-scale integration of the rigorous models within the water network.

Naik, Sakshi↗

Evaluation and Development of Phase Array Ultrasonic Testing (PAUT) System for Additively Manufactured Parts

This research focuses on the application of advanced ultrasonic testing techniques developed by The Phased Array Company (TPAC) for inspecting defects in additive manufacturing (AM) parts. Traditionally, X-ray computed tomography is the standard for inspecting AM components. Although, the long inspection and analysis time, along with relatively high cost make implementation difficult. Thus, an alternative nondestructive evaluation (NDE) approach is necessary to support quality assurance efforts within the field of AM. TPAC is recognized as a leader in ultrasonic testing innovation, deploying sophisticated algorithms such as Total Focusing Method (TFM) and Phased Wave Imaging (PWI) for ultrasonic data processing and interpretation. This work will explore how the TFM and PWI algorithms can assist defect detection within polymer AM parts. The AM field is seeking novel NDE methods to provide support within quality control and assurance efforts. Advanced ultrasonics inspection have the potential to fulfill this need.

99 GENERAL AND MISCELLANEOUS↗

Evaluation and Development of Phase Array Ultrasonic Testing (PAUT) System for Additively Manufactured Parts

This research focuses on the application of advanced ultrasonic testing techniques developed by The Phased Array Company (TPAC) for inspecting defects in additive manufacturing (AM) parts. Traditionally, X-ray computed tomography is the standard for inspecting AM components. Although, the long inspection and analysis time, along with relatively high cost make implementation difficult. Thus, an alternative nondestructive evaluation (NDE) approach is necessary to support quality assurance efforts within the field of AM. TPAC is recognized as a leader in ultrasonic testing innovation, deploying sophisticated algorithms such as Total Focusing Method (TFM) and Phased Wave Imaging (PWI) for ultrasonic data processing and interpretation. This work will explore how the TFM and PWI algorithms can assist defect detection within polymer AM parts. The AM field is seeking novel NDE methods to provide support within quality control and assurance efforts. Advanced ultrasonics inspection have the potential to fulfill this need.

36 MATERIALS SCIENCE↗

HPC I/O innovations in the exascale era

As high performance computing architecture evolves to deliver ever-increasing performance, the middleware tools also need to adapt in order for applications to better use these higher-performance features. Here, the Adaptable Input Output System (ADIOS), which provides scalable IO performance for exascale HPC applications is one such middleware. During the Exascale Computing Project (ECP), key portions of the ADIOS environment were adapted to respond to ongoing developments in exascale computing and the stresses and opportunities inherent in those changes. This paper examines those changes and where appropriate compares them to pre-exascale implementations.

ADIOS↗

Terra-Populus v0.1: A Python Library for LandScan High-Definition Population Analysis and Modeling

The terra-populus library is designed for use by the LandScan HD technical team, offering a streamlined set of tools for generating and updating LandScan HD datasets from foundational building-level data, referred to as 'molecules,' provided by the building-level attribution team. This document serves as the primary technical documentation for terra-populus. Version 0.1 of the library includes the core modeling components necessary for LandScan HD production. It enables the generation of the LandScan HD Baseline dataset as well as corresponding confidence measures for the occupancy rates used. Parameters have been included for incorporating damaged building indicators and changes in population, to faciliate the creation of rapid updates for LandScan HD. Future iterations of terra-populus will introduce tools for creating a confidence index, and quantifying and propagating uncertainty, facilitating the creation of probabilistic LandScan HD outputs. This report provides an overview of the tools available in the library and the corresponding code implementations. One of the key advancements implemented in terra-populus is a redefinition of the atomic modeling unit for LandScan HD. Traditionally, the LandScan HD vector analytical framework has generated population estimates at the building sub-component (molecule) level. However, terra-populus adopts a building-level modeling approach. This shift is an operational decision aimed at aligning LandScan HD outputs with confidence measures, which are computed and validated at the building level (confidence measures are not included in this version of terra-populus, aside from those associated with the occupancy rates). Additional advancements to the LandScan HD modeling, as implemented by terra-populus, include a minimum population value parameter and an auto assignment of building floor counts. The population minimum value was implemented to prevent buildings and subsequent LandScan HD pixels that contained small values that may not rasterize in production. An 'auto' value has been included as a method for dealing with buildings lacking floor count information, where it is the average floor count of all other buildings with a residential building use type tag. The logic behind this is to remain consistent with the current logic employed for dealing with building use type null instances, where a null use type is defaulted to residential since it is the most common building type. The auto logic is intended to apply the most common building floor count of the most common type of buildings. The tools provided in terra-populus represent a significant step forward in improving the efficiency, reproducibility, and transparency of the LandScan HD modeling process. As the library evolves, it will continue to serve as a foundational resource for high-resolution population modeling.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Regularized Differentiation for Bioburden Density Estimation in Planetary Protection

In this paper, we propose and investigate the performance of two novel shrinkage estimators for bioburden density estimation in planetary protection. The estimators are based on the regularized differentiation of a cumulative count of colony forming units collected throughout the data collecting session or the life cycle of the entire mission. The regularized differentiation recasts the problem of bioburden density estimation as a linear least squares problem. The least squares problem is then solved through regularization techniques, such as truncated singular value decomposition and penalized least squares. The regularization is necessary to avoid noise amplification during the differentiation of noisy data. The two regularization estimators are compared with four other commonly used estimators to simultaneously evaluate the means of multivariable independent Poisson distributions: the maximum likelihood, noninformative Bayes estimator with Jeffreys prior, Empirical Bayes using conjugate gamma-Poisson model with gamma parameters selected by method of moments, and the Clevenson-Zidek estimator. It is shown through computer-simulated data that the regularized differentiation based on ridge regression has the smallest mean-squared error among all estimators. The analysis of shrinkage mechanism implemented by regularized differentiation is performed, and it is shown that the regularized differentiation amounts to performing a weighted averaging of all the samples. The weights are determined by the regularization parameter automatically selected by the L-curve technique. Since the method of least squares makes no distributional assumptions about the data, it presents an attractive technique for bioburden density estimation when there are concerns about the misspecification of the distributional model. The paper concludes with the analysis of the bioburden data collected during InSight mission and directions for future work.

97 - MATHEMATICS AND COMPUTING↗

AMR-Wind: A Performance-Portable, High-Fidelity Flow Solver for Wind Farm Simulations

We present AMR-Wind, a verified and validated high-fidelity computational-fluid-dynamics code for wind farm flows. AMR-Wind is a block-structured, adaptive-mesh, incompressible-flow solver that enables predictive simulations of the atmospheric boundary layer and wind plants. It is a highly scalable code designed for parallel high-performance computing with a specific focus on performance portability for current and future computing architectures, including graphical processing units (GPUs). In this paper, we detail the governing equations, the numerical methods, and the turbine models. Establishing a foundation for the correctness of the code, we present the results of formal verification and validation. The verification studies, which include a novel actuator line test case, indicate that AMR-Wind is spatially and temporally second-order accurate. The validation studies demonstrate that the key physics capabilities implemented in the code, including actuator disk models, actuator line models, turbulence models, and large eddy simulation (LES) models for atmospheric boundary layers, perform well in comparison to reference data from established computational tools and theory. We conclude with a demonstration simulation of a 12-turbine wind farm operating in a turbulent atmospheric boundary layer, detailing computational performance and realistic wake interactions.

17 WIND ENERGY↗

A Scalable Multi-Modal Framework for High-Fidelity Distributed Human Mobility Simulations

The development of data-driven models for human mobility in urban settings requires access to substantial and diverse real-world data. However, existing historical data often presents challenges such as limited volume, variety, and veracity, as well as missing data and privacy preservation concerns. Also, urban mobility modeling is inherently time-variant, complex, and multi-modal, encompassing everything from individual walking and running to private road travel and large-scale public transportation. These challenges call for innovative solutions to overcome data limitations and compute needs to model mobility behaviors accurately. To address these challenges, we propose a distributed, co-simulation-based architecture DURMOSim that integrates real-world data with scalable, high-fidelity simulations, demonstrating distributed co-simulation feasibility with existing mobility models. DURMOSim underpins a modular integration that would enable using any available mobility simulators for greater extensibility and scalability in performing various urban scenarios. In this paper, we present the design, implementation, and performance evaluation of DURMOSim, highlighting its capability to model population-scale mobility patterns. Our initial results show its ability to dynamically synchronize multiple simulation models at runtime with negligible computational overhead. We believe DURMOSim could be a robust tool for advancing urban mobility research and intelligent transportation systems.

Yoginath, Srikanth [ORNL] (ORCID:0000000184236050)↗

Integration of Waveform Simulation Methods

The generation of synthetic seismograms through simulation is a fundamental tool of seismology required to run quantitative hypothesis tests. A variety of approaches have been developed throughout the seismological community and each has their own specific user interface based on their implementation. This causes a challenge to researchers who will need to learn new interfaces with each new software they wish to use and create substantial challenges when attempting to compare results from different tools. Here we provide a unified interface that facilitates interoperability amongst several simulation tools through a modern containerized Python package. Further, this package includes post-processing analysis modules designed to facilitate end-to-end analysis of synthetic seismograms. In this report we present the conceptual guidance and an example implementation of the new Waveform Simulation Framework.

58 GEOSCIENCES↗

Accelerating Multivariate Functional Approximation Computation with Domain Decomposition Techniques⋆

Modeling large datasets through Multivariate Functional Approximations (MFA) provide an elegant way to handle many visualization and scientific analysis workflows. The process necessitates scalable data partitioning methods to compute MFA representations efficiently without compromising the accuracy or continuity of the reconstructed solution. We propose a domain -decomposed method for computing the MFA with B -spline bases, which reduces the total work per task and uses a restricted Additive Schwarz (RAS) method to converge the control point data degrees -of -freedom along subdomain boundaries. We provide an in-depth analysis of the parallel approach with domain decomposition solvers, aiming to minimize local subdomain error residuals and recover high -order continuity at subdomain interfaces with appropriate choices of knot overlaps. The communication cost, determined by the overlap regions in the RAS implementation, is optimized to recover the numerical error profile of the single subdomain case. Our proposed method stands in contrast to previous methods, which typically only recover either C 0 or at best C 1 continuity for arbitrary B -spline degree expansions, or those that require post -processing to blend discontinuities in the reconstructed data. We demonstrate the effectiveness of our approach using analytical and real -world datasets in 1D, 2D, and 3D through both strong and weak scaling studies. The performance results indicate that the overall cost of computing the approximation is directly proportional to the underlying nearest -neighbor communication implementation, and is only weakly dependent on the overlap region size that determines the size of the messages. This finding underscores the efficiency and scalability of our proposed method, making it a promising solution for handling large datasets in scientific workflows.

additive Schwarz solvers↗

Enabling Parallel Performance and Portability of Solid Mechanics Simulations Across CPU and GPU Architectures

Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.

Morgan, Nathaniel (ORCID:0000000276118449)↗

Optimal Twirling Depth for Classical Shadows in the Presence of Noise

The classical shadows protocol is an efficient strategy for estimating properties of an unknown state p using a small number of state copies and measurements. In its original form, it involves twirling the state with unitaries from some ensemble and measuring the twirled state in a fixed basis. It was recently shown that for computing local properties, optimal sample complexity (copies of the state required) is remarkably achieved for unitaries drawn from shallow depth circuits composed of local entangling gates, as opposed to purely local (zero depth) or global twirling (infinite depth) ensembles. Here, we consider the sample complexity as a function of the depth of the circuit, in the presence of noise. We find that this noise has important implications for determining the optimal twirling ensemble. Under fairly general conditions, we (i) show that any single-site noise can be accounted for using a depolarizing noise channel with an appropriate damping parameter f, (ii) compute thresholds f th at which optimal twirling reduces to local twirling for Pauli operators, (iii) nth order Renyi entropies (n ≥2), and (iv) provide a meaningful upper bound t max on the optimal circuit depth for any finite noise strength f, which applies to observables and entanglement entropy measurements. In conclusion, these thresholds strongly constrain the search for optimal strategies to implement shadow tomography and are easily tailored to the experimental system at hand.

97 MATHEMATICS AND COMPUTING↗

Extension of Clad Damage Propagation Model for Fission Gas Dispersal and Two-Phase Flow Effects in MOOSE SubChannel Module

This report presents an extension of the Clad Damage Propagation (CDAP) model implemented in the MOOSE SubChannel Module (SCM) to capture post-failure fission-gas dispersal and two-phase flow effects in sodium-cooled fast reactor assemblies. The extended model tracks discharged gas axially and radially, computes channel-averaged flow quality and void fraction using a Lockhart–Martinelli framework, evaluates two-phase frictional pressure-drop multipliers, determines inlet mass-flow degradation under fixed core pressures, and applies an intensified-void-based heat-transfer degradation to affected fuel pins. Radial plume expansion is parameterized using mineral-oil jet experiments mapped to sodium conditions via Reynolds–Weber similarity. Implementation details are documented, along with the new methods and user inputs needed to control plume mapping and two-phase behavior. Demonstration simulations for 19- and 37-pin bundles show that breach size and inlet velocity strongly influence propagation potential: small breaches (≤0.5 mm) produce limited degradation while larger breaches (~1 mm) can drive oscillatory temperature spikes and enhanced failure propagation, especially at higher velocities. These results demonstrate that the extended CDAP model provides a more complete framework for quantifying cladding damage propagation and evaluating propagation potential in transient scenarios. The approach remains computationally efficient, consistent with subchannel-level analysis, yet incorporates sufficient physics to bridge localized post-failure effects with bundle- and assembly-scale degradation.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Efficient Mixed-Precision Matrix Factorization of the Inverse Overlap Matrix in Electronic Structure Calculations with AI-Hardware and GPUs

In recent years, a new kind of accelerated hardware has gained popularity in the artificial intelligence (AI) community which enables extremely high-performance tensor contractions in reduced precision for deep neural network calculations. In this article, we exploit Nvidia Tensor cores, a prototypical example of such AI-hardware, to develop a mixed precision approach for computing a dense matrix factorization of the inverse overlap matrix in electronic structure theory, S –1 . This factorization of S –1 , written as ZZT = S –1 , is used to transform the general matrix eigenvalue problem into a standard matrix eigenvalue problem. Here we present a mixed precision iterative refinement algorithm where Z is given recursively using matrix–matrix multiplications and can be computed with high performance on Tensor cores. To understand the performance and accuracy of Tensor cores, comparisons are made to GPU-only implementations in single and double precision. Additionally, we propose a nonparametric stopping criteria which is robust in the face of lower precision floating point operations. The algorithm is particularly useful when we have a good initial guess to Z, for example, from previous time steps in quantum-mechanical molecular dynamics simulations or from a previous iteration in a geometry optimization.

36 MATERIALS SCIENCE↗

The Functor system: a new on-the-fly take on Material Properties based on C++ functions

In the context of solving multiphysics problems, the discretization of the partial differential equations (PDE) at hand often takes the spotlight. However, for most engineering users and even application developers, the discretization of the equations has already been performed. Instead, they are tasked with implementing specific closure relations and material properties. MOOSE has long enabled this using the Materials system. This system relied on the pre-computation of all properties before they are used in the PDE or in postprocessing. In this talk we will introduce the Functor system, which was deployed in MOOSE in 2021, then present a few applications of functors in flow modeling simulations by the NEAMS program. Functors first offer great flexibility in their evaluation. Rather than storing various arrays for material properties, they are evaluated on the fly at the location and state, e.g. current or old value, requested. Unlike regular material properties, several operations such as the time derivative, the divergence and the curl can be requested from a functor. Similar to material properties, functors can be made to depend on arbitrary combinations of variables, functions, postprocessors and other properties. However, unlike material properties, any of these can be substituted for a functor material property. Thanks to this, objects no longer need to be duplicated based on the types of their parameters.

97 - MATHEMATICS AND COMPUTING↗