Search NASA⌕ Search

SEARCH · Search NASA

Results for “KERNEL FUNCTION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Single-Event Effects Test Report Linux Operating System Configurations on TUL PYNQ-Z2

This study was undertaken to determine the single-event functional interrupt (SEFI) susceptibility on different Linux operating system configurations. The device-under-test (DUT) was the Xilinx Zynq-7020 SoC on the TUL PYNQ-Z2 board. The device was monitored for kernel panics or hangs, classified as SEFIs, to observe any differences in the SEFI cross sections between operating system configurations. The primary purpose of this experiment was to observe if the number of drivers installed in a Linux system affects its overall execution reliability.

Seth S Roffe↗

Transition operators in acoustic-wave diffraction theory. I - General theory. II - Short-wavelength behavior, dominant singularities of Zk0 and Zk0 exp -1

A formal theory of the scattering of time-harmonic acoustic scalar waves from impenetrable, immobile obstacles is established. The time-independent formal scattering theory of nonrelativistic quantum mechanics, in particular the theory of the complete Green's function and the transition (T) operator, provides the model. The quantum-mechanical approach is modified to allow the treatment of acoustic-wave scattering with imposed boundary conditions of impedance type on the surface (delta-Omega) of an impenetrable obstacle. With k0 as the free-space wavenumber of the signal, a simplified expression is obtained for the k0-dependent T operator for a general case of homogeneous impedance boundary conditions for the acoustic wave on delta-Omega. All the nonelementary operators entering the expression for the T operator are formally simple rational algebraic functions of a certain invertible linear radiation impedance operator which maps any sufficiently well-behaved complex-valued function on delta-Omega into another such function on delta-Omega. In the subsequent study, the short-wavelength and the long-wavelength behavior of the radiation impedance operator and its inverse (the 'radiation admittance' operator) as two-point kernels on a smooth delta-Omega are studied for pairs of points that are close together.

Hahne, G. E.↗

An Optimized Multicolor Point-Implicit Solver for Unstructured Grid Applications on Graphics Processing Units

In the field of computational fluid dynamics, the Navier-Stokes equations are often solved using an unstructuredgrid approach to accommodate geometric complexity. Implicit solution methodologies for such spatial discretizations generally require frequent solution of large tightly-coupled systems of block-sparse linear equations. The multicolor point-implicit solver used in the current work typically requires a significant fraction of the overall application run time. In this work, an efficient implementation of the solver for graphics processing units is proposed. Several factors present unique challenges to achieving an efficient implementation in this environment. These include the variable amount of parallelism available in different kernel calls, indirect memory access patterns, low arithmetic intensity, and the requirement to support variable block sizes. In this work, the solver is reformulated to use standard sparse and dense Basic Linear Algebra Subprograms (BLAS) functions. However, numerical experiments show that the performance of the BLAS functions available in existing CUDA libraries is suboptimal for matrices representative of those encountered in actual simulations. Instead, optimized versions of these functions are developed. Depending on block size, the new implementations show performance gains of up to 7x over the existing CUDA library functions.

Zubair, Mohammad↗

Static Analysis Using Abstract Interpretation

Short presentation about static analysis and most particularly abstract interpretation. It starts with a brief explanation on why static analysis is used at NASA. Then, it describes the IKOS (Inference Kernel for Open Static Analyzers) tool chain. Results on NASA projects are shown. Several well known algorithms from the static analysis literature are then explained (such as pointer analyses, memory analyses, weak relational abstract domains, function summarization, etc.). It ends with interesting problems we encountered (such as C++ analysis with exception handling, or the detection of integer overflow).

Static Analysis↗

Divide and conquer: separating the two probabilities in seismic phase picking

There are two fundamental probabilities in the seismic phase picking process—the probability of the existence of a seismic phase (detection probability) and the probability associated with the phase arrival time estimation (timing probability). The nearly ubiquitous approach in developing deep learning phase picking models is to use a kernel, such as a truncated Gaussian, to mask the labelled phase arrival time and train a segmentation model. Once a model is trained, the times of the peaks in the output are taken as phase arrival times (picks), and the height of the peaks are taken as ‘probability’ of the picks. Here, we show that this ‘probability’ represents neither the detection nor the timing probability because this approach forces the output to follow the shape of the kernel. We introduce an approach using two models to estimate these two distinct probabilities. We use a binary classifier with a calibrated confidence to address the detection probability and a multiclass classifier to obtain a probability mass function to address the timing probability. This new approach can make the deep learning-based phase picking process more interpretable and provide options to logically control seismic monitoring workflows.

58 GEOSCIENCES↗

Definition of an auxiliary processor dedicated to real-time operating system kernels

In order to increase the efficiency of process control data processing, it is necessary to enhance the productivity of real time high level languages and to automate the task administration, because presently 60 percent or more of the applications are still programmed in assembly languages. This may be achieved by migrating apt functions for the support of process control oriented languages into the hardware, i.e., by new architectures. Whereas numerous high level languages have already been defined or realized, there are no investigations yet on hardware assisted implementation of real time features. The requirements to be fulfilled by languages and operating systems in hard real time environment are summarized. A comparison of the most prominent languages, viz. Ada, HAL/S, LTR, Pearl, as well as the real time extensions of FORTRAN and PL/1, reveals how existing languages meet these demands and which features still need to be incorporated to enable the development of reliable software with predictable program behavior, thus making it possible to carry out a technical safety approval. Accordingly, Pearl proved to be the closest match to the mentioned requirements.

Halang, Wolfgang A.↗

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING↗

Cooperative Data Sharing: Simple Support for Clusters of SMP Nodes

Libraries like PVM and MPI send typed messages to allow for heterogeneous cluster computing. Lower-level libraries, such as GAM, provide more efficient access to communication by removing the need to copy messages between the interface and user space in some cases. still lower-level interfaces, such as UNET, get right down to the hardware level to provide maximum performance. However, these are all still interfaces for passing messages from one process to another, and have limited utility in a shared-memory environment, due primarily to the fact that message passing is just another term for copying. This drawback is made more pertinent by today's hybrid architectures (e.g. clusters of SMPs), where it is difficult to know beforehand whether two communicating processes will share memory. As a result, even portable language tools (like HPF compilers) must either map all interprocess communication, into message passing with the accompanying performance degradation in shared memory environments, or they must check each communication at run-time and implement the shared-memory case separately for efficiency. Cooperative Data Sharing (CDS) is a single user-level API which abstracts all communication between processes into the sharing and access coordination of memory regions, in a model which might be described as "distributed shared messages" or "large-grain distributed shared memory". As a result, the user programs to a simple latency-tolerant abstract communication specification which can be mapped efficiently to either a shared-memory or message-passing based run-time system, depending upon the available architecture. Unlike some distributed shared memory interfaces, the user still has complete control over the assignment of data to processors, the forwarding of data to its next likely destination, and the queuing of data until it is needed, so even the relatively high latency present in clusters can be accomodated. CDS does not require special use of an MMU, which can add overhead to some DSM systems, and does not require an SPMD programming model. unlike some message-passing interfaces, CDS allows the user to implement efficient demand-driven applications where processes must "fight" over data, and does not perform copying if processes share memory and do not attempt concurrent writes. CDS also supports heterogeneous computing, dynamic process creation, handlers, and a very simple thread-arbitration mechanism. Additional support for array subsections is currently being considered. The CDS1 API, which forms the kernel of CDS, is built primarily upon only 2 communication primitives, one process initiation primitive, and some data translation (and marshalling) routines, memory allocation routines, and priority control routines. The entire current collection of 28 routines provides enough functionality to implement most (or all) of MPI 1 and 2, which has a much larger interface consisting of hundreds of routines. still, the API is small enough to consider integrating into standard os interfaces for handling inter-process communication in a network-independent way. This approach would also help to solve many of the problems plaguing other higher-level standards such as MPI and PVM which must, in some cases, "play OS" to adequately address progress and process control issues. The CDS2 API, a higher level of interface roughly equivalent in functionality to MPI and to be built entirely upon CDS1, is still being designed. It is intended to add support for the equivalent of communicators, reduction and other collective operations, process topologies, additional support for process creation, and some automatic memory management. CDS2 will not exactly match MPI, because the copy-free semantics of communication from CDS1 will be supported. CDS2 application programs will be free to carefully also use CDS1. CDS1 has been implemented on networks of workstations running unmodified Unix-based operating systems, using UDP/IP and vendor-supplied high- performance locks. Although its inter-node performance is currently unimpressive due to rudimentary implementation technique, it even now outperforms highly-optimized MPI implementation on intra-node communication due to its support for non-copy communication. The similarity of the CDS1 architecture to that of other projects such as UNET and TRAP suggests that the inter-node performance can be increased significantly to surpass MPI or PVM, and it may be possible to migrate some of its functionality to communication controllers.

DiNucci, David C.↗

GP Cosmology Surrogate v1.0

GP Cosmology Surrogate is a Python library for building and training a generalized multi-output Gaussian process (GP) framework of @takhtaganov2021cosmic. In this approach, the surrogate is constructed sequentially, guided by a Bayesian optimization acquisition function that targets reduction of emulation error in the regions most consistent with the observational data. This adaptive design concentrates computational resources where they have the greatest impact on inference accuracy. The library supports efficient training for separable GP kernels, which allows the use of Kronecker algebra to handle high-dimensional input spaces and large numbers of correlated outputs. This makes it well suited for applications such as modeling cosmological power spectra, large-scale physical simulations, and multi-output hyperparameter tuning. By combining scalable multi-output GP modeling with data-driven adaptive sampling, GPsurrogate enables parameter inference and optimization with substantially fewer simulations than conventional space-filling designs.

Lukic, Zarija [Lawrence Berkeley National Laborato↗

Tomography of the rho meson in the QCD instanton vacuum: Transverse momentum dependent parton distribution functions

We analyze the rho meson unpolarized and polarized transverse momentum dependent parton distribution functions (TMDPDFs) in the instanton liquid model (ILM). The corresponding TMDs in ILM are approximated by a constituent quark beam function in the leading Fock state multiplied by a rapidity-dependent factor resulting from the staple-shaped Wilson lines, for fixed longitudinal momentum, transverse separation, and rapidity. At the resolution of the ILM, all of the rho meson TMDs are symmetric in parton 𝑥 for fixed transverse momentum, and Gaussian-like in the transverse momentum for fixed parton 𝑥. The latter is a direct consequence of the profiling of the quark zero modes in the ILM. The evolved TMDs at higher rapidity using the Collins-Soper kernel, and higher resolution using the renormalization group, show substantial skewness towards low parton 𝑥.

Gluons↗

On comparing helioseismic two-dimensional inversion methods

We consider inversion techniques for investigating the structure and dynamics of the solar interior as functions of radius and latitude. In particular, we look at the problem of inferring the radial and latitudinal dependence of the Sun's internal rotation, using a fully two-dimensional least-squares inversion algorithm. Concepts such as averaging kernels, measures of resolution, and trade-off curves, which have previously been used in the one-dimensional case, are generalized to facilitate a comparison of two-dimensional methods. We investigate the weighting given to different modes and discuss the implications of this for observational strategies. As an illustration we use a mode set whose properties are similar to those expected for data from the GONG network.

Schou, J.↗

Shock Hugoniot calculations using on-the-fly machine learned force fields with ab initio accuracy

We present a framework for computing the shock Hugoniot using on-the-fly machine learned force field (MLFF) molecular dynamics simulations. In particular, we employ an MLFF model based on the kernel method and Bayesian linear regression to compute the free energy, atomic forces, and pressure, in conjunction with a linear regression model between the internal and free energies to compute the internal energy, with all training data generated from Kohn–Sham density functional theory (DFT). We verify the accuracy of the formalism by comparing the Hugoniot for carbon with recent Kohn–Sham DFT results in the literature. In so doing, we demonstrate that Kohn–Sham calculations for the Hugoniot can be accelerated by up to two orders of magnitude, while retaining ab initio accuracy. We apply this framework to calculate the Hugoniots of 14 materials in the FPEOS database, comprising 9 single elements and 5 compounds, between temperatures of 10 kK and 2 MK. We find good agreement with first principles results in the literature while providing tighter error bars. In addition, we confirm that the inter-element interaction in compounds decreases with temperature.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Recovering pointwise values of discontinuous data within spectral accuracy

The pointwise values of a function, f(x), can be accurately recovered either from its spectral or pseudospectral approximations, so that the accuracy solely depends on the local smoothness of f in the neighborhood of the point x. Most notably, given the equidistant function grid values, its intermediate point values are recovered within spectral accuracy, despite the possible presence of discontinuities scattered in the domain. (Recall that the usual spectral convergence rate decelerates otherwise to first order, throughout). To this end, a highly oscillatory smoothing kernel is employed in contrast to the more standard positive unit-mass mollifiers. In particular, post-processing of a stable Fourier method applied to hyperbolic equations with discontinuous data, recovers the exact solution modulo a spectrally small error. Numerical examples are presented.

Gottlieb, D.↗

Recovering pointwise values of discontinuous data within spectral accuracy

The pointwise values of a function, f(x), can be accurately recovered either from its spectral or pseudospectral approximations, so that the accuracy solely depends on the local smoothness of f in the neighborhood of the point x. Most notably, given the equidistant function grid values, its intermediate point values are recovered within spectral accuracy, despite the possible presence of discontinuities scattered in the domain. (Recall that the usual spectral convergence rate decelerates otherwise to first order, throughout). To this end, a highly oscillatory smoothing kernel is employed in contrast to the more standard positive unit-mass mollifiers. In particular, post-processing of a stable Fourier method applied to hyperbolic equations with discontinuous data, recovers the exact solution modulo a spectrally small error. Numerical examples are presented.

Gottlieb, D.↗

Sharp front tracking with geometric interface reconstruction

Here, this paper presents a novel sharp front-tracking method designed to address limitations in classical front-tracking approaches, specifically their reliance on smooth interpolation kernels and extended stencils for coupling the front and fluid mesh. In contrast, the proposed method employs exclusively sharp, localized interpolation and spreading kernels, restricting the coupling to the interfacial fluid cells–those containing the interface/front. This localized coupling is achieved by integrating a divergence-preserving velocity interpolation method with a piecewise parabolic interface calculation (PPIC) and a polyhedron intersection algorithm to compute the indicator function and local interface curvature. Surface tension is computed using the Continuum Surface Force (CSF) method, maintaining consistency with the sharp representation. Additionally, we propose an efficient local roughness smoothing implementation to account for surface mesh undulations, which is easily applicable to any triangulated surface mesh. Building on our previous work, the primary innovation of this study lies in the localization of the coupling for both the indicator function and surface tension calculations. By reducing the interface thickness on the fluid mesh to a single cell, as opposed to the 4–5 cell spans typical in classical methods, the proposed sharp front-tracking method achieves a highly localized and accurate representation of the interface. This sharper representation mitigates parasitic currents and improves force balancing, making it particularly suitable for scenarios where the interface plays a critical role, such as microfluidics, fluid-fluid interactions, and fluid-structure interactions. The proposed method is comprehensively validated and tested on canonical interfacial flow problems, including stationary and translating Laplace equilibria, oscillating droplets, and rising bubbles. The presented results demonstrate that the sharp front-tracking method significantly outperforms the classical approach in terms of accuracy, stability, and computational efficiency. Notably, parasitic currents are reduced by approximately two orders of magnitude and stable results are obtained for parameter ranges where classical front tracking fails to converge.

42 ENGINEERING↗

Aeroelastic Response of Nonlinear Wing Section by Functional Series Technique

This paper addresses the problem of the determination of the subcritical aeroelastic response and flutter instability of nonlinear two-dimensional lifting surfaces in an incompressible flow-field via indicial functions and Volterra series approach. The related aeroelastic governing equations are based upon the inclusion of structural and damping nonlinearities in plunging and pitching, of the linear unsteady aerodynamics and consideration of an arbitrary time-dependent external pressure pulse. Unsteady aeroelastic nonlinear kernels are determined, and based on these, frequency and time histories of the subcritical aeroelastic response are obtained, and in this context the influence of the considered nonlinearities is emphasized. Conclusions and results displaying the implications of the considered effects are supplied.

Silva, Walter A.↗

Aeroelastic Response of Nonlinear Wing Section By Functional Series Technique

This paper addresses the problem of the determination of the subcritical aeroelastic response and flutter instability of nonlinear two-dimensional lifting surfaces in an incompressible flow-field via indicial functions and Volterra series approach. The related aeroelastic governing equations are based upon the inclusion of structural and damping nonlinearities in plunging and pitching, of the linear unsteady aerodynamics and consideration of an arbitrary time-dependent external pressure pulse. Unsteady aeroelastic nonlinear kernels are determined, and based on these, frequency and time histories of the subcritical aeroelastic response are obtained, and in this context the influence of the considered nonlinearities is emphasized. Conclusions and results displaying the implications of the considered effects are supplied.

Marzocca, Piergiovanni↗

A Discrete Hankel Transform Approach to Nuclear Data Processing for Fusion Applications

This study introduces advancements to the numerical solutions employed in the processing of nuclear data for fusion applications. It leverages the convolution theorem and Fourier transform techniques to enhance computational efficiency and broaden applicability. Building upon a previously reported discrete Hankel transform approach for Doppler broadening, this work refines the solution of convolution integrals central to these applications. The methodology provides a general and unified framework for evaluating any convolution operation, regardless of whether the underlying problem involves temperature effects in nuclear reactions. The applicability to the nuclear data processing for fusion is demonstrated by deriving the convolution integrals for some of the fusion-related quantities. As before, the convolution operation utilizes a Gaussian-based kernel; however, the discrete Hankel transform of order $𝛼$ = $\frac{1}{2}$ is now applied to the forward Fourier transform of the nonkernel argument, rather than the inverse Fourier transform. This modification eliminates the need for the integration of the nonkernel, cross section–based function, which is a step that posed challenges for certain pointwise cross-section representations. It also removes the requirement for cross-section linearization. Optimized for graphics processing unit architectures, the approach significantly improves computational performance. These advancements are currently under evaluation as the foundation for the next-generation thermonuclear data file processing codes being developed at Lawrence Livermore National Laboratory.

Nuclear science and engineering↗