Search NASA⌕ Search

SEARCH · Search NASA

Results for “Kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Catalytic Microtube Rocket Igniter

Devices that generate both high energy and high temperature are required to ignite reliably the propellant mixtures in combustion chambers like those present in rockets and other combustion systems. This catalytic microtube rocket igniter generates these conditions with a small, catalysis-based torch. While traditional spark plug systems can require anywhere from 50 W to multiple kW of power in different applications, this system has demonstrated ignition at less than 25 W. Reactants are fed to the igniter from the same tanks that feed the reactants to the rest of the rocket or combustion system. While this specific igniter was originally designed for liquid methane and liquid oxygen rockets, it can be easily operated with gaseous propellants or modified for hydrogen use in commercial combustion devices. For the present cryogenic propellant rocket case, the main propellant tanks liquid oxygen and liquid methane, respectively are regulated and split into different systems for the individual stages of the rocket and igniter. As the catalyst requires a gas phase for reaction, either the stored boil-off of the tanks can be used directly or one stream each of fuel and oxidizer can go through a heat exchanger/vaporizer that turns the liquid propellants into a gaseous form. For commercial applications, where the reactants are stored as gases, the system is simplified. The resulting gas-phase streams of fuel and oxidizer are then further divided for the individual components of the igniter. One stream each of the fuel and oxidizer is introduced to a mixing bottle/apparatus where they are mixed to a fuel-rich composition with an O/F mass-based mixture ratio of under 1.0. This premixed flow then feeds into the catalytic microtube device. The total flow is on the order of 0.01 g/s. The microtube device is composed of a pair of sub-millimeter diameter platinum tubes connected only at the outlet so that the two outlet flows are parallel to each other. The tubes are each approximately 10 cm long and are heated via direct electric resistive heating. This heating brings the gasses to their minimum required ignition temperature, which is lower than the auto-thermal ignition temperature, and causes the onset of both surface and gas phase ignition producing hot temperatures and a highly reacting flame. The combustion products from the catalytic tubes, which are below the melting point of platinum, are injected into the center of another combustion stage, called the primary augmenter. The reactants for this combustion stage come from the same source but the flows of non-premixed methane and oxygen gas are split off to a secondary mixing apparatus and can be mixed in a near-stoichiometric to highly lean mixture ratio. The primary augmenter is a component that has channels venting this mixed gas to impinge on each other in the center of the augmenter, perpendicular to the flow from the catalyst. The total crosssectional area of these channels is on a similar order as that of the catalyst. The augmenter has internal channels that act as a manifold to distribute equally the gas to the inward-venting channels. This stage creates a stable flame kernel as its flows, which are on the order of 0.01 g/s, are ignited by the combustion products of the catalyst. This stage is designed to produce combustion products in the flame kernel that exceed the autothermal ignition temperature of oxygen and methane.

Schneider, Steven J.↗

Parametric Deformation of Discrete Geometry for Aerodynamic Shape Design

We present a versatile discrete geometry manipulation platform for aerospace vehicle shape optimization. The platform is based on the geometry kernel of an open-source modeling tool called Blender and offers access to four parametric deformation techniques: lattice, cage-based, skeletal, and direct manipulation. Custom deformation methods are implemented as plugins, and the kernel is controlled through a scripting interface. Surface sensitivities are provided to support gradient-based optimization. The platform architecture allows the use of geometry pipelines, where multiple modelers are used in sequence, enabling manipulation difficult or impossible to achieve with a constructive modeler or deformer alone. We implement an intuitive custom deformation method in which a set of surface points serve as the design variables and user-specified constraints are intrinsically satisfied. We test our geometry platform on several design examples using an aerodynamic design framework based on Cartesian grids. We examine inverse airfoil design and shape matching and perform lift-constrained drag minimization on an airfoil with thickness constraints. A transport wing-fuselage integration problem demonstrates the approach in 3D. In a final example, our platform is pipelined with a constructive modeler to parabolically sweep a wingtip while applying a 1-G loading deformation across the wingspan. This work is an important first step towards the larger goal of leveraging the investment of the graphics industry to improve the state-of-the-art in aerospace geometry tools.

Anderson, George R.↗

SBUV version 8.6 Retrieval Algorithm: Error Analysis and Validation Technique

SBUV version 8.6 algorithm was used to reprocess data from the Back Scattered Ultra Violet (BUV), the Solar Back Scattered Ultra Violet (SBUV) and a number of SBUV/2 instruments, which 'span a 41-year period from 1970 to 2011 (except a 5-year gap in the 1970s)[see Bhartia et al, 2012]. In the new version Daumont et al. [1992] ozone cross section were used, and new ozone [McPeters et ai, 2007] and cloud climatologies Doiner and Bhartia, 1995] were implemented. The algorithm uses the Optimum Estimation technique [Rodgers, 2000] to retrieve ozone profiles as ozone layer (partial column, DU) on 21 pressure layers. The corresponding total ozone values are calculated by summing ozone columns at individual layers. The algorithm is optimized to accurately retrieve monthly zonal mean (mzm) profiles rather than an individual profile, since it uses monthly zonal mean ozone climatology as the A Priori. Thus, the SBUV version 8.6 ozone dataset is better suited for long-term trend analysis and monitoring ozone changes rather than for studying short-term ozone variability. Here we discuss some characteristics of the SBUV algorithm and sources of error in the SBUV profile and total ozone retrievals. For the first time the Averaging Kernels, smoothing errors and weighting functions (or Jacobians) are included in the SBUV metadata. The Averaging Kernels (AK) represent the sensitivity of the retrieved profile to the true state and contain valuable information about the retrieval algorithm, such as Vertical Resolution, Degrees of Freedom for Signals (DFS) and Retrieval Efficiency [Rodgers, 2000]. Analysis of AK for mzm ozone profiles shows that the total number of DFS for ozone profiles varies from 4.4 to 5.5 out of 6-9 wavelengths used for retrieval. The number of wavelengths in turn depends on solar zenith angles. Between 25 and 0.5 hPa, where SBUV vertical resolution is the highest, DFS for individual layers are about 0.5.

Kramarova, N. A.↗

Design Evolutuion of Hot Isotatic Press Cans for NTP Cermet Fuel Fabrication

Nuclear Thermal Propulsion (NTP) is under consideration for potential use in deep space exploration missions due to desirable performance properties such as a high specific impulse (> 850 seconds). Tungsten (W)-60vol%UO2 cermet fuel elements are under development, with efforts emphasizing fabrication, performance testing and process optimization to meet NTP service life requirements [1]. Fuel elements incorporate design features that provide redundant protection from crack initiation, crack propagation potentially resulting in hot hydrogen (H2) reduction of UO2 kernels. Fuel erosion and fission product retention barriers include W coated UO2 fuel kernels, W clad internal flow channels and fuel element external W clad resulting in a fully encapsulated fuel element design as shown.

Mireles, O. R.↗

Collaborative WorkBench for Researchers - Work Smarter, Not Harder

It is important to define some commonly used terminology related to collaboration to facilitate clarity in later discussions. We define provisioning as infrastructure capabilities such as computation, storage, data, and tools provided by some agency or similarly trusted institution. Sharing is defined as the process of exchanging data, programs, and knowledge among individuals (often strangers) and groups. Collaboration is a specialized case of sharing. In collaboration, sharing with others (usually known colleagues) is done in pursuit of a common scientific goal or objective. Collaboration entails more dynamic and frequent interactions and can occur at different speeds. Synchronous collaboration occurs in real time such as editing a shared document on the fly, chatting, video conference, etc., and typically requires a peer-to-peer connection. Asynchronous collaboration is episodic in nature based on a push-pull model. Examples of asynchronous collaboration include email exchanges, blogging, repositories, etc. The purpose of a workbench is to provide a customizable framework for different applications. Since the workbench will be common to all the customized tools, it promotes building modular functionality that can be used and reused by multiple tools. The objective of our Collaborative Workbench (CWB) is thus to create such an open and extensible framework for the Earth Science community via a set of plug-ins. Our CWB is based on the Eclipse [2] Integrated Development Environment (IDE), which is designed as a small kernel containing a plug-in loader for hundreds of plug-ins. The kernel itself is an implementation of a known specification to provide an environment for the plug-ins to execute. This design enables modularity, where discrete chunks of functionality can be reused to build new applications. The minimal set of plug-ins necessary to create a client application is called the Eclipse Rich Client Platform (RCP) [3]; The Eclipse RCP also supports thousands of community-contributed plug-ins, making it a popular development platform for many diverse applications including the Science Activity Planner developed at JPL for the Mars rovers [4] and the scientific experiment tool Gumtree [5]. By leveraging the Eclipse RCP to provide an open, extensible framework, a CWB supports customizations via plug-ins to build rich user applications specific for Earth Science. More importantly, CWB plug-ins can be used by existing science tools built off Eclipse such as IDL or PyDev to provide seamless collaboration functionalities.

Ramachandran, Rahul↗

Comptonization in Ultra-Strong Magnetic Fields: Numerical Solution to the Radiative Transfer Problem

We consider the radiative transfer problem in a plane-parallel slab of thermal electrons in the presence of an ultra-strong magnetic field (B approximately greater than B(sub c) approx. = 4.4 x 10(exp 13) G). Under these conditions, the magnetic field behaves like a birefringent medium for the propagating photons, and the electromagnetic radiation is split into two polarization modes, ordinary and extraordinary, that have different cross-sections. When the optical depth of the slab is large, the ordinary-mode photons are strongly Comptonized and the photon field is dominated by an isotropic component. Aims. The radiative transfer problem in strong magnetic fields presents many mathematical issues and analytical or numerical solutions can be obtained only under some given approximations. We investigate this problem both from the analytical and numerical point of view, provide a test of the previous analytical estimates, and extend these results with numerical techniques. Methods. We consider here the case of low temperature black-body photons propagating in a sub-relativistic temperature plasma, which allows us to deal with a semi-Fokker-Planck approximation of the radiative transfer equation. The problem can then be treated with the variable separation method, and we use a numerical technique to find solutions to the eigenvalue problem in the case of a singular kernel of the space operator. The singularity of the space kernel is the result of the strong angular dependence of the electron cross-section in the presence of a strong magnetic field. Results. We provide the numerical solution obtained for eigenvalues and eigenfunctions of the space operator, and the emerging Comptonization spectrum of the ordinary-mode photons for any eigenvalue of the space equation and for energies significantly lesser than the cyclotron energy, which is on the order of MeV for the intensity of the magnetic field here considered. Conclusions. We derived the specific intensity of the ordinary photons, under the approximation of large angle and large optical depth. These assumptions allow the equation to be treated using a diffusion-like approximation.

acceleration of particles↗

SpF: Enabling Petascale Performance for Pseudospectral Dynamo Models

Pseudospectral (PS) methods possess a number of characteristics (e.g., efficiency, accuracy, natural boundary conditions) that are extremely desirable for dynamo models. Unfortunately, dynamo models based upon PS methods face a number of daunting challenges, which include exposing additional parallelism, leveraging hardware accelerators, exploiting hybrid parallelism, and improving the scalability of global memory transposes. Although these issues are a concern for most models, solutions for PS methods tend to require far more pervasive changes to underlying data and control structures. Further, improvements in performance in one model are difficult to transfer to other models, resulting in significant duplication of effort across the research community.We have developed an extensible software framework for pseudospectral methods called SpF that is intended to enable extreme scalability and optimal performance. High-level abstractions provided by SpF unburden applications of the responsibility of managing domain decomposition and load balance while reducing the changes in code required to adapt to new computing architectures. The key design concept in SpF is that each phase of the numerical calculation is partitioned into disjoint numerical kernels that can be performed entirely in-processor. The granularity of domain-decomposition provided by SpF is only constrained by the data-locality requirements of these kernels. SpF builds on top of optimized vendor libraries for common numerical operations such as transforms, matrix solvers, etc., but can also be configured to use open source alternatives for portability. SpF includes several alternative schemes for global data redistribution and is expected to serve as an ideal testbed for further research into optimal approaches for different network architectures.In this presentation, we will describe the basic architecture of SpF as well as preliminary performance data and experience with adapting legacy dynamo codes. We will conclude with a discussion of planned extensions to SpF that will provide pseudospectral applications with additional flexibility with regard to time integration, linear solvers, and discretization in the radial direction.

Pseudospectral (PS)↗

Using SpF to Achieve Petascale for Legacy Pseudospectral Applications

Pseudospectral (PS) methods possess a number of characteristics (e.g., efficiency, accuracy, natural boundary conditions) that are extremely desirable for dynamo models. Unfortunately, dynamo models based upon PS methods face a number of daunting challenges, which include exposing additional parallelism, leveraging hardware accelerators, exploiting hybrid parallelism, and improving the scalability of global memory transposes. Although these issues are a concern for most models, solutions for PS methods tend to require far more pervasive changes to underlying data and control structures. Further, improvements in performance in one model are difficult to transfer to other models, resulting in significant duplication of effort across the research community. We have developed an extensible software framework for pseudospectral methods called SpF that is intended to enable extreme scalability and optimal performance. Highlevel abstractions provided by SpF unburden applications of the responsibility of managing domain decomposition and load balance while reducing the changes in code required to adapt to new computing architectures. The key design concept in SpF is that each phase of the numerical calculation is partitioned into disjoint numerical kernels that can be performed entirely inprocessor. The granularity of domain decomposition provided by SpF is only constrained by the datalocality requirements of these kernels. SpF builds on top of optimized vendor libraries for common numerical operations such as transforms, matrix solvers, etc., but can also be configured to use open source alternatives for portability. SpF includes several alternative schemes for global data redistribution and is expected to serve as an ideal testbed for further research into optimal approaches for different network architectures. In this presentation, we will describe our experience in porting legacy pseudospectral models, MoSST and DYNAMO, to use SpF as well as present preliminary performance results provided by the improved scalability.

DYNAMO↗

Surface Reflectance Product from Geostationary Satellite

We have generated provisional Himawari-8 AHI surface reflectance (SR) product for land and vegetation monitoring. The Himawari-8 AHI surface reflectance product is part of our GeoNEX land products, which integrate level 2 and higher remote sensing data from a set of geostationary satellite sensors (i.e. GOES-16, -17 ABI, Himawari-8 AHI, FY4-A AGRI, and MTG-I). Adapted Multiangle Implementation of Atmospheric Correction (MAIAC) algorithm is used to process time series Himawari-8 AHI observations. Himawari-8 AHI SR provides gridded and tiled land SR in 1-km resolution with high frequency (every 10 minutes during daylight time). There are three subdatasets: 1) retrieved atmospheric properties (e.g. column water vapor at 0.86 m, aerosol optical depth at 0.47m and 0.51m); 2) spectral (AHI bands 1-6) surface reflectance, kernels of RTLS BRDF model; 3)spectral BRDF kernel weights, and extensive quality assurance flags. The evaluation results show that Himawari-8 AHI data yield much more valid pixels in a single day in the characterization of land surface, when compare to NASA flagship satellite MODIS Terra/Aqua. This observation frequency and resolution of geostationary data should allow for using continuous ecosystem monitoring in diurnal studies at continental scale. Initial evaluations indicate a stable Himawari-8 AHI land SR product.

Li, Shuang↗

A Radiometric Consistent Spectral Fingerprinting Algorithm for Continuity Products of Hyperspectral Sounders

A radiometric consistent climate fingerprinting methodology has been developed to derive long-term temperature, water vapor, cloud, trace gases, and surface skin temperature anomaly time series from the hyper-spectral sounder measurements of multiple platforms. The spectral fingerprinting methodology requires the use of radiative kernels that are radiometrically consistent with observations. Radiative kernels are built using space-time averaged Jacobians that are physically retrieved from observations under all sky conditions. The physical retrieval algorithm uses the Principal Component based Radiative Transfer Model (PCRTM) for the forward simulation. The incorporation of multiple scattering simulation in PCRTM allows the direct radiative relationship between single field-of-view (FOV) radiance observations and corresponding thermal dynamic variables including cloud properties to be established. Therefore, radiance ?closure? can be achieved under all-sky conditions by the fingerprinting scheme. This methodology has been used to derive climate anomalies from the space-time averaged spectra of AIRS/AMSU and CrIS/ATMS. The use of a consistent fingerprinting scheme provides an effective mean of generating continuity product by merging observations from different platforms and therefore facilitating the long-term climate trend study.

Wan Wu↗

Uncertainty in Observational Estimates of the Aerosol Direct Radiative Effect and Forcing

Aerosols continue to be responsible for the largest uncertainty in determining the anthropogenic radiative forcing of the climate. To both reconcile the large range in satellite-based estimates of the aerosol direct radiative effect (DRE, the direct interaction with solar radiation by all aerosols) and to optimize the design of future observing systems, we build a framework for assessing uncertainty in aerosol DRE and the aerosol direct radiative forcing (DRF, the radiative effect of just anthropogenic aerosols, RF_ari). Shortwave aerosol radiative kernels (Jacobians) were derived using the MERRA-2 reanalysis data. These radiative kernels are used to compute a lower-bound on the systematic uncertainty in observational estimates of the aerosol DRE/DRF by making the optimistic assumption that global aerosol observations can be made with the accuracy found in the Aerosol Robotic Network (AERONET) sun photometer retrievals. The total uncertainty is shown to be dominated by contributions from the aerosol single scattering albedo uncertainty. These uncertainty estimates were compared to a literature survey of mostly satellite-based aerosol DRE/DRF values. Comparisons to previous studies reveal that most have significantly underestimated the aerosol DRE uncertainty. Past estimates of the aerosol DRF uncertainty are smaller (on average) than our optimistic observational estimates, including the aerosol DRF uncertainty given in the Intergovernmental Panel on Climate Change (IPCC) fifth assessment report (AR5).

Tyler James Thorsen↗

Uncertainty in Observational Estimates of the Aerosol Direct Radiative Effect and Forcing

Aerosols continue to be responsible for the largest uncertainty in determining the anthropogenic radiative forcing of the climate. To both reconcile the large range in satellite-based estimates of the aerosol direct radiative effect (DRE, the direct interaction with solar radiation by all aerosols) and to optimize the design of future observing systems, we build a framework for assessing uncertainty in aerosol DRE and the aerosol direct radiative forcing (DRF, the radiative effect of just anthropogenic aerosols, RF_ari). Shortwave aerosol radiative kernels (Jacobians) were derived using the MERRA-2 reanalysis data. These radiative kernels are used to compute a lower-bound on the systematic uncertainty in observational estimates of the aerosol DRE/DRF by making the optimistic assumption that global aerosol observations can be made with the accuracy found in the Aerosol Robotic Network (AERONET) sun photometer retrievals. The total uncertainty is shown to be dominated by contributions from the aerosol single scattering albedo uncertainty. These uncertainty estimates were compared to a literature survey of mostly satellite-based aerosol DRE/DRF values. Comparisons to previous studies reveal that most have significantly underestimated the aerosol DRE uncertainty. Past estimates of the aerosol DRF uncertainty are smaller (on average) than our optimistic observational estimates, including the aerosol DRF uncertainty given in the Intergovernmental Panel on Climate Change (IPCC) fifth assessment report (AR5).

Tyler James Thorsen↗

NASA and Blue Origin Collaborative Assessment of Precision Landing Algorithms and Computing

NASA’s Safe and Precise Landing Integrated Capabilities Evolution (SPLICE) project is developing sensor, algorithm, and compute technologies for precision landing and hazard avoidance. These technologies are being tested as an integrated Precision Landing and Hazard Avoidance (PL&HA) system on Blue Origin’s New Shephard suborbital vehicle. A key goal for the computing element of this technology development is to characterize the performance of the SPLICE software workloads on the project’s Descent and Landing Computer (DLC). The DLC is a multi-core processor designed as a surrogate for NASA’s High-Performance Space Computer (HPSC). Measurements of the SPLICE workload performance on the DLC provides NASA insight on how PL&HA capabilities will perform on the HPSC, and guidance on how the SPLICE algorithms can be implemented to best utilize the DLC platform. This insight can also be used to derive requirements to guide trade studies on candidate computing architectures, for use on platforms like Blue Moon. NASA and Blue Origin are collaborating under an agreement to pursue this mutual benefit. Performance metrics collected are based on measurement of common compute resources such as percentage used of memory bandwidth, I/O utilization, interrupt latency, and kernel vs. user space code residency. Where possible existing performance counters and metrics that are part of the operating system kernel are used. As the design has a significant FPGA component, performance counters are identified and instantiated in the fabric to measure DMA performance and interface metrics. Collection of metrics is performed on the DLC with a representative workload that simulates a full landing cycle of the Blue Origin New Shepard vehicle. Consideration is given to the other compute implementations and whether they can run SPLICE algorithms at the same rate and with the same latency as the DLC. One option being considered is the use of a RISC-V soft core instantiated in a radiation resilient FPGA fabric such as the Xilinx KU60. Select algorithms from the SPLICE code will be run for comparison with the DLC. This paper describes how the DLC is instrumented to collect performance measurements of the SPLICE workloads, preliminary results from these measurements, and their implications on SPLICE algorithm implementation. The results of experimentation to derive candidate requirements for architecture trades on a PL&HA computing system are also presented.

computer performance↗

WebGeocalc and Cosmographia: Modern Tools to Access OPS SPICE Data

For more than two decades navigation and other ancillary data from most US and international planetary science missions have been packaged using "SPICE" (Spacecraft, Planet, Instrument, Camera-matrix, Events) system data files (a.k.a. SPICE kernels) and, in conjunction with SPICE Toolkit software used by scientists and engineers to compute observation geometry in various ground system tools ranging from mission planning and analysis applications to data production pipelines to science analysis tools. The traditional way for accessing SPICE data is by downloading necessary SPICE kernels to a user’s workstation, installing the SPICE Toolkit software available from NAIF, and writing an application calling APIs from the SPICE Toolkit library to compute numeric geometric parameters of interest. While this approach did and still does provide the greatest flexibility in implementing geometric computations of interest, it proved to be complicated for users with little programming abilities, required data to be always copied to the users’ workstations, and lacked any out-of-the-box visualization capabilities. To address these shortcomings NAIF developed the WebGeocalc (WGC) tool and extended the publicly available Cosmographia program to use SPICE. Employing these two new tools in mission operations enables easier access to SPICE computations and SPICE-based visualizations for a wider variety of mission personnel.

Semenov, Boris V.↗

Computing observation geometry for small satellites

Most solar system science missions need a variety of observation geometry–quantities such as position and velocity, range and altitude, viewing latitude and longitude, and lighting angles– to support mission engineering, science planning, and science data analysis activities. NASA's "SPICE" system offers one popular, multi-mission means for doing just that. SPICE comprises both data files, called kernels, and a SPICE software Toolkit that is available in many popular languages. A mission operations center produces the SPICE kernel files. Scientists and engineers write their own applications programs to address some need, and they include a few SPICE subroutines within that code to do the needed geometry computations. The SPICE system has been in use throughout NASA’s planetary science mission domain since 1991, and it has slowly spread to most major space agencies around the globe since then. The SPICE software is available in most popular languages, and for most popular platforms. The code is thoroughly tested before being released, and new versions of the Toolkit are always backwards compatible. The SPICE components are freely offered to everyone, and have no export, licensing or similar restrictions. Maybe using SPICE would work for your CubeSat or SmallSat mission?

Acton, Charles H.↗

Attribution of Chemistry-Climate Model Initiative (CCMI) Ozone Radiative Flux Bias from Satellites

The top-of-atmosphere (TOA) outgoing longwave flux over the 9.6-μm ozone band is a fundamental quantity for understanding chemistry-climate coupling. However, observed TOA fluxes are hard to estimate as they exhibit considerable variability in space and time that depend on the distributions of clouds, ozone (O3), water vapor (H2O), air temperature (Ta), and surface temperature (Ts). Benchmarking present day fluxes and quantifying the relative influence of their drivers is the first step for estimating climate feedbacks from ozone radiative forcing and predicting radiative forcing evolution. To that end, we constructed observational instantaneous radiative kernels (IRKs) under clear-sky conditions, representing the sensitivities of the TOA flux in the 9.6-μm ozone band to the vertical distribution of geophysical variables, including O3, H2O, Ta, and Ts based upon the Aura Tropospheric Emission Spectrometer (TES) measurements. Applying these kernels to present-day simulations from the Chemistry-Climate Model Initiative (CCMI) project as compared to a 2006 reanalysis assimilating satellite observations, we show that the models have large differences in TOA flux, attributable to different geophysical variables. In particular, model simulations continue to diverge from observations in the tropics, as reported in previous studies of the Atmospheric Chemistry Climate Model Inter-comparison Project (ACCMIP) simulations. The principal culprits are tropical mid and upper tropospheric ozone followed by tropical lower tropospheric H2O. Five models out of the eight studied here have TOA flux biases exceeding 100 mWm-2 attributable to tropospheric ozone bias. Another set of five models have flux biases over 50 mWm-2 due to H2O. On the other hand, Ta radiative bias is negligible in all models (no more than 30 mWm-2). We found that AM3 and CMAM have the lowest TOA flux biases globally but are a result of cancellation of opposite biases due to difference processes. Overall, the multi-model ensemble mean bias is –133±98 mWm-2, indicating that they are too atmospherically opaque due to trapping too much radiation in the atmosphere by overestimated tropical tropospheric O3 and H2O. Having too much O3 and H2O in the troposphere would have different impacts on the sensitivity of TOA flux to O3 and these competing effects add more uncertainties on the ozone radiative forcing. We find that the inter-model TOA outgoing longwave radiation (OLR) difference is well anti-correlated with their ozone band flux bias. This suggests that there is significant radiative compensation in the calculation of model outgoing longwave radiation.

Aura Tropospheric Emission Spectrometer (TES) meas↗

Memory Optimizations for Sparse Linear Algebra on GPU Hardware

An effort to maximize memory bandwidth utilization for a sparse linear algebra kernel executing on NVIDIA® Tesla V100 and A100 Graphics Processing Units (GPUs) is described. The kernel consists of a block-sparse matrix-vector product and a series of forward/backward triangular solves. The computation is memory-bound and exhibits low arithmetic intensity. Along with a relatively small block size, the data layout poses a challenge to effectively utilize the available memory bandwidth on common GPU architectures. An earlier implementation using a warp to process a single row of the matrix was found to yield good memory performance on the V100 architecture. However, anew approach, which assigns a warp to six rows of the matrix, is proposed for the A100. In addition, two new features offered by the A100 architecture are explored.L2residency control enables a portion of theL2cache to be used for persistent data access, and the asynchronous copy instruction allows data to be loaded directly from main memory into shared memory. Demonstrations show that the new implementation improves memory bandwidth utilization from 71.5% to 81.2% of the peak available on theA100 architecture.

GPU↗

Assessment of Edge-Based Viscous Method for Corner-Flow Solutions on Graphics Processing Units

A highly efficient, edge-based viscous (EBV) discretization method has been recently implemented in a practical, unstructured-grid, node-centered, finite-volume flow solver and evaluated for Reynolds-averaged Navier-Stokes (RANS) formulations. In comparison to a well-established cell-based viscous (CBV) method, the EBV method has demonstrated multifold acceleration of all viscous-kernel computations on general unstructured mixed-element grids. The viscous kernels include evaluation of viscous fluxes, diffusion terms in turbulence models, and the corresponding Jacobian terms. In this paper, an EBV implementation of a nonlinear extension of the Spalart-Allmaras turbulence model, SA-neg-QCR2000, is presented and verified. The SA-neg-QCR2000 model is used for simulating turbulent corner flows. Previously reported EBV computations have been conducted on traditional computing architectures based on central processing units (CPU). This paper assesses benefits of the EBV method on modern high-performance computing architectures based on graphics processing units (GPU). The GPU implementations of the CBV and EBV methods are verified by comparing solutions and iterative convergence with those observed in CPU computations on the same grids. A comprehensive assessment of the EBV speedup on CPU and GPU architectures is presented for established benchmark corner flows, namely, a supersonic flow through a long square duct and a subsonic flow around a NASA juncture flow model.

CFD↗