Search NASA⌕ Search

SEARCH · Search NASA

Results for “Partitioned scheme”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Improved Convective Ice Microphysics Parameterization in the NCAR CAM Model

Partitioning deep convective cloud condensates into components that sediment and detrain, known to be a challenge for global climate models, is important for cloud vertical distribution and anvil cloud formation. In this study, we address this issue by improving the convective microphysics scheme in the National Center for Atmospheric Research Community Atmosphere Model version 5.3 (CAM5.3). The improvements include: (1) considering sedimentation for cloud ice crystals that do not fall in the original scheme, (2) applying a new terminal velocity parameterization that depends on the environmental conditions for convective snow, (3) adding a new hydrometeor category, “rimed ice,” to the original four-class (cloud liquid, cloud ice, rain, and snow) scheme, and (4) allowing convective clouds to detrain snow particles into stratiform clouds. Results from the default and modified CAM5.3 models were evaluated against observations from the U.S. Department of Energy Tropical Warm Pool-International Cloud Experiment (TWP-ICE) field campaign. The default model overestimates ice amount, which is largely attributed to the underestimation of convective ice particle sedimentation. By considering cloud ice sedimentation and rimed ice particles and applying a new convective snow terminal velocity parameterization, the vertical distribution of ice amount is much improved in the midtroposphere and upper troposphere when compared to observations. The vertical distribution of ice condensate also agrees well with observational best estimates upon considering snow detrainment. Comparison with observed convective updrafts reveals that current bulk model fails to reproduce the observed updraft magnitude and occurrence frequency, suggesting spectral distributions be required to simulate the subgrid updraft heterogeneity.

terminal velocity↗

Partition-based Feasible Integer Solution Pre-computation for Hybrid Model Predictive Control

For multiparametric mixed-integer convex programming problems such as those encountered in hybrid model predictive control, we propose an algorithm for generating a feasible partition of a subset of the parameter space. The result is a static map from the current parameter to a suboptimal integer solution such that the remaining convex program is feasible. Convergence is proved with a new insight that the overlap among the feasible parameter sets of each integer solution governs the partition complexity. The partition is stored as a tree which makes querying the feasible solution efficient. The algorithm can be used to warm start a mixed integer solver with a real-time guarantee or to provide a reference integer solution in several suboptimal MPC schemes. The algorithm is tested on randomly generated systems with up to six states, demonstrating the effectiveness of the approach.

Bayard, David S.↗

Parallelized solvers for heat conduction formulations

Based on multilevel partitioning, this paper develops a structural parallelizable solution methodology that enables a significant reduction in computational effort and memory requirements for very large scale linear and nonlinear steady and transient thermal (heat conduction) models. Due to the generality of the formulation of the scheme, both finite element and finite difference simulations can be treated. Diverse model topologies can thus be handled, including both simply and multiply connected (branched/perforated) geometries. To verify the methodology, analytical and numerical benchmark trends are verified in both sequential and parallel computer environments.

Padovan, Joe↗

A Partitioned -Task Parallel Implementation of the NASA Multiscale Analysis Tool for High Performance Computing

The NASA Multiscale Analysis Tool (NASMAT) is a “plug and play” software package that allows users to conduct massively multiscale modeling of hierarchical and nonlinear materials. This work extends the scalability and improves the High Performance Computing friendliness of NASMAT by adopting a Partitioned Task-Parallel approach. Interoperability of NASMAT with external software is enhanced through preCICE, a open source library for multiphysics coupling in a partitioned manner. Enhancement through preCICE allows for easy integration of NASMAT to other macro solvers and dissociates the parallelization strategy adopted within NASMAT from the macro solver. The task-parallel framework based on Master-Worker approach is implemented as the parallelization scheme. The scheme accounts for hierarchy of multiple scales (task-dependence) and heterogeneous nature (dynamic load balancing) of computations. The applicability and scalability of the framework will be evaluated by analyzing large scale engineering problems through massively multiscale methods.

NASMAT↗

Additive Runge-Kutta Schemes for Convection-Diffusion-Reaction Equations

Additive Runge-Kutta (ARK) methods are investigated for application to the spatially discretized one- dimensional convection-diffusion-reaction (CDR) equations. Accuracy, stability, conservation, and dense-output are first considered for the general case when N different Runge-Kutta methods are grouped into a single composite method. Then, implicit-explicit, (N = 2), additive Runge-Kutta (ARK(sub 2)) methods from third- to fifth-order are presented that allow for integration of stiff terms by an L-stable, stiffly-accurate explicit, singly diagonally implicit Runge-Kutta (ESDIRK) method while the nonstiff terms are integrated with a traditional explicit Runge-Kutta method (ERK). Coupling error terms of the partitioned method are of equal order to those of the elemental methods. Derived ARK(sub 2) methods have vanishing stability functions for very large values of the stiff scaled eigenvalue, z['] yields -infinity, and retain high stability efficiency in the absence of stiffness, z['] yield 0. Extrapolation-type stage- value predictors are provided based on dense-output formulae. Optimized methods minimize both leading order ARK(sub 2) error terms and Butcher coefficient magnitudes as well as maximize conservation properties. Numerical tests of the new schemes on a CDR problem show negligible stiffness leakage and near classical order convergence rates. However, tests on three simple singular-perturbation problems reveal generally predictable order reduction. Error control is best managed with a PID-controller. While results for the fifth-order method are disappointing, both the new third- and fourth-order methods are at least as efficient as existing ARK(sub 2) methods.

Kennedy, Christopher A.↗

Anomaly inflow, dualities, and quantum simulation of Abelian lattice gauge theories induced by measurements

Previous work [] has demonstrated that quantum simulation of Abelian lattice gauge theories (Wegner models including the toric code in a limit) in general dimensions can be achieved by local adaptive measurements on symmetry-protected topological (SPT) states with higher-form generalized global symmetries. The entanglement structure of the resource SPT state reflects the geometric structure of the gauge theory. In this work we explicitly demonstrate the anomaly inflow mechanism between the deconfining phase of the simulated gauge theory on the boundary and the SPT state in the bulk by showing that the anomalous gauge variation of the boundary state obtained by bulk measurement matches that of the bulk theory. Moreover, we construct the resource state and the measurement pattern for the measurement-based quantum simulation of a lattice gauge theory with a matter field (Fradkin-Shenker model), where a simple scheme to protect gauge invariance of the simulated state against errors is proposed. We further consider taking an overlap between the wave function of the resource state for lattice gauge theories and that of a parameterized product state, and we derive precise dualities between partition functions with insertion of defects corresponding to gauging higher-form global symmetries, as well as measurement-induced phases where states induced by a partial overlap possess different (symmetry-protected) topological orders. Measurement-assisted operators to dualize quantum Hamiltonians of lattice gauge theories and their noninvertibility are also presented. Published by the American Physical Society 2024

Okuda, Takuya↗

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learn- ing on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15–40% improvement in end-to-end training performance on the NERSC Perlmutter supercomputer for various OGB datasets.

Machine Leanring, high performance comptuing, grap↗

Porting Gravitational Wave Signal Extraction to Parallel Virtual Machine (PVM)

Laser Interferometer Space Antenna (LISA) is a planned NASA-ESA mission to be launched around 2012. The Gravitational Wave detection is fundamentally the determination of frequency, source parameters, and waveform amplitude derived in a specific order from the interferometric time-series of the rotating LISA spacecrafts. The LISA Science Team has developed a Mock LISA Data Challenge intended to promote the testing of complicated nested search algorithms to detect the 100-1 millihertz frequency signals at amplitudes of 10E-21. However, it has become clear that, sequential search of the parameters is very time consuming and ultra-sensitive; hence, a new strategy has been developed. Parallelization of existing sequential search algorithms of Gravitational Wave signal identification consists of decomposing sequential search loops, beginning with outermost loops and working inward. In this process, the main challenge is to detect interdependencies among loops and partitioning the loops so as to preserve concurrency. Existing parallel programs are based upon either shared memory or distributed memory paradigms. In PVM, master and node programs are used to execute parallelization and process spawning. The PVM can handle process management and process addressing schemes using a virtual machine configuration. The task scheduling and the messaging and signaling can be implemented efficiently for the LISA Gravitational Wave search process using a master and 6 nodes. This approach is accomplished using a server that is available at NASA Ames Research Center, and has been dedicated to the LISA Data Challenge Competition. Historically, gravitational wave and source identification parameters have taken around 7 days in this dedicated single thread Linux based server. Using PVM approach, the parameter extraction problem can be reduced to within a day. The low frequency computation and a proxy signal-to-noise ratio are calculated in separate nodes that are controlled by the master using message and vector of data passing. The message passing among nodes follows a pattern of synchronous and asynchronous send-and-receive protocols. The communication model and the message buffers are allocated dynamically to address rapid search of gravitational wave source information in the Mock LISA data sets.

Thirumalainambi, Rajkumar↗

An Extension of the Time-Spectral Method to Overset Solvers

Relative motion in the Cartesian or overset framework causes certain spatial nodes to move in and out of the physical domain as they are dynamically blanked by moving solid bodies. This poses a problem for the conventional Time-Spectral approach, which expands the solution at every spatial node into a Fourier series spanning the period of motion. The proposed extension to the Time-Spectral method treats unblanked nodes in the conventional manner but expands the solution at dynamically blanked nodes in a basis of barycentric rational polynomials spanning partitions of contiguously defined temporal intervals. Rational polynomials avoid Runge's phenomenon on the equidistant time samples of these sub-periodic intervals. Fourier- and rational polynomial-based differentiation operators are used in tandem to provide a consistent hybrid Time-Spectral overset scheme capable of handling relative motion. The hybrid scheme is tested with a linear model problem and implemented within NASA's OVERFLOW Reynolds-averaged Navier- Stokes (RANS) solver. The hybrid Time-Spectral solver is then applied to inviscid and turbulent RANS cases of plunging and pitching airfoils and compared to time-accurate and experimental data. A limiter was applied in the turbulent case to avoid undershoots in the undamped turbulent eddy viscosity while maintaining accuracy. The hybrid scheme matches the performance of the conventional Time-Spectral method and converges to the time-accurate results with increased temporal resolution.

Leffell, Joshua Isaac↗

Computation of the inviscid supersonic flow about cones at large angles of attack by a floating discontinuity approach

The technique of floating shock fitting is adapted to the computation of the inviscid flowfield about circular cones in a supersonic free stream at angles of attack that exceed the cone half-angle. The resulting equations are applicable over the complete range of free-stream Mach numbers, angles of attack and cone half-angles for which the bow shock is attached. A finite difference algorithm is used to obtain the solution by an unsteady relaxation approach. The bow shock, embedded cross-flow shock, and vortical singularity in the leeward symmetry plane are treated as floating discontinuities in a fixed computational mesh. Where possible, the flowfield is partitioned into windward, shoulder, and leeward regions with each region computed separately to achieve maximum computational efficiency. An alternative shock fitting technique which treats the bow shock as a computational boundary is developed and compared with the floating-fitting approach. Several surface boundary condition schemes are also analyzed.

Daywitt, J.↗

A Vertically Resolved Canopy Improves Chemical Transport Model Predictions of Ozone Deposition to North Temperate Forests

Abstract Dry deposition is the second largest tropospheric ozone (O 3 ) sink and occurs through stomatal and nonstomatal pathways. Current O 3 uptake predictions are limited by the simplistic big‐leaf schemes commonly used in chemical transport models (CTMs) to parameterize deposition. Such schemes fail to reproduce observed O 3 fluxes over terrestrial ecosystems, highlighting the need for more realistic treatment of surface‐atmosphere exchange in CTMs. We address this need by linking a resolved canopy model (1D Multi‐Layer Canopy CHemistry and Exchange Model, MLC‐CHEM) to the GEOS‐Chem CTM and use this new framework to simulate O 3 fluxes over three north temperate forests. We compare results with in situ measurements from four field studies and with standalone, observationally constrained MLC‐CHEM runs to test current knowledge of O 3 deposition and its drivers. We show that GEOS‐Chem overpredicts observed O 3 fluxes across all four studies by up to 2×, whereas the resolved‐canopy models capture observed diel profiles of O 3 deposition and in‐canopy concentrations to within 10%. Relative humidity and solar irradiance are strong O 3 flux drivers over these forests, and uncertainties in those fields provide the largest remaining source of model deposition biases. Flux partitioning analysis shows that: (a) nonstomatal loss accounts for 60% of O 3 deposition on average; (b) in‐canopy chemistry makes only a small contribution to total O 3 fluxes; and (c) the CTM big‐leaf treatment overestimates O 3 ‐driven stomatal loss and plant phytotoxicity in these temperate forests by up to 7×. Results motivate the application of fully online vertically explicit canopy schemes in CTMs for improved O 3 predictions.

Vermeuel, Michael P. [Department of Soil, Water, a↗

Interplay Between Time and Energy in Bosonic Noisy Quantum Metrology

Quantum entanglement and coherence often allow for protocols that outperform classical ones in estimating a system’s parameter. When using infinite-dimensional probes (such as a bosonic mode), one could, in principle, obtain infinite precision in a finite time for both classical and quantum protocols, which makes it hard to quantify potential quantum advantage. However, such a situation is unphysical, as it would require infinite resources, so one needs to impose some additional constraint: typically the average energy employed by the probe is finite. Here we treat both energy and time as a resource, showing that, in the presence of noise, there is a nontrivial interplay between the average energy and the time devoted to the estimation. Our results are valid for the most general metrological schemes (e.g., adaptive schemes, which may involve entanglement with external ancillae or any kind of continuous measurement). We apply recently derived precision bounds for all parameters characterizing the paradigmatic case of a bosonic mode, subject to Lindbladian noise. We show how the time employed in the estimation should be partitioned in order to achieve the best possible precision. In most cases, the optimal performance may be obtained without the necessity of adaptivity or entanglement with ancilla. We compare results with classical strategies. Interestingly, for temperature estimation, applying a fast-prepare-and-measure protocol with Fock states provides better scaling with the number of photons than any classical strategy.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Effects of partitioning and scheduling sparse matrix factorization on communication and load balance

A block based, automatic partitioning and scheduling methodology is presented for sparse matrix factorization on distributed memory systems. Using experimental results, this technique is analyzed for communication and load imbalance overhead. To study the performance effects, these overheads were compared with those obtained from a straightforward 'wrap mapped' column assignment scheme. All experimental results were obtained using test sparse matrices from the Harwell-Boeing data set. The results show that there is a communication and load balance tradeoff. The block based method results in lower communication cost whereas the wrap mapped scheme gives better load balance.

Venugopal, Sesh↗

Hierarchical Parallelism in Finite Difference Analysis of Heat Conduction

Based on the concept of hierarchical parallelism, this research effort resulted in highly efficient parallel solution strategies for very large scale heat conduction problems. Overall, the method of hierarchical parallelism involves the partitioning of thermal models into several substructured levels wherein an optimal balance into various associated bandwidths is achieved. The details are described in this report. Overall, the report is organized into two parts. Part 1 describes the parallel modelling methodology and associated multilevel direct, iterative and mixed solution schemes. Part 2 establishes both the formal and computational properties of the scheme.

Padovan, Joseph↗

E-beam generated holographic masks for optical vector-matrix multiplication

An optical vector matrix multiplication scheme that encodes the matrix elements as a holographic mask consisting of linear diffraction gratings is proposed. The binary, chrome on glass masks are fabricated by e-beam lithography. This approach results in a fairly simple optical system that promises both large numerical range and high accuracy. A partitioned computer generated hologram mask was fabricated and tested. This hologram was diagonally separated outputs, compact facets and symmetry about the axis. The resultant diffraction pattern at the output plane is shown. Since the grating fringes are written at 45 deg relative to the facet boundaries, the many on-axis sidelobes from each output are seen to be diagonally separated from the adjacent output signals.

Arnold, S. M.↗

Out-of-Core Streamline Visualization on Large Unstructured Meshes

It's advantageous for computational scientists to have the capability to perform interactive visualization on their desktop workstations. For data on large unstructured meshes, this capability is not generally available. In particular, particle tracing on unstructured grids can result in a high percentage of non-contiguous memory accesses and therefore may perform very poorly with virtual memory paging schemes. The alternative of visualizing a lower resolution of the data degrades the original high-resolution calculations. This paper presents an out-of-core approach for interactive streamline construction on large unstructured tetrahedral meshes containing millions of elements. The out-of-core algorithm uses an octree to partition and restructure the raw data into subsets stored into disk files for fast data retrieval. A memory management policy tailored to the streamline calculations is used such that during the streamline construction only a very small amount of data are brought into the main memory on demand. By carefully scheduling computation and data fetching, the overhead of reading data from the disk is significantly reduced and good memory performance results. This out-of-core algorithm makes possible interactive streamline visualization of large unstructured-grid data sets on a single mid-range workstation with relatively low main-memory capacity: 5-20 megabytes. Our test results also show that this approach is much more efficient than relying on virtual memory and operating system's paging algorithms.

Ueng, Shyh-Kuang↗

Trellis-coded multidimensional phase modulation

A 2L-dimensional multiple phase-shift keyed (MPSK) (L x MPSK) signal set is obtained by forming the Cartesian product of L two-dimensional MPSK signal sets. A systematic approach to partitioning L x MPSK signal sets that is based on block coding is used. An encoder system approach is developed. It incorporates the design of a differential precoder, a systematic convolutional encoder, and a signal set mapper. Trellis-coded L x 4PSK, L x 8PSK, and L x 16PSK modulation schemes are found for L = 1-4 and a variety of code rates and decoder complexities, many of which are fully transparent to discrete phase rotations of the signal set. The new codes achieve asymptotic coding gains up to 5.85 dB.

Pietrobon, Steven S.↗

Hierarchial implicit dynamic least-square solution algorithm

This paper develops an implicit type transient solution strategy which possesses hierarchial levels of application. In particular, due to the manner of formulation, stiffness updating, assembly inversion, solution constraint, as well as iteration are all performed at a localized level. The level of iterative calculations depends on the type of hierarchial partitioning employed, namely degree of freedom, nodal, elemental, material/nonlinear group, substructural, and so on. Since the iterative solution process and application of constraints are applied at a local level, the resulting so-called hierarchial implicit solution algorithm possesses very stable and efficient numerical properties and is highly storage efficient. To demonstrate the scheme, the results of several benchmark examples are presented. These enable comparisons with the Newton-Raphson solved implicit transient solution method. Overall the comparisons illustrate the superior stability and efficiency of the hierarchial scheme.

Padovan, J.↗