Search NASA⌕ Search

SEARCH · Search NASA

Results for “Memory Optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Optimal Estimation of Rain Rate Profiles from Single-Frequency Radar Echoes

The significant ambiguities inherent in the determination of a particular vertical rain intensity profile from a given time profile of radar echo powers measured by a downward-looking (spaceborne or airborne) radar at a single attenuating frequency are well-documented. Indeed, one already knows that by appropriately varying the parameters of the reflectivity-rain-rate (Z - R) and/or attenuation-rain-rate (k - R) relationships, one can produce several substantially different hypothetical rain rate profiles which would have the same radar power profile. Imposing the additional constraint that the path-averaged rain-rate be a given fixed number does reduce the ambiguities but falls far short of eliminating them. While we now know how to generate as many mutually ambiguous rain-rate profiles from a given profile of received radar reflectivities as we like, there remains to produce a quantitative measure to assess how likely each of these profiles is, what the appropriate 'average' profile should be, and what the 'variance' of these multiple solutions is. Of course, in order to do this, one needs to spell out the stochastic constraints that can allow us to make sense of the words 'average' and 'variance' in a mathematically rigorous way. Such a quantitative approach would be particularly well-suited for such systems as the proposed Precipitation Radar of the Tropical Rainfall Measuring Mission (TRMM). Indeed, one would then be able to use the radar reflectivities measured by the TRMM radar from one particular look in order to estimate the most likely rain-rate profile that would have produced the measurements, as well as the uncertainty in the estimated rain-rates as a function of range. Such an optimal approach is described in this paper.

deep space flight computer real-time advanced spac↗

Advanced Engine Cycles Analyzed for Turbofans With Variable-Area Fan Nozzles Actuated by a Shape Memory Alloy

Advanced, large commercial turbofan engines using low-fan-pressure-ratio, very high bypass ratio thermodynamic cycles can offer significant fuel savings over engines currently in operation. Several technological challenges must be addressed, however, before these engines can be designed. To name a few, the high-diameter fans associated with these engines pose a significant packaging and aircraft installation challenge, and a large, heavy gearbox is often necessary to address the differences in ideal operating speeds between the fan and the low-pressure turbine. Also, the large nacelles contribute aerodynamic drag penalties and require long, heavy landing gear when mounted on conventional, low wing aircraft. Nevertheless, the reduced fuel consumption rates of these engines are a compelling economic incentive, and fans designed with low pressure ratios and low tip speeds offer attractive noise-reduction benefits. Another complication associated with low-pressure-ratio fans is their need for variable flow-path geometry. As the design fan pressure ratio is reduced below about 1.4, an operational disparity is set up in the fan between high and low flight speeds. In other words, between takeoff and cruise there is too large a swing in several key fan parameters-- such as speed, flow, and pressure--for a fan to accommodate. One solution to this problem is to make use of a variable-area fan nozzle (VAFN). However, conventional, hydraulically actuated variable nozzles have weight, cost, maintenance, and reliability issues that discourage their use with low-fan-pressure-ratio engine cycles. United Technologies Research, in cooperation with NASA, is developing a revolutionary, lightweight, and reliable shape memory alloy actuator system that can change the on-demand nozzle exit area by up to 20 percent. This "smart material" actuation technology, being studied under NASA's Ultra-Efficient Engine Technology (UEET) Program and Revolutionary Concepts in Aeronautics (RevCon) Program, has the potential to enable the next generation of efficient, quiet, very high bypass ratio turbofans. NASA Glenn Research Center's Propulsion Systems Analysis Office, along with NASA Langley Research Center's Systems Analysis Branch, conducted an independent analytical assessment of this new technology to provide strategic guidance to UEET and RevCon. A 2010-technology-level high-spool engine core was designed for this evaluation. Two families of low-spool components, one with and one without VAFN's, were designed to operate with the core. This "constant core" approach was used to hold most design parameters constant so that any performance differences between the VAFN and fixed nozzle cycles could be attributed to the VAFN technology alone. In this manner, the cycle design regimes that offer a performance payoff when VAFN's are used could be identified. The NASA analytical model of a performance-optimized VAFN turbofan with a fan pressure ratio of 1.28 is shown. Mission analyses of the engines were conducted using the notional, long-haul, advanced commercial twinjet shown. A high wing design was used to accommodate the large high-bypassratio engines. The mission fuel reduction benefit of very high bypass shape-memory-alloy VAFN aircraft was calculated to be 8.3 percent lower than a moderate bypass cycle using a conventional fixed nozzle. Shape-memory-alloy VAFN technology is currently under development in NASA's UEET and RevCon Programs.

Berton, Jeffrey J.↗

Capturing, Analyzing, Maintaining, and Disseminating Shape Memory Material Data Between Information Management Systems

With an increased demand on reducing the time, cost, and effort to develop new materials, Integrated Computational Materials Engineering (ICME) has received widespread attention in various engineering disciplines as a catalyst for significantly reducing experimental testing during the material design process. An ICME approach to design can enable ‘fit-for-purpose’ materials to be realized in engineering applications by incorporating well-understood process-property-performance relationships between the various length and time scales in a material’s structure, enabling material optimization. However, such an approach requires validated multiscale models at the various length scales for a material, which in turn requires a large amount of data, a robust means of storing the data, and the ability to link data to developed material models. The NASA Vision 2040 [1] has identified nine key elements to enabling ICME approaches in system level design, with one being “Data, Information, and Visualization”, thus outlining the importance of a robust information management system for ICME. As the relationship between microstructure, properties, and material performance become better understood and incorporated into multiscale models that can be leveraged in application design, the emergence of new materials with application-driven properties can be realized. One such new material class that has seen growing attention are shape memory materials (SMM), in which a material can transition between a deformed and undeformed state via a reversible phase transformation when subject to a thermal, mechanical, or magnetic load [2]. SMMs have been used widely in aerospace and biomedical industries, including applications such as actuators, low-shock mechanisms, medical staples, braces, and stents [3, 4]. These materials exhibit unique behavior due to their ability to transition between phases, and thus the mechanisms that enable this transition must be captured in a data information management system and incorporated into SMM material models. At NASA Glenn Research Center, the Shape Memory Materials Database (SMMD) Tool has been developed to capture the necessary information that governs SMM material behavior and provide users the ability to select and visualize various SMMs for a specific application [5]. The database contains point-wise data for published SMM materials, along with the pedigree metadata for traceability necessary for a robust information management system. The database is also capable of storing in-house test data performed at NASA GRC by interacting with the developed Shape Memory Alloy (SMA) Analytics tool to extract the necessary point-wise values and populate the database. Although the SMMD Tool offers its users a single, authoritative source for SMM material data that is critical for model development and material design, the full material pedigree of the in-house test data for SMMs is not currently captured and is out of the scope for the SMMD tool. In this work, the schema for capturing SMM test data within the larger NASA GRC ICME Schema [6, 7, 8, 9] will be developed and implemented for thermomechanical tests conducted at NASA GRC. The developed schema will not only store the relevant data needed for the SMMD tool, but also the material pedigree (i.e., production of the bulk material, bulk material analysis, sample cut-out diagrams, sample fabrication procedure, etc.), test pedigree (i.e., test equipment used, measurement systems used, raw test data), and analysis pedigree (i.e., how the data in the SMMD tool is calculated). Furthermore, a Python-based framework will be developed to seamlessly interact between the SMA Analytics and SMMD tools, which will write the full dataset and associated metadata to the GRC Information Management System before passing the required point-wise data to the SMMD tool. Data informatics is a key element of the NASA Vision 2040, which requires not only that data is stored and maintained throughout the material lifecycle, but that the data is also accessible and reusable such that material development efforts can be minimized. Therefore, for an ICME design approach to be realized, a centralized information management system that drives the ICME process must be able to communicate with other databases. The work that will be presented in this presentation will therefore not only demonstrate the ability of NASA GRC’s information management system to capture SMM data, but also its ability to interact with pre-existing tools specialized for such materials.

Data management↗

The ParaScope parallel programming environment

The ParaScope parallel programming environment, developed to support scientific programming of shared-memory multiprocessors, includes a collection of tools that use global program analysis to help users develop and debug parallel programs. This paper focuses on ParaScope's compilation system, its parallel program editor, and its parallel debugging system. The compilation system extends the traditional single-procedure compiler by providing a mechanism for managing the compilation of complete programs. Thus, ParaScope can support both traditional single-procedure optimization and optimization across procedure boundaries. The ParaScope editor brings both compiler analysis and user expertise to bear on program parallelization. It assists the knowledgeable user by displaying and managing analysis and by providing a variety of interactive program transformations that are effective in exposing parallelism. The debugging system detects and reports timing-dependent errors, called data races, in execution of parallel programs. The system combines static analysis, program instrumentation, and run-time reporting to provide a mechanical system for isolating errors in parallel program executions. Finally, we describe a new project to extend ParaScope to support programming in FORTRAN D, a machine-independent parallel programming language intended for use with both distributed-memory and shared-memory parallel computers.

Cooper, Keith D.↗

Comparison of Machine Learning-Based Predictive Models of the Nutrient Loads Delivered from the Mississippi/Atchafalaya River Basin to the Gulf of Mexico

Predicting nutrient loads is essential to understanding and managing one of the environmental issues faced by the northern Gulf of Mexico hypoxic zone, which poses a severe threat to the Gulf’s healthy ecosystem and economy. The development of hypoxia in the Gulf of Mexico is strongly associated with the eutrophication process initiated by excessive nutrient loads. Due to the complexities in the excessive nutrient loads to the Gulf of Mexico, it is challenging to understand and predict the underlying temporal variation of nutrient loads. The study was aimed at identifying an optimal predictive machine learning model to capture and predict nonlinear behavior of the nutrient loads delivered from the Mississippi/Atchafalaya River Basin (MARB) to the Gulf of Mexico. For this purpose, monthly nutrient loads (N and P) in tons were collected from US Geological Survey (USGS) monitoring station 07373420 from 1980 to 2020. Machine learning models—including autoregressive integrated moving average (ARIMA), gaussian process regression (GPR), single-layer multilayer perceptron (MLP), and a long short-term memory (LSTM) with the single hidden layer—were developed to predict the monthly nutrient loads, and model performances were evaluated by standard assessment metrics—Root Mean Square Error (RMSE) and Correlation Coefficient (R). The residuals of predictive models were examined by the Durbin–Watson statistic. The results showed that MLP and LSTM persistently achieved better accuracy in predicting monthly TN and TP loads compared to GPR and ARIMA. In addition, GPR models achieved slightly better test RMSE score than ARIMA models while their correlation coefficients are much lower than ARIMA models. Moreover, MLP performed slightly better than LSTM in predicting monthly TP loads while LSTM slightly outperformed for TN loads. Furthermore, it was found that the optimizer and number of inputs didn’t show effects on the LSTM performance while they exhibited impacts on MLP outcomes. This study explores the capability of machine learning models to accurately predict nonlinearly fluctuating nutrient loads delivered to the Gulf of Mexico. Further efforts focus on improving the accuracy of forecasting using hybrid models which combine several machine learning models with superior predictive performance for nutrient fluxes throughout the MARB.

54 ENVIRONMENTAL SCIENCES↗

Direct numerical simulations for hybrid rocket boundary layers: Performance modeling and scaling

This paper presents a comprehensive performance and scaling analysis of direct numerical simulations for reacting boundary layers, focusing on slab burner configurations. Using a PETSc-based finite volume CFD framework, the study evaluates the scalability and computational cost of flow, chemistry, and radiation evaluations across 2D and 3D simulations. Polymethyl methacrylate (PMMA) is the fuel with pure O 2 as the oxidizer, modeled using a detailed chemical kinetics mechanism with 113 species and 660 reactions. A ray-tracing-based radiation solver, designed for distributed memory applications, is implemented to model radiation heat transfer. Parallel scalability is analyzed for the coupled flow, chemistry, and radiation heat transfer processes. Weak and strong scaling studies are conducted on up to 15,000 computational ranks, revealing robust performance when flow cells exceed 200 per rank. Chemistry evaluations dominate the computational cost in large 3D simulations, accounting for approximately 40% of the total runtime, while flow processes contribute around 35%, and radiation solver contributions remain below 10% due to reduced evaluation frequencies. GPU accelerated chemistry evaluation, implemented with Zero-RK, demonstrates significant promise, achieving up to a 4x speedup for workloads exceeding 30,000 cells per GPU. However, diminishing returns are observed for smaller workloads due to CPU-GPU communication overhead. This study identifies key challenges, including memory bottlenecks and the effects of domain partitioning on flow scalability, while highlighting the potential of GPU-accelerated chemistry to reduce computational costs. In conclusion, these findings provide realizable run configurations for 2D, 3D, and GPU-accelerated cases, offering insights for optimizing reactive flow solvers.

CFD Scalability↗

Score-based deterministic density sampling

We propose a deterministic sampling framework using Score-Based Transport Modeling for sampling an unnormalized target density π given only its score ∇ log π. Our method approximates the Wasserstein gradient flow on KL($f_t$∥π) by learning the time-varying score ∇ log $f_t$ on the fly using score matching. While having the same marginal distribution as Langevin dynamics, our method produces smooth deterministic trajectories, resulting in monotone noise-free convergence. We prove that our method dissipates relative entropy at the same rate as the exact gradient flow, provided sufficient training. Numerical experiments validate our theoretical findings: our method converges at the optimal rate, has smooth trajectories, and is often more sample efficient than its stochastic counterpart. Experiments on high-dimensional image data show that our method produces high-quality generations in as few as 15 steps and exhibits natural exploratory behavior. The memory and runtime scale linearly in the sample size.

97 MATHEMATICS AND COMPUTING↗

Partitioning problems in parallel, pipelined and distributed computing

The problem of optimally assigning the modules of a parallel program over the processors of a multiple computer system is addressed. A Sum-Bottleneck path algorithm is developed that permits the efficient solution of many variants of this problem under some constraints on the structure of the partitions. In particular, the following problems are solved optimally for a single-host, multiple satellite system: partitioning multiple chain structured parallel programs, multiple arbitrarily structured serial programs and single tree structured parallel programs. In addition, the problems of partitioning chain structured parallel programs across chain connected systems and across shared memory (or shared bus) systems are also solved under certain constraints. All solutions for parallel programs are equally applicable to pipelined programs. These results extend prior research in this area by explicitly taking concurrency into account and permit the efficient utilization of multiple computer architectures for a wide range of problems of practical interest.

Bokhari, S.↗

Conditional Entropy-Constrained Residual VQ with Application to Image Coding

This paper introduces an extension of entropy-constrained residual vector quantization (VQ) where intervector dependencies are exploited. The method, which we call conditional entropy-constrained residual VQ, employs a high-order entropy conditioning strategy that captures local information in the neighboring vectors. When applied to coding images, the proposed method is shown to achieve better rate-distortion performance than that of entropy-constrained residual vector quantization with less computational complexity and lower memory requirements. Moreover, it can be designed to support progressive transmission in a natural way. It is also shown to outperform some of the best predictive and finite-state VQ techniques reported in the literature. This is due partly to the joint optimization between the residual vector quantizer and a high-order conditional entropy coder as well as the efficiency of the multistage residual VQ structure and the dynamic nature of the prediction.

Kossentini, Faouzi↗

High-Order Coronagraphic Wavefront Control With Algorithmic Differentiation: First Experimental Demonstration

Future space-based coronagraphs will rely critically on focal-plane wavefront sensing and control with deformable mirrors to reach deep contrast by mitigating optical aberrations in the primary beam path. Until now, most focal-plane wavefront control algorithms have been formulated in terms of Jacobian matrices, which encode the predicted effect of each deformable mirror actuator on the focal-plane electric field. A disadvantage of these methods is that Jacobian matrices can be cumbersome to compute and manipulate, particularly when the number of deformable mirror actuators is large. Recently, we proposed a new class of focal-plane wavefront control algorithms that utilize gradient-based optimization with algorithmic differentiation to compute wavefront control solutions while avoiding the explicit computation and manipulation of Jacobian matrices entirely. In simulations using a coronagraph design for the proposed Large UV/Optical/Infrared Surveyor (LUVOIR), we showed that our approach reduces overall CPU time and memory consumption compared to a Jacobian-based algorithm. Here, we expand on these results by implementing the proposed algorithm on the High Contrast Imager for Complex Aperture Telescopes (HiCAT) testbed at the Space Telescope Science Institute (STScI) and present initial experimental results, demonstrating contrast suppression capabilities equivalent to Jacobian-based methods.

wavefront control↗

Towards a NEAMS-based high-fidelity model of the MARVEL reactor

This report outlines the progress of Idaho National Laboratory in developing a high-fidelity and high-resolution model of the Microreactor Applications Research Validation and Evaluation reactor. The model was developed under the Nuclear Energy Advanced Modeling and Simulation microreactor application driver at Idaho National Laboratory. The overarching objective of this activity is the development of a high-fidelity multiphysics MARVEL model using NEAMS tools, and to verify and validate NEAMS tools against MARVEL reference simulation and experimental data, respectively. This is a unique opportunity to conduct multiphysics analysis on a soon-to-be-deployed microreactor. This multiphysics model developed under the NEAMS-funded INL microreactor application driver leverages three single-physics models coupled via the MOOSE’s MultiApp and Transfer systems. The latter systems enable in-memory data transfer between MOOSE-based and MOOSE-wrapped applications. The first single-physics model, that functions as main application, leverages Griffin to model the neutron transport in the core through the discontinuous finite element (DFEM) discrete ordinates solver (SN). Several optimization flags that were developed by the Griffin developer team were beta-tested to enhance the solver’s performance. These include the combined use of using_average_xs and update_averaged_xs_on that enable to avoid expensive on-the-fly cross sections evaluations at each linear iterations in favor of evaluations of the macroscopic cross sections at each Picard iteration. The second single-physics model uses BISON to handle solid heat transfer and asymptotic hydrogen redistribution analysis in the fuel. While the model returns consistent results for the temperature and hydrogen distribution in the fuel, a mismatch was noticed in the calculated temperature in the reflector due to the value of the gap conductance used in our model. Ongoing investigations are being performed to assess the origin of this discrepancy. Finally, the System Analysis Module (SAM) was used to model the flow of the sodium-potassium eutectic in the primary loop. A first verification was also performed showing good agreement in terms of mass flow rate and inlet temperature. All mesh files were generated using the MOOSE Reactor module, removing the need for external meshing tools. Notably, this workscope represents one of the initial applications of the MOOSE Reactor module for modeling highly irregular geometries. The use of the reactor module significantly streamlined the mesh generation process. The full multiphysics mode, that combines all the single physics models, was leveraged to conduct initial steady-state multiphysics simulations to compute power, and temperature distribution in the reactor. Initial testing was performed for transient simulations as well. In this case, the new checkpoint restart capability for eigenvalue calculations was tested showing the capability for streamlined restart of transient calculations. Future work will focus on improving the fidelity of the model by performing comprehensive code-to-code comparisons. For instance, the full-core Griffin neutronics model will be benchmarked against MCNP reference results, that were provided by the MARVEL design team. Additionally, the SAM T/H model will be verified against reference RELAP-5 results for selected accident scenarios. Besides code-to-code verification exercises, the model fidelity will be improved by replacing the single-channel SAM model with a more complex SAM-Pronghorn coupled model, in which the sub-channel capability is deployed to obtain radial temperature resolution in the coolant. This model will be developed in synergy with the NEAMS thermal hydraulics team.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Optically intraconnected computer employing dynamically reconfigurable holographic optical element

An optically intraconnected computer and a reconfigurable holographic optical element employed therein. The basic computer comprises a memory for holding a sequence of instructions to be executed; logic for accessing the instructions in sequence; logic for determining for each the instruction the function to be performed and the effective address thereof; a plurality of individual elements on a common support substrate optimized to perform certain logical sequences employed in executing the instructions; and, element selection logic connected to the logic determining the function to be performed for each the instruction for determining the class of each function and for causing the instruction to be executed by those the elements which perform those associated the logical sequences affecting the instruction execution in an optimum manner. In the optically intraconnected version, the element selection logic is adapted for transmitting and switching signals to the elements optically.

Bergman, Larry A.↗

High-Rate Communications Outage Recorder Operations for Optimal Payload and Science Telemetry Management Onboard the International Space Station

All International Space Station (ISS) Ku-band telemetry transmits through the High-Rate Communications Outage Recorder (HCOR). The HCOR provides the recording and playback capability for all payload, science, and International Partner data streams transmitting through NASA's Ku-band antenna system. The HCOR is a solid-state memory recorder that provides recording capability to record all eight ISS high-rate data during ISS Loss-of-Signal periods. NASA payloads in the Destiny module are prime users of the HCOR; however, NASDA and ESA will also utilize the HCOR for data capture and playback of their high data rate links from the Kibo and Columbus modules. Marshall Space Flight Center's Payload Operations Integration Center manages the HCOR for nominal functions, including system configurations and playback operations. The purpose of this paper is to present the nominal operations plan for the HCOR and the plans for handling contingency operations affecting payload operations. In addition, the paper will address HCOR operation limitations and the expected effects on payload operations. The HCOR is manifested for ISS delivery on flight 9A with the HCOR backup manifested on flight 11A. The HCOR replaces the Medium-Rate Communications Outage Recorder (MCOR), which has supported payloads since flight 5A.1.

Shell, Michael T.↗

Data traffic reduction schemes for Cholesky factorization on asynchronous multiprocessor systems

Communication requirements of Cholesky factorization of dense and sparse symmetric, positive definite matrices are analyzed. The communication requirement is characterized by the data traffic generated on multiprocessor systems with local and shared memory. Lower bound proofs are given to show that when the load is uniformly distributed the data traffic associated with factoring an n x n dense matrix using n to the alpha power (alpha less than or equal 2) processors is omega(n to the 2 + alpha/2 power). For n x n sparse matrices representing a square root of n x square root of n regular grid graph the data traffic is shown to be omega(n to the 1 + alpha/2 power), alpha less than or equal 1. Partitioning schemes that are variations of block assignment scheme are described and it is shown that the data traffic generated by these schemes are asymptotically optimal. The schemes allow efficient use of up to O(n to the 2nd power) processors in the dense case and up to O(n) processors in the sparse case before the total data traffic reaches the maximum value of O(n to the 3rd power) and O(n to the 3/2 power), respectively. It is shown that the block based partitioning schemes allow a better utilization of the data accessed from shared memory and thus reduce the data traffic than those based on column-wise wrap around assignment schemes.

Naik, Vijay K.↗

Efficient generation of grids and traversal graphs in compositional spaces towards exploration and path planning

Abstract Diverse disciplines across science and engineering deal with problems related to compositions, which exist in non-Euclidean simplex spaces, rendering many standard tools inaccurate or inefficient. This work explores such spaces conceptually in the context of materials discovery, quantifies their computational feasibility, and implements several essential methods specific to simplex spaces through a new high-performance open-source library . Most significantly, we derive and implement an algorithm for constructing a novel n-dimensional simplex graph data structure, containing all discretized compositions and possible neighbor-to-neighbor transitions. Critically, no distance or neighborhood calculations are performed, instead leveraging pure combinatorics and order in procedurally generated simplex grids, keeping the algorithm $${\mathcal{O}}(N)$$ O ( N ) , with minimal memory, enabling rapid construction of graphs with billions of transitions in seconds. Additionally, we demonstrate how such graph representations can be combined to homogeneously express complex path-planning problems, while facilitating efficient deployment of existing high-performance gradient descent, graph traversal, and other optimization algorithms.

Krajewski, Adam M. (ORCID:0000000222660099)↗

Interference Lattice-based Loop Nest Tilings for Stencil Computations

A common method for improving performance of stencil operations on structured multi-dimensional discretization grids is loop tiling. Tile shapes and sizes are usually determined heuristically, based on the size of the primary data cache. We provide a lower bound on the numbers of cache misses that must be incurred by any tiling, and a close achievable bound using a particular tiling based on the grid interference lattice. The latter tiling is used to derive highly efficient loop orderings. The total number of cache misses of a code is the sum of (necessary) cold misses and misses caused by elements being dropped from the cache between successive loads (replacement misses). Maximizing temporal locality is equivalent to minimizing replacement misses. Temporal locality of loop nests implementing stencil operations is optimized by tilings that avoid data conflicts. We divide the loop nest iteration space into conflict-free tiles, derived from the cache miss equation. The tiling involves the definition of the grid interference lattice an equivalence class of grid points whose images in main memory map to the same location in the cache-and the construction of a special basis for the lattice. Conflicts only occur on the boundaries of the tiles, unless the tiles are too thin. We show that the surface area of the tiles is bounded for grids of any dimensionality, and for caches of any associativity, provided the eccentricity of the fundamental parallelepiped (the tile spanned by the basis) of the lattice is bounded. Eccentricity is determined by two factors, aspect ratio and skewness. The aspect ratio of the parallelepiped can be bounded by appropriate array padding. The skewness can be bounded by the choice of a proper basis. Combining these two strategies ensures that pathologically thin tiles are avoided. They do not, however, minimize replacement misses per se. The reason is that tile visitation order influences the number of data conflicts on the tile boundaries. If two adjacent tiles are visited successively, there will be no replacement misses on the shared boundary. The iteration space may be covered with pencils larger than the size of the cache while avoiding data conflicts if the pencils are traversed by a scanning-face method. Replacement misses are incurred only on the boundaries of the pencils, and the number of misses is minimized by maximizing the volume of the scanning face, not the volume of the tile. We present an algorithm for constructing the most efficient scanning face for a given grid and stencil operator. In two dimensions it is based on a continued fraction algorithm. In three dimensions it follows Voronoi's successive minima algorithm. We show experimental results of using the scanning face, and compare with canonical loop orderings.

VanderWijngaart, Rob F.↗

Optimal Padding for the Two-Dimensional Fast Fourier Transform

One-dimensional Fast Fourier Transform (FFT) operations work fastest on grids whose size is divisible by a power of two. Because of this, padding grids (that are not already sized to a power of two) so that their size is the next highest power of two can speed up operations. While this works well for one-dimensional grids, it does not work well for two-dimensional grids. For a two-dimensional grid, there are certain pad sizes that work better than others. Therefore, the need exists to generalize a strategy for determining optimal pad sizes. There are three steps in the FFT algorithm. The first is to perform a one-dimensional transform on each row in the grid. The second step is to transpose the resulting matrix. The third step is to perform a one-dimensional transform on each row in the resulting grid. Steps one and three both benefit from padding the row to the next highest power of two, but the second step needs a novel approach. An algorithm was developed that struck a balance between optimizing the grid pad size with prime factors that are small (which are optimal for one-dimensional operations), and with prime factors that are large (which are optimal for two-dimensional operations). This algorithm optimizes based on average run times, and is not fine-tuned for any specific application. It increases the amount of times that processor-requested data is found in the set-associative processor cache. Cache retrievals are 4-10 times faster than conventional memory retrievals. The tested implementation of the algorithm resulted in faster execution times on all platforms tested, but with varying sized grids. This is because various computer architectures process commands differently. The test grid was 512 512. Using a 540 540 grid on a Pentium V processor, the code ran 30 percent faster. On a PowerPC, a 256x256 grid worked best. A Core2Duo computer preferred either a 1040x1040 (15 percent faster) or a 1008x1008 (30 percent faster) grid. There are many industries that can benefit from this algorithm, including optics, image-processing, signal-processing, and engineering applications.

Dean, Bruce H.↗

Deep learning-assisted modeling for χ (2) nonlinear optics

Modeling second-order (χ(2)) nonlinear optical processes remains computationally expensive due to the need to resolve fast field oscillations and simulate wave propagation using methods such as the split-step Fourier method (SSFM). This can become a bottleneck in real-time applications, such as high-repetition-rate laser systems requiring rapid feedback and control. We present a long short-term memory-based surrogate model trained on SSFM simulations generated from a start-to-end model of the photocathode drive laser at SLAC National Accelerator Laboratory’s Linac Coherent Light Source II. The model achieves over 250× speedup while maintaining high fidelity, enabling future real-time optimization and laying the foundation for data-integrated modeling frameworks and digital twins of laser systems.

Accelerator Physics (physics.acc-ph)↗