Search NASA⌕ Search

SEARCH · Search NASA

Results for “computational cost”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

A physics-constrained neural ordinary differential equations approach for robust learning of stiff chemical kinetics

The high computational cost associated with solving for detailed chemistry poses a significant challenge for predictive computational fluid dynamics (CFD) simulations of turbulent reacting flows. While deep learning techniques have been explored to develop faster surrogate models, they often fail to integrate reliably with CFD solvers. This instability arises because traditional deep learning approaches optimize for training error without ensuring compatibility with ordinary differential equation (ODE) solvers, resulting in accumulation of errors over time. Recently, neuralODE (NODE) based approaches have been shown to be a promising technique to emulate and accelerate detailed chemistry computations. Here, in the present work, we extend this NODE framework for stiff chemical kinetics by incorporating mass conservation constraints directly into the loss function during training. This ensures that the total mass as well as the individual elemental species masses are conserved in an a-posteriori manner. Proof-of-concept studies are performed with the novel physics-constrained NODE (PC-NODE) approach for homogeneous autoignition of hydrogen-air mixture over a range of composition and thermodynamic conditions. It is demonstrated that the PC-NODE framework not only improves the physical consistency of the resulting data-driven model with respect to mass conservation criteria, but also improves training efficiency. PC-NODE is shown to achieve 2–100× speedup relative to the hydrogen-air detailed chemical mechanism depending on the type of the ODE solver (implicit or explicit) used during autoregressive inference tests. Lastly, a-posteriori studies are performed wherein the trained PC-NODE model is coupled with a CFD solver. It is shown that higher accuracy is achieved with PC-NODE relative to the purely data-driven NODE approach. Moreover, PC-NODE also exhibits robustness and generalizability to unseen initial conditions from within (interpolative capability) as well as outside (extrapolative capability) the training regime.

computational combustion↗

RNS Applications for Interacting Sub- and Supersonic Flows

A solution based grid adaptation method that combines elements of the multigrid method for solution acceleration and the domain decomposition philosophy for grid optimization is described. Unlike other solution based adaptive gridding schemes, wherein the overhead of recomputing the grid and re-evaluating the solution on the adapted grid leads to higher computational costs compared to a non-adapted calculation, the present methodology reduces the computational time required to obtain the solution. The computational effort involved in the present calculation is significantly lower than a non-adapted calculation that utilizes the multigrid method purely as a convergence acceleration tool. In addition to convergence acceleration, the multigrid framework provides a mechanism of information transfer from regions wherein grid refinement is specified to unrefined coarse grid regions. The basis for domain decomposition in the current procedure is the variation in grid refinement requirements for each coordinate direction in different portions of the flow field. The method is demonstrated herein on an efficient set of governing equations termed the reduced Navier Stokes equations, applied in conjunction with a set of physical boundary conditions. The governing equations are discretized through a pressure based flux splitting procedure that is uniformly applicable from incompressible to supersonic Mach numbers.

Rubin, Stanley G.↗

Effects of Spatial Resolution on Retropropulsion Aerodynamics in an Atmospheric Environment

Development of a powered descent capability for atmospheric environments is heavily reliant on computational simulation. The prohibitive computational cost of such simulations motivates an improvement in the understanding of the minimum computational fidelity re-quired to accurately characterize aerodynamic-propulsive interference for such applications. This work examines the applicability of detached eddy simulation methods for retropropulsion in atmospheric environments through utilization of a GPU-accelerated computational framework, yielding data that are largely unachievable with conventional high-performance computing resources. This effort was specifically designed to quantitatively assess the effects of spatial resolution on vehicle aerodynamics for nominal operation of a low lift-to-drag ratio, human-scale Mars lander concept. The test matrix and scaling approach span relevant nozzle expansion conditions as well as mid-supersonic to high-subsonic operating conditions. Solutions were generated using computational grids ranging from 143 million to 1.14 billion grid points (degrees of freedom). This paper will provide an overview of the computational campaign, approach, and discussion of preliminary results focused on a range of operating conditions for a conceptual low lift-to-drag, human-scale Mars lander.

Ashley M Korzun↗

Modeling Deformable Linear Objects for Autonomous Robotic Outfitting of Lunar Surface Systems

This paper presents structural models of deformable linear objects (DLOs). DLOs are a subclass of deformable objects that encompasses common outfitting elements such as cables and ropes. Models are validated through hardware experiments, and integration in a robotic autonomy architecture for space environments is discussed. A persistent human presence on the lunar surface is one of the next major milestones in space exploration. This requires the development of robust extraplanetary construction technologies including structures and materials modeling and robotic systems. Previous robotic construction technology development has primarily focused on structural assembly, with significantly less focus on robotically performed outfitting tasks to instantiate subsystems providing power, data, life support, etc. These tasks involve manipulation of highly flexible elements, which are difficult to model, such as cable harnesses, ropes, and hoses. Robotic manipulation of DLOs, especially cable harnesses, is an active area of research as cable harnesses are essential for providing power and data to space assets. DLO models that can be used for robot manipulator trajectory generation are necessary for autonomous operation of lunar infrastructure. There are many proposed methods for modeling DLOs, and they primarily fall into three types: 1) discrete model-based, 2) continuum model-based, and 3) Neural Network-based. These types each have pros and cons, and the tradeoff between model accuracy and computational speed informs which type should be used. An understanding of this trade-off is imperative for real-time control of autonomous systems. High computational requirements reduce the speed of the model, making real-time control difficult, while accuracy is critical to preventing collisions. Discrete models, such as a mass-spring multibody representation, require relatively few calculations, and accuracy is directly tied to the step size of the discretization. Continuum models, such as a B-spline representation or a Cosserat rod model (a mix of continuous and discrete), are more informed of the structural properties of the cable and are much more accurate than a rigid body mass-spring model, but at significant computational cost. A Neural Network approach can provide an online solution with very few computational steps, but properly generating training data can be difficult and validation for an in-space application is not trivial. This paper explores the trade-off between different modeling approaches and compares accuracy and computational speed/complexity of the three types mentioned above. Model accuracy is evaluated using a cable in a static configuration. True cable shape is obtained using a depth camera for RGB images and point-cloud segmentation. The purpose of this experiment is to evaluate the trade-offs of different approaches to the DLO modeling problem. Understanding the tradeoffs between different cable modeling techniques paves the way for developing robotic control and planning architectures necessary for real-time manipulation of DLOs for lunar infrastructure outfitting. Real-time control is required for robotic systems to be able to actively manipulate a cable in a harsh environment where model and sensor errors compound, and environmental conditions can cause significant disturbances. Cable routing must be performed in areas with high density of objects/obstacles: through truss structures, near solar panels or mirror arrays, next to bundles of electrical equipment. Understanding the best way to plan and manipulate a cable without disrupting the environment or damaging the cable is imperative to robotic outfitting operations on the lunar surface.

Amy M Quartaro↗

Simulation of the Rotorwash Induced by a Quadrotor Urban Air Taxi in Ground Effect

A rotorcraft hovering near the ground causes downwash and outwash. The impact of these induced velocities on urban air mobility operations has not been extensively quantified. This paper explores the opportunity of using CFD to perform dedicated studies on the aerodynamics of rotorcraft in ground effect (IGE). For this purpose, we compare load and flow predictions obtained with the high-fidelity CFD solver OVERFLOW and with a medium-fidelity vortex particle-mesh (VPM) method, for a single rotor IGE. We show that the computational cost of high-fidelity simulation makes it impractical to sweep through a large number of configurations or operating conditions. However, a small set of cases can be used to verify a medium-fidelity tool. The latter provides a better trade-off between “accuracy” and computational intensity, which enables numerous, longer simulations at a more affordable cost. As an example, the outwash flow is computed for two different designs of a quadrotor air taxi in hover and reveals the existence of increased velocities between the rotors. This work illustrates how CFD can help identify dangerous areas for passengers and ground personnel when they approach the vehicle.

HECC↗

Enhancing the accuracy of XPS calculations: Exploring hybrid basis set schemes for CVS-EOMIP-CCSD calculations

Reliable computational methodologies and basis sets for modeling x-ray spectra are essential for extracting and interpreting electronic and structural information from experimental x-ray spectra. In particular, the trade-off between numerical accuracy and computational cost due to the size of the basis set is a major challenge, since molecular orbitals undergo extreme relaxation in the core-hole state. To gain clarity on the changes in electronic structure induced by the formation of a core-hole, the use of sufficiently flexible basis for expanding the orbitals, particularly for the core region, has been shown to be essential. This work focuses on the refinement of core-hole ionized state calculations using the equation-of-motion coupled cluster family of methods through an extensive analysis on the effectiveness of “hybrid” and mixed basis sets. In this investigation, we utilize the CVS-EOMIP-CCSD method in combination and construct hybrid basis sets piecewise from readily available Dunning’s correlation consistent basis sets in order to calculate x-ray ionization energies (IEs) for a set of small gas phase molecules. Our results provide insights into the impact of basis sets on the CVS-EOMIP-CCSD calculations of K-edge IEs of first-row p-block elements. Furthermore, these insights enable us to understand more about the basis set dependence of the core IEs computed and allow us to establish a protocol for deriving reliable and cost-effective theoretical estimates for computing IEs of small molecules containing such elements.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Taming the virtual space for incremental full configuration interaction

Incremental full configuration interaction (iFCI) closely approximates the FCI limit with polynomial cost through a many-body expansion of the correlation energy, providing highly accurate total energies within a given basis set. To extend iFCI beyond previous basis set limitations, this work introduces a novel natural orbital (NO) screening approach, incremental NO full configuration interaction (iNO-FCI). By consideration of the importance of virtual orbital selection in the convergence of iFCI, iNO-FCI maximizes the consistency between orbitals selected for each correlated body. iNO-FCI employs a principle of cancellation of errors and ensures that the same set of virtual NOs is used for interdependent terms. Here, this strategy significantly reduces computational cost without compromising precision. Computational savings of up to 95% are demonstrated, allowing access to larger basis sets that were previously computationally prohibitive. iNO-FCI is herein introduced and benchmarked for several difficult test cases involving double-bond dissociation, biradical systems, conjugated π systems, and the spin gap of a Cu-based transition metal complex.

Correlation energy↗

Uncertainty Quantification and Certification Prediction of Low-Boom Supersonic Aircraft Configurations

The primary objective of this work was to develop and demonstrate a process for accurate and efficient uncertainty quantification and certification prediction of low-boom, supersonic, transport aircraft. High-fidelity computational fluid dynamics models of multiple low-boom configurations were investigated including the Lockheed Martin SEEB-ALR body of revolution, the NASA 69 Delta Wing, and the Lockheed Martin 1021-01 configuration. A nonintrusive polynomial chaos surrogate modeling approach was used for reduced computational cost of propagating mixed, inherent (aleatory) and model-form (epistemic) uncertainty from both the computation fluid dynamics model and the near-field to ground level propagation model. A methodology has also been introduced to quantify the plausibility of a design to pass a certification under uncertainty. Results of this study include the analysis of each of the three configurations of interest under inviscid and fully turbulent flow assumptions. A comparison of the uncertainty outputs and sensitivity analyses between the configurations is also given. The results of this study illustrate the flexibility and robustness of the developed framework as a tool for uncertainty quantification and certification prediction of low-boom, supersonic aircraft.

West, Thomas K., IV↗

Automated Hybrid Variance Reduction on Advanced Architectures in the Shift Monte Carlo Code

Monte Carlo transport methods are the most accurate schemes for solving problems with complex energy and spatial features, but they come with a high computational cost. Although hybrid methods have enabled the use of Monte Carlo transport for a large class of problems, they still require significant computing resources. Modern multicore CPUs with large numbers of compute cores and graphical processing units (GPUs) provide opportunities to optimize the memory and run-time costs of hybrid Monte Carlo methods. This paper documents the development and analysis of three Monte Carlo transport algorithms that support hybrid transport using the consistent adjoint-driven importance sampling (CADIS) and forward-weighted CADIS methods in the Shift Monte Carlo code: history-based transport using static and dynamic threading on multicore CPUs and event-based transport enabling weight window tracking on GPUs. The results are shown for two challenging hybrid problems on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility. The results show that all three methods yield good performance and enable solutions of difficult fixed-source transport problems in less than 2 min on 20 nodes of Frontier. Dynamic threading was observed to give up to 20% better scaling behavior than static threading. Moreover, the AMD Instinct 250X GPU was found to give 9 to 11 times greater throughput per graphics compute die than the best CPU performance. In conclusion, additional opportunities for optimization of hybrid transport on GPUs are discussed.

Denovo↗

A High-Efficiency Delayed Update Algorithm for Evaluating Slater Determinants in Quantum Monte Carlo

For quantum Monte Carlo simulations of molecular systems or supercells with thousands of electrons, matrix operations related to Slater determinants lead the computational cost. McDaniel et al. [J. Chem. Phys. 2017, 147, 174107] proposed a delayed update algorithm to increase computational efficiency by using matrix–matrix multiplication when updating the inverse matrices of Slater determinants. However, preparing intermediate matrices for applying the Sherman–Morrison–Woodbury formula remained a bottleneck. Here, in this work, we introduce an improved algorithm for CPUs and GPUs that (1) reduces this bottleneck by iteratively updating the intermediate matrices and (2) is efficient at any acceptance ratio, with no cost for rejected moves on CPUs and minimal cost on GPUs. We show the full scheme of integrating the delayed update algorithm into a single-electron move. The high efficiency of our algorithm is demonstrated on CPUs and GPUs for a 512 atom/6144 valence electron calculation, with 12× and 2× overall speed-up compared to traditional rank-1 update schemes in diffusion quantum Monte Carlo, respectively.

Luo, Ye [Argonne National Laboratory (ANL), Argonn↗

Accelerate microstructure evolution simulation using graph neural networks with adaptive spatiotemporal resolution

Abstract Surrogate models driven by sizeable datasets and scientific machine-learning methods have emerged as an attractive microstructure simulation tool with the potential to deliver predictive microstructure evolution dynamics with huge savings in computational costs. Taking 2D and 3D grain growth simulations as an example, we present a completely overhauled computational framework based on graph neural networks with not only excellent agreement to both the ground truth phase-field methods and theoretical predictions, but enhanced accuracy and efficiency compared to previous works based on convolutional neural networks. These improvements can be attributed to the graph representation, both improved predictive power and a more flexible data structure amenable to adaptive mesh refinement. As the simulated microstructures coarsen, our method can adaptively adopt remeshed grids and larger timesteps to achieve further speedup. The data-to-model pipeline with training procedures together with the source codes are provided.

36 MATERIALS SCIENCE↗

Testing convolutional neural network based deep learning systems: a statistical metamorphic approach

Machine learning technology spans many areas and today plays a significant role in addressing a wide range of problems in critical domains,i.e., healthcare, autonomous driving, finance, manufacturing, cybersecurity,etc. Metamorphic testing (MT) is considered a simple but very powerful approach in testing such computationally complex systems for which either an oracle is not available or is available but difficult to apply. Conventional metamorphic testing techniques have certain limitations in verifying deep learning-based models (i.e., convolutional neural networks (CNNs)) that have a stochastic nature (because of randomly initializing the network weights) in their training. In this article, we attempt to address this problem by using a statistical metamorphic testing (SMT) technique that does not require software testers to worry about fixing the random seeds (to get deterministic results) to verify the metamorphic relations (MRs). We propose seven MRs combined with different statistical methods to statistically verify whether the program under test adheres to the relation(s) specified in the MR(s). We further use mutation testing techniques to show the usefulness of the proposed approach in the healthcare space and test two CNN-based deep learning models (used for pneumonia detection among patients). The empirical results show that our proposed approach uncovers 85.71% of the implementation faults in the classifiers under test (CUT). Furthermore, we also propose an MRs minimization algorithm for the CUT, thus saving computational costs and organizational testing resources.

Computer Science↗

Investigation of Main Bearing Fatigue Estimate Sensitivity to Synthetic Turbulence Models Using a Novel Drivetrain Model Implemented in OpenFAST

ABSTRACT A coupled medium‐fidelity drivetrain model is developed and implemented in OpenFAST for a 10‐MW land‐based reference turbine. The implementation is verified against a fully coupled multibody wind turbine model, including a detailed drivetrain. The new model can simultaneously and accurately estimate main bearing loads and represent elastic bending of the drivetrain. It has low computational cost and is useful for early design phases, sensitivity analyses and complex systems like wind farms (where computational expense must be expended elsewhere). Here, the model is implemented for a monopile offshore wind turbine and used to investigate the sensitivity of main bearing basic rating life to different synthetic turbulence models. Large‐eddy simulations (LES) targeting stable, neutral, and unstable atmospheric conditions at below‐, near‐ and above‐rated wind speeds are used as a reference. The turbulence models recommended by the International Electrotechnical Commission, the Mann spectral tensor model, and the Kaimal spectral model with exponential coherence are fitted to the LES data. Additionally, a constrained turbulence generator, PyConTurb (short for Python Constrained Turbulence ), based on LES data, is applied in the aero‐hydro‐servo‐elastic simulations. Taking PyConTurb as the baseline, the Kaimal model significantly underestimates fatigue of the downwind main bearing, with between 10% and 40% less damage. The Mann model also underestimates the downwind main bearing fatigue by up to 30%. The upwind main bearing damage is driven by mean loads, and differences between models are less significant, although the trends are similar. Reasons for these discrepancies are investigated and attributed to differences in spatial and temporal variations among the turbulence models.

17 WIND ENERGY↗

Machine learning of 27Al NMR electric field gradient tensors for crystalline structures from DFT

NMR crystallography has emerged as a promising technique for the determination and refinement of atomic coordinates in crystal structures. The crystal structure of compounds containing quadrupolar nuclei, such as 27Al, can be improved by directly comparing solid-state NMR measurements to DFT computations of the electric field gradient (EFG) tensor. The non-negligible computational cost of these first-principles calculations limits the applicability of this method to all but the most well-defined structures. We developed a fast, low-cost machine learning model to predict EFG parameters based on local structural motifs and elemental parameters. We computed 8081 EFG tensors from 1681 27Al crystalline solids using DFT and benchmarked them against 105 experimentally measured 27Al sites. Surprisingly, simple local geometric features dominate the predictive performance of the resulting random-forest model, yielding an R2 value of 0.98 and an RMSE of 0.61 MHz for CQ, the quadrupolar coupling constant. This model accuracy should enable pre-refining future structural assignments before finally validating with first-principles calculations. Such a catalogue of 27Al NMR tensors can serve as a tool for researchers assigning complex NMR spectra influenced by the nuclear electric quadrupole interaction.

Sun, He↗

Boosting efficiency and reducing graph reliance: Basis adaptation integration in Bayesian multi-fidelity networks

The computational cost of high-fidelity numerical models makes outer-loop analysis, which requires repeated interrogation of the model such as uncertainty quantification, computationally demanding. Multi-fidelity methods, which construct a surrogate model using data from an ensemble of models of varying cost and accuracy, can substantially reduce the cost of outer-loop analysis. However, these methods can be difficult to apply when the model ensemble does not admit a clear hierarchy a priori and the correlations between models are low. Consequently, in this paper, we present a multi-fidelity method that leverages dimension reduction to enhance the correlation between models, thereby reducing the amount of data needed to train a surrogate from an unordered ensemble of models. Our method utilizes basis adaptation to build low-dimensional polynomial chaos expansions of each model and employs Multi-fidelity Networks to encode the relationships among models. We show that the resulting method exhibit two notable advantages over its counterpart: (1) enhanced accuracy (both reduced bias and variance); and (2) reduced dependency on the graph structure encoding relationships among models. We demonstrate the approach on an analytical test problem and a challenging finite element model for a spent nuclear fuel. Our method produces a surrogate model that is significantly more accurate than either a single-fidelity surrogate or a multi-fidelity surrogate constructed without basis adaptation.

42 ENGINEERING↗

Estimating the Impacts of Increasing Temperatures and the Efficacy of Climate Adaptation Strategies in Urban Microclimates with Deep Learning

As urbanization and climate change progress, understanding and addressing urban heat becomes a priority for climate adaptation efforts. High temperatures concentrated in the urban core can drive increased risk of heat-related death and illness as well as increased energy demand for cooling. However, modeling the urban microclimate is an ongoing field of research typically burdened by an imprecise description of the built environment, incomplete observational records, significant computational cost, and a lack of high-resolution estimates of the impacts of increasing temperatures. Here, we present computationally efficient machine learning methods that can improve the accuracy of urban temperature estimates when compared to historical reanalysis data. These models are applied to a neighborhood in Los Angeles, and we compare the energy benefits of heat mitigation strategies to the impacts of climate change. We find that cooling demand is likely to increase substantially through midcentury, but engineered high-albedo surfaces could lessen this increase by more than 50 %. The corresponding increase in winter gas heating offsets the summer cooling benefit in the current climate, but total annual energy use from combined heating and cooling with electric heat pumps benefits from the engineered heat mitigation strategies under both current and future climates.

54 ENVIRONMENTAL SCIENCES↗

The Singlet–Triplet Gap of Cyclobutadiene: The CIPSI-Driven CC( P ; Q ) Study

An accurate determination of singlet−triplet gaps in biradicals, including cyclobutadiene in the automerization barrier region where one has to balance the substantial nondynamical many-electron correlation effects characterizing the singlet ground state with the predominantly dynamical correlations of the lowest-energy triplet, remains a challenge for many quantum chemistry methods. High-level coupled-cluster (CC) approaches, such as the CC method with a full treatment of singly, doubly, and triply excited clusters (CCSDT), are often capable of providing reliable results, but routine application of such methods is hindered by their high computational costs. We have recently proposed a practical alternative to converging the CCSDT energetics at small fractions of the computational effort, even when electron correlations become stronger and connected triply excited clusters are larger and nonperturbative, by merging the CC(P;Q) moment expansions with the selected configuration interaction methodology abbreviated as CIPSI. We demonstrate that one can accurately approximate the highly accurate CCSDT potential surfaces characterizing the lowest singlet and triplet states of cyclobutadiene along the automerization coordinate and the gap between them using tiny fractions of triply excited cluster amplitudes identified with the help of relatively inexpensive CIPSI Hamiltonian diagonalizations.

Basis sets↗

FPGA Implementation of Stereo Disparity with High Throughput for Mobility Applications

High speed stereo vision can allow unmanned robotic systems to navigate safely in unstructured terrain, but the computational cost can exceed the capacity of typical embedded CPUs. In this paper, we describe an end-to-end stereo computation co-processing system optimized for fast throughput that has been implemented on a single Virtex 4 LX160 FPGA. This system is capable of operating on images from a 1024 x 768 3CCD (true RGB) camera pair at 15 Hz. Data enters the FPGA directly from the cameras via Camera Link and is rectified, pre-filtered and converted into a disparity image all within the FPGA, incurring no CPU load. Once complete, a rectified image and the final disparity image are read out over the PCI bus, for a bandwidth cost of 68 MB/sec. Within the FPGA there are 4 distinct algorithms: Camera Link capture, Bilinear rectification, Bilateral subtraction pre-filtering and the Sum of Absolute Difference (SAD) disparity. Each module will be described in brief along with the data flow and control logic for the system. The system has been successfully fielded upon the Carnegie Mellon University's National Robotics Engineering Center (NREC) Crusher system during extensive field trials in 2007 and 2008 and is being implemented for other surface mobility systems at JPL.

Random access memory↗