Search NASA⌕ Search

SEARCH · Search NASA

Results for “computing frameworks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Continuous integration data-driven platform of industrial-scale subsurface storage for real-time analytics

This project helped address the growing need for efficient and scalable models to support geological carbon and energy storage, which are crucial for achieving net-zero emissions. Traditionally accurate high-fidelity numerical models have been used to simulate relevant storage processes under a handful of processes, however such models are computationally demanding, making uncertainty quantification impractical. Consequently, we first developed a machine learning framework, based on Graph Neural Operators (GNOs), to improving the accuracy of model predictions for a fixed computational budget. We then developed an Ensemble of Improved Neural Operators (ENO), which uses bagging and Monte Carlo dropout techniques, to further improve prediction accuracy. Lastly, we developed the way to explain progressive transfer learning methods to reduce the amount of training data and computational cost of training (i.e., reduce trainable parameters) when using our models for multiple storage sites. Our numerical investigation, which used real-world case studies, demonstrated that our framework can significantly improve the safety and efficiency of geological storage operations, with potential applications in other domains such as geothermal reservoirs and climate modeling.

54 ENVIRONMENTAL SCIENCES↗

Computer-Simulation Surrogates for Optimization: Application to Trapezoidal Ducts and Axisymmetric Bodies

Engineering design and optimization efforts using computational systems rapidly become resource intensive. The goal of the surrogate-based approach is to perform a complete optimization with limited resources. In this paper we present a Bayesian-validated approach that informs the designer as to how well the surrogate performs; in particular, our surrogate framework provides precise (albeit probabilistic) bounds on the errors incurred in the surrogate-for-simulation substitution. The theory and algorithms of our computer{simulation surrogate framework are first described. The utility of the framework is then demonstrated through two illustrative examples: maximization of the flowrate of fully developed ow in trapezoidal ducts; and design of an axisymmetric body that achieves a target Stokes drag.

Otto, John C.↗

A Benchmark Suite for Evaluating Scientific AI Workloads on GPUs

AI applications have been steadily increasing in the allocation portfolio among leadership computing facilities. These applications depend on deep learning frameworks with hardware acceleration and underlying software systems. With the rapid development of applications, software stacks, and hardware devices, it is essential to evaluate the performance of core operations in AI workloads for direction of optimizations and procurement of next-generation high-performance computing (HPC) infrastructures. Currently, most benchmarks lack scientific AI workloads. So, we present DeepKernelBench and the experimental results of evaluating the benchmark suite for early observations and performance comparisons on datacenter GPUs using representative workloads for scientific AI, including Attentions, General matrix multiplications, Geometrics and Fourier neural operations.

Jin, Zheming [Advanced Micro Devices (AMD)]↗

Algorithms and software for nonlinear structural dynamics

The objective of this research is to develop efficient methods for explicit time integration in nonlinear structural dynamics for computers which utilize both concurrency and vectorization. As a framework for these studies, the program WHAMS, which is described in Explicit Algorithms for the Nonlinear Dynamics of Shells (T. Belytschko, J. I. Lin, and C.-S. Tsay, Computer Methods in Applied Mechanics and Engineering, Vol. 42, 1984, pp 225 to 251), is used. There are two factors which make the development of efficient concurrent explicit time integration programs a challenge in a structural dynamics program: (1) the need for a variety of element types, which complicates the scheduling-allocation problem; and (2) the need for different time steps in different parts of the mesh, which is here called mixed delta t integration, so that a few stiff elements do not reduce the time steps throughout the mesh.

Belytschko, Ted↗

Exploring Flood Predictability in Taiwan through Coupled Atmospheric–Hydrological and High-Performance Hydrodynamic Models

Effective flood simulation capabilities can tremendously support early warning and disaster prevention. To examine the applicability of a fully physics-based and high-performance flood simulation and forecasting modeling framework for a flood-prone region in Taiwan, we conduct a numerical experiment that couples the Weather Research and Forecasting (WRF) Model, WRF-Hydrological modeling system (WRF-Hydro), and the Two-Dimensional Runoff Inundation Toolkit for Operational Needs (TRITON) to perform integrated rainfall, streamflow, and flood simulations. Furthermore, we first use the coupled WRF and WRF-Hydro (WWH) to predict rainfall and streamflow and then drive TRITON with the predicted streamflow hydrographs to simulate flood depth and inundation area. With the refined spatial resolution and parameterization, this framework can better predict rainfall with reasonable spatial patterns. Although WWH could overestimate the amount of rainfall in some areas, the uncertain rainfall–streamflow predictions produce reasonable flood maps able to pinpoint regions at risk of flooding. In terms of model efficiency, the graphics processing unit–based computation can yield a speed-up factor as high as ∼13 compared to the central processing unit–based computation, promoting the efficacy of the coupled modeling framework in practical real-time flood forecasting.

Coupled models↗

Ambient and Initial Temperature Effects on Energy Consumption Rate Modeled in FASTSim

Ambient and initial temperatures significantly impact the energy consumption rate (ECR) of battery electric vehicles (BEVs) due to auxiliary loads and the temperature dependence of battery efficiency. This study introduces a streamlined, physics-based thermal modeling approach within the FASTSim tool that bridges the gap between oversimplified constant-load models and computationally expensive high-fidelity simulations. By employing a lumped thermal mass framework, the model captures fundamental energy balances and critical non-linear energy penalties while maintaining the computational efficiency required for expansive sensitivity studies. The simulations evaluated a compact BEV hatchback with a resistive heater over city (UDDS) and highway (HWFET) test cycles. Compared to a 22 degrees Celsius initial and ambient temperature baseline, a -7 degrees Celsius initial/ambient temperature resulted in a 221% increase in the ECR for the city cycle and a 100% increase for the highway cycle. Conversely, a 45 degrees Celsius initial / 40 degrees Celsius ambient temperature resulted in a 40% increase for UDDS and an 18% increase for HWFET. These results demonstrate that while cold conditions impose the most severe energy penalties due to resistive heating, the impact is consistently more pronounced in city driving where auxiliary loads represent a larger proportion of total energy. This lightweight yet robust framework enables researchers to rapidly quantify BEV thermal sensitivity across diverse climates without the need for high-overhead simulation environments.

33 ADVANCED PROPULSION SYSTEMS↗

PLUM: Parallel Load Balancing for Unstructured Adaptive Meshes

Dynamic mesh adaption on unstructured grids is a powerful tool for computing large-scale problems that require grid modifications to efficiently resolve solution features. Unfortunately, an efficient parallel implementation is difficult to achieve, primarily due to the load imbalance created by the dynamically-changing nonuniform grid. To address this problem, we have developed PLUM, an automatic portable framework for performing adaptive large-scale numerical computations in a message-passing environment. First, we present an efficient parallel implementation of a tetrahedral mesh adaption scheme. Extremely promising parallel performance is achieved for various refinement and coarsening strategies on a realistic-sized domain. Next we describe PLUM, a novel method for dynamically balancing the processor workloads in adaptive grid computations. This research includes interfacing the parallel mesh adaption procedure based on actual flow solutions to a data remapping module, and incorporating an efficient parallel mesh repartitioner. A significant runtime improvement is achieved by observing that data movement for a refinement step should be performed after the edge-marking phase but before the actual subdivision. We also present optimal and heuristic remapping cost metrics that can accurately predict the total overhead for data redistribution. Several experiments are performed to verify the effectiveness of PLUM on sequences of dynamically adapted unstructured grids. Portability is demonstrated by presenting results on the two vastly different architectures of the SP2 and the Origin2OOO. Additionally, we evaluate the performance of five state-of-the-art partitioning algorithms that can be used within PLUM. It is shown that for certain classes of unsteady adaption, globally repartitioning the computational mesh produces higher quality results than diffusive repartitioning schemes. We also demonstrate that a coarse starting mesh produces high quality load balancing, at a fraction of the cost required a fine initial mesh. Results indicate that our parallel load balancing strategy will remain viable on large numbers of processors.

Oliker, Leonid↗

Analyses of requirements for computer control and data processing experiment subsystems. Volume 2: ATM experiment S-056 image data processing system software development

The IDAPS (Image Data Processing System) is a user-oriented, computer-based, language and control system, which provides a framework or standard for implementing image data processing applications, simplifies set-up of image processing runs so that the system may be used without a working knowledge of computer programming or operation, streamlines operation of the image processing facility, and allows multiple applications to be run in sequence without operator interaction. The control system loads the operators, interprets the input, constructs the necessary parameters for each application, and cells the application. The overlay feature of the IBSYS loader (IBLDR) provides the means of running multiple operators which would otherwise overflow core storage.

Source record↗

Multi-physics melt pool modeling and process optimization for laser direct energy deposition of Nb-based refractory C103: Defect formation, geometric precision, and process mapping

Recent developments in additive manufacturing (AM) technology have reignited interest in the fabrication of the Nb-based refractory C103 alloy offering solutions to the challenges posed by traditional manufacturing methods. However, the limited numerical and experimental studies on laser direct energy deposition (DED) of C103 have hindered the understanding of the relationships between process parameters and build quality. This has made it challenging to consistently produce parts with the desired quality and microstructure suitable for critical applications. In this study, we focus on optimizing the laser DED process for C103 by employing a hybrid approach that combines experimental techniques and computational fluid dynamics (CFD). This approach facilitates the development of process maps for defect detection and geometric precision. To achieve this, multi-layer C103 samples were fabricated using laser DED under various process parameters, enabling the creation of a process map for defect detection. Additionally, a multi-physics, multiphase simulation framework was developed within a high-performance computing (HPC) environment to establish process maps for geometric precision. Using these process maps, printability windows were identified for achieving both the desired geometric accuracy and defect-free prints. It was observed that prints with a power-to-velocity (P/V) ratio close to unity resulted in defect-free outcomes. This study provides a foundation for reducing design lead time and rejected parts, ultimately optimizing the laser DED process for C103.

Defect formation and geometric precision↗

Computational Inference of Vibratory System with Incomplete Modal Information Using Parallel, Interactive and Adaptive Markov Chains

Inverse analysis of vibratory system is an important subject in fault identification, model updating, and robust design and control. It is challenging subject because 1) the problem is oftentimes underdetermined while the measurements are limited and/or incomplete; 2) many combinations of parameters may yield results that are similar with respect to actual response measurements; and 3) uncertainties inevitably exist. The aim of this research is to leverage upon computational intelligence through statistical inference to facilitate an enhanced, probabilistic framework using incomplete modal response measurement. This new framework is built upon efficient inverse identification through optimization, whereas Bayesian inference is employed to account for the effect of uncertainties. To overcome the computational cost barrier, we adopt Markov chain Monte Carlo (MCMC) to characterize the target function/distribution. Instead of using single Markov chain in conventional Bayesian approach, we develop a new sampling theory with multiple parallel, interactive and adaptive Markov chains and incorporate it into Bayesian inference. This can harness the collective power of these Markov chains to realize the concurrent search of multiple local optima. The number of required Markov chains and their respective initial model parameters are automatically determined via Monte Carlo simulation-based sample pre-screening followed by K-means clustering analysis. These enhancements can effectively address the aforementioned challenges in finite element inverse analysis. The validity of this framework is systematically demonstrated through case studies.

K Zhou↗

Envisioning an Optimal Network of Space-Based Lasers for Orbital Debris Remediation

The rapid increase in resident space objects, including satellites and orbital debris, poses a significant threat to the safety and sustainability of space missions. This paper explores orbital debris remediation using a network of collaborative space-based lasers, leveraging laser ablation for momentum transfer on debris. A novel delta-v vector analysis framework quantifies the e↵ects of multiple simultaneous laser-to-debris (L2D) engagements by using vector composition of the imparted delta-v vectors. The paper introduces the Concurrent LocationScheduling Problem (CLSP), which optimizes the placement of laser platforms and the scheduling of L2D engagements to maximize debris remediation capacity. Due to the computational complexity of the CLSP, it is decomposed into two sequential subproblems: (1) optimal laser platform locations are determined using the Maximal Covering Location Problem, and (2) a novel integer linear programming-based approach schedules L2D engagements within the network configuration to maximize remediation capacity. Computational experiments are conducted to evaluate the proposed framework’s e↵ectiveness under various mission scenarios, demonstrating key network functions such as collaborative nudging, deorbiting, and just-in-time collision avoidance. A sensitivity analysis further examines how varying the number and distribution of laser platforms a↵ects debris remediation capacity, providing insights into optimizing the performance of space-based laser networks.

David O Williams Rogers↗

Practical Implementation of GPU-based Computing at the Grid Edge for Resilience Scenarios

This paper presents a practical implementation of GPU-accelerated computing at the grid edge to enhance power system resilience through next-generation smart meters. Advanced Metering Infrastructure (AMI) systems rely predominantly on centralized processing architectures, which limit real-time response capabilities during grid disturbances. This work proposes the integration of GPU-enabled computational platforms directly within smart meter to enable local execution support for power system analytics, fault detection algorithms, and optimization routines. The proposed framework uses the Julia programming language to leverage highperformance parallel computing capabilities while maintaining code portability and development efficiency. We use two experimental scenarios to benchmark the computational feasibility of this approach: sparse linear system solutions representative of power flow analyses, and multi-stage production cost simulations incorporating unit commitment and economic dispatch operations. Results demonstrate that computationally intensive power system algorithms, such as those supporting resilience scenario calculations, can be effectively executed at the distribution edge using commercially available embedded GPU hardware. Keywords—GPU acceleration, edge computing, smart meters, grid resilience, AMI, resilience.

De Souza, Reubun [School of Electrical Engineering↗

Algorithmic construction of SSA-compatible extreme rays of the subadditivity cone and the N = 6 solution

We compute the set of all extreme rays of the 6-party subadditivity cone that are compatible with strong subadditivity. In total, we identify 208 new (genuine 6-party) orbits, 52 of which violate at least one known holographic entropy inequality. For the remaining 156 orbits, which do not violate any such inequalities, we construct holographic graph models for 150 of them. For the final 6 orbits, it remains an open question whether they are holographic. Consistent with the strong form of the conjecture in [1], 148 of these graph models are trees. However, 2 of the graphs contain a “bulk cycle”, leaving open the question of whether equivalent models with tree topology exist, or if these extreme rays are counterexamples to the conjecture. The paper includes a detailed description of the algorithm used for the computation, which is presented in a general framework and can be applied to any situation involving a polyhedral cone defined by a set of linear inequalities and a partial order among them to find extreme rays corresponding to down-sets in this poset.

AdS-CFT correspondence↗

Dynamics and lipid membrane coupling of the RAS-RAF complex revealed via multiscale simulations

To gain molecular and mechanistic insights into initiation of the RAS-RAF signaling cascade, we developed and used a combination of multiscale simulation and experimental approaches. The influence and impact of the membrane on RAS and RAF proteins is a factor we are just beginning to understand and appreciate in more detail. Molecular simulation is an ideal methodology to further study this complicated relationship between the membrane and associated proteins. Our previous work using Multiscale Machine-learned Modeling Infrastructure investigated different lipid compositions solely around the KRAS4b protein and the interplay between protein behavior and these membrane environments. Multiscale Machine-learned Modeling Infrastructure uses machine learning to couple adjacent simulation scales and has been efficiently scaled across some of the world’s largest high-performance computers. Recently, we have expanded this multiresolution framework to include the all-atom simulation scale and to incorporate the RAF RBDCRD domains. Here, we present the overall analysis results from this new simulation campaign comprising a mixture of RAS and RAF RBDCRD proteins. Approximately 35,000 coarse-grained and 10,000 all-atom molecular dynamics simulations were completed, sampled from a variety of protein/lipid composition configurations that were generated from a micron-scale continuum simulation containing hundreds of copies of the proteins. Our studies suggest that orientations of the RAS-RBDCRD complex on the membrane occupy distinct configurational states, and the spatial patterns of lipid arrangements around these different protein states are unique to each state. The extent and size of lipid “fingerprints” imposed on the membrane by the RAS-RBDCRD protein complex are significantly larger than observed for just the RAS protein on its own. These protein complexes strongly associate, but we do not observe statistically significant preferred protein-protein orientations. These observations indicate that spatial colocalization of RAS-RBDCRD proteins in the same vicinity may be assisted by specific membrane environments, acting to increase the probability of signaling complex formation.

Carpenter, Timothy S. [Lawrence Livermore National↗

Validation of Phasor-Domain Transmission and Distribution Co-simulation Against Electromagnetic Transient Simulation

The rapid deployment of renewable energy resources has led to the widespread use of power electronics in modern power systems. As these systems transition from being dominated by large synchronous machines to increasingly incorporating inverter-based resources (IBRs), traditional methods are becoming inadequate. Addressing this challenge, this paper introduces a scalable phasor-domain T\&D co-simulation framework based on open-source software. It focuses on the framework's validation against the PSCAD Electromagnetic Transient (EMT) analysis tool. The validation results demonstrate the framework's high-fidelity and a computational time speed-up of 60 to 100 times, marking a pioneering validation effort in T\&D co-simulation research.

Inverter-based resources, co-simulation, Electroma↗

Preparation of Neptunyl and Plutonyl Acetates To Access Nonaqueous Transuranium Coordination Chemistry

Uranyl diacetate dihydrate is a useful reagent for the preparation of uranyl (UO 2 2+ ) coordination complexes, as it is a well-defined stoichiometric compound featuring moderately basic acetates that can facilitate protonolysis reactivity, unlike other anions commonly used in synthetic actinide chemistry such as halides or nitrate. Despite these attractive features, analogous neptunium (Np) and plutonium (Pu) compounds are unknown to date. Here, in this study, a modular synthetic route is reported for accessing stoichiometric neptunyl(VI) and plutonyl(VI) diacetate compounds that can serve as starting materials for transuranic coordination chemistry. The new NpO 2 2+ and PuO 2 2+ complexes, as well as a corresponding molecular UO 2 2+ complex, are isomorphous in the solid state, and in solution show similar solubility properties that facilitate their use in synthesis. In both solid and solution state, the +VI oxidation state (O.S.) is maintained, as demonstrated by vibrational and optical spectroscopy, confirming that acetate anions stabilize the oxidizing, high-valent +VI states of Np and Pu as they do for the more stable U(VI). All three acetate salts readily react with a model diprotic ligand, affording incorporation of U(VI), Np(VI), and Pu(VI) cores into molecular coordination compounds that occurs concomitantly with elimination of acetic acid; the new complexes are high-valent, yet overall charge neutral, facilitating entry into nonaqueous chemistry by rational synthesis. Computational studies reveal that the dianionic ligand framework assists in stabilizing the +VI O.S. via donation to the 5f shells of the actinides, highlighting the potential usefulness of protonolysis reactivity toward preparation of stabilized high-valent transuranic species.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Computational investigation of water glasses using machine-learning potentials

The molecular origins of water’s anomalous properties have long been a subject of scientific inquiry. The liquid–liquid phase transition hypothesis, which posits the existence of distinct low-density and high-density liquid states separated by a first-order phase transition terminating at a critical point, has gained increasing experimental and computational support and offers a thermodynamically consistent framework for many of water’s anomalies. However, experimental challenges in avoiding crystallization near the postulated liquid–liquid critical point have focused attention to water’s canonical glassy states: low-density and high-density amorphous ice. Here, we use two Deep Potential machine-learning models, trained on the Strongly Constrained and Appropriately Normed density functional and the highly accurate Many-Body Polarizable potential, to conduct an investigation of water’s glassy phenomenology based on quantum mechanical calculations. Despite not being explicitly trained on amorphous ices, both models accurately capture the structure and transformation of the water glasses, including their interconversion along different thermodynamic paths. Isobaric quenching of liquid water at various pressures generates a continuum of intermediate amorphous ices and density fluctuations increase near the liquid–liquid critical pressure. The glass transition temperatures of the amorphous ices produced at different pressures exhibit two distinct branches, corresponding to low-density and high-density amorphous ice behaviors, consistent with experiment and the liquid–liquid transition hypothesis. Extrapolating transformation pressures from isothermal compressions to experimental compression rates brings our simulations into excellent agreement with data. Our findings demonstrate that machine-learning potentials trained on equilibrium phases can effectively model nonequilibrium glassy behavior and pave the way for studying long-timescale, out-of-equilibrium processes with quantum mechanical accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗