Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer architecture”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Continuous Counter‐Current Microfluidic Liquid–Liquid Extraction Achieved Using a Pair of Wettable Screen Meshes

Continuous counter‐current microfluidic liquid–liquid extraction performs separations by flowing immiscible liquids in opposing directions within a single flow channel. In principle, this flow arrangement enables a large number of theoretical separation units in a small footprint, without using interstage valving, pumping, and phase separation. Despite its potential for excellent separation performance, this microfluidic scheme rarely appears in literature due to the requirement for capillary forces to be greater than hydrodynamic forces for stable flow. We present a novel microfluidic device and flow approaches that overcome this force‐balance challenge, enabling stable, long‐duration continuous counter‐current flow. Additionally, we cover a suite of methodologies for quantifying the performance of the microfluidic device, revealing the number of theoretical equilibrium stages achieved. The enabling technologies include a woven mesh screen‐based microfluidic device architecture that is easily fabricated outside of a clean room, surface functionalization strategies to promote conjugate (organic/aqueous) wettability, flow approaches to eliminate bubbles and carryover, and computer‐aided flow automation with optical measurement of extraction performance. The reported experiments lasted for over 36 h, terminated only at experiment conclusion, where the device still exhibited good performance. Automated Raman spectroscopy was used for solute quantitation of the ternary system tert‐butanol in a toluene/water matrix, a ternary system that was specifically chosen to analyze the device's performance with a small solute partition ratio and to enable in‐line Raman measurements of solute concentrations in both phases. The microfluidic device possessed a 55 mm contact length and a 38.5 µL internal volume. During counter‐current flow, we observed approximately 37 equilibrium stages (37 ± 13) based on a best‐fit of the solute fraction remaining in the aqueous phase using a Kremser Group Method analysis.

36 MATERIALS SCIENCE↗

An end-to-end deep learning method for solving nonlocal Allen–Cahn and Cahn–Hilliard phase-field models

Here, we propose an efficient end-to-end deep learning method for solving nonlocal Allen–Cahn (AC) and Cahn–Hilliard (CH) phase-field models. One motivation for this effort emanates from the fact that discretized partial differential equation-based AC or CH phase-field models result in diffuse interfaces between phases, with the only recourse for remediation is to severely refine the spatial grids in the vicinity of the true moving sharp interface whose width is determined by a grid-independent parameter that is substantially larger than the local grid size. In this work, we introduce non-mass conserving nonlocal AC or CH phase-field models with regular, logarithmic, or obstacle double-well potentials. Because of non-locality, some of these models feature totally sharp interfaces separating phases. The discretization of such models can lead to a transition between phases whose width is only a single grid cell wide. Another motivation is to use deep learning approaches to ameliorate the otherwise high cost of solving discretized nonlocal phase-field models. To this end, loss functions of the customized neural networks are defined using the residual of the fully discrete approximations of the AC or CH models, which results from applying a Fourier collocation method and a temporal semi-implicit approximation. To address the long-range interactions in the models, we tailor the architecture of the neural network by incorporating a nonlocal kernel as an input channel to the neural network model. We then provide the results of extensive computational experiments to illustrate the accuracy, predictive capabilities, and cost reductions of the proposed method.

42 ENGINEERING↗

Cost-efficient finite-volume high-order schemes for compressible magnetohydrodynamics

We present an efficient dimension-by-dimension finite-volume method which solves the adiabatic magnetohydrodynamics equations at high discretization order, using the constrained-transport approach on Cartesian grids. Results are presented up to tenth order of accuracy. The algorithmic architecture of this method is very close to that of commonly employed second-order schemes: it requires only one reconstructed value per face for each computational cell, independently of the scheme's order. This property is highly beneficial for the numerical efficiency. It results from reusing the required values already available in neighboring grid cells, in contrast to standard algorithms that require a number of reconstructions and evaluations which increases with the scheme's order of accuracy. At a given resolution, these high-order schemes present significantly less numerical dissipation than commonly employed lower-order approaches. Thus, results of comparable accuracy are achievable at a substantially coarser resolution, yielding overall performance gains. We also present a way to include physical dissipative terms: viscosity, magnetic diffusivity and cooling functions, respecting the finite-volume and constrained-transport frameworks. Benefits of this method are shown through applications in turbulent flows.

97 MATHEMATICS AND COMPUTING↗

Signal propagation in reversible digital mechanics

Digital mechanics explores information processing through binary, mechanical circuits. This work demonstrates a flexural, mechanical integrated circuit (m-IC) that achieves reversible, non-reciprocal signal propagation through integrated AND logic and memory. Our approach exploits sequential bistable transitions with symmetric energy wells, tunable stiffness, impedance matching, and AND gate non-linearity, to enable signal propagation, repeatability, and reversibility. We present a generalized model of logic kinematics and energetics, validated experimentally, to study energy flows, quantify energetic limits, and identify operating regimes for reversible logic. Macro-scale experiments confirm propagation dynamics, and new fabrication methods extend the architecture to micro-scale devices. By achieving controlled, reversible signal transmission across interconnected logic and memory, this work establishes a scalable platform for robust mechanical computing and adaptive sensing.

Johnson, Hilary A. [Lawrence Livermore National La↗

NeuroCoreX: An Open-Source FPGA-Based Spiking Neural Network Emulator with On-Chip Learning

Spiking Neural Networks (SNNs) are computational models inspired by the event-driven communication and connectivity patterns of biological neural circuits. They enable high energy efficiency and natural support for diverse architectures ranging from layered networks to small-world and graphstructured topologies. In this work, we introduce NeuroCoreX, an open-source, FPGA-based spiking neural network emulator that provides real-time, on-chip learning and flexible network organization. NeuroCoreX supports both feedforward sensory inputs streamed directly from sensors or PCs via UART and recurrent on-chip connectivity, enabling simultaneous processing and learning from external stimuli and internal network dynamics-capabilities rarely available in existing FPGA SNN platforms. The system implements a Leaky Integrate-and-Fire (LIF) neuron model with current-based synapses and supports pair-based STDP learning on both feedforward and recurrent synapses. A lightweight Python interface enables interactive configuration, live monitoring, weight read-back, and experiment control. Importantly, NeuroCoreX is tightly integrated with the SuperNeuroMAT simulator, allowing SNN models to be transferred seamlessly from software to hardware for hardware-in-the-loop development. By combining real-time plasticity, flexible connectivity, and an open-source VHDL implementation, NeuroCoreX provides an extensible and accessible platform for neuromorphic research, algorithm-hardware co-design, and energy-efficient edge intelligence.

Gautam, Ashish [ORNL]↗

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol↗

Simplifying activations with linear approximations in neural networks

A key step in Neural Networks is activation. Among the different types of activation functions, sigmoid, tanh, and others involve the usage of exponents for calculation. From a hardware perspective, exponential implementation implies the usage of Taylor series or repeated methods involving many addition, multiplication, and division steps, and as a result are power-hungry and consume many clock cycles. We implement a piecewise linear approximation of the sigmoid function as a replacement for standard sigmoid activation libraries. This approach provides a practical alternative by leveraging piecewise segmentation, which simplifies hardware implementation and improves computational efficiency. In this paper, we detail piecewise functions that can be implemented using linear approximations and their implications for overall model accuracy and performance gain. Our results show that for the DenseNet, ResNet, and GoogLeNet architectures, the piecewise linear approximation of the sigmoid function provides faster execution times compared to the standard TensorFlow sigmoid implementation while maintaining comparable accuracy. Specifically, for MNIST with DenseNet, accuracy reaches 99.91% (Piecewise) vs. 99.97% (Base) with up to 1.31x speedup in execution time. For CIFAR-10 with DenseNet, accuracy improves to 98.97% (Piecewise) vs. 99.40% (Base) while achieving 1.24x faster execution. Similarly, for CIFAR-100 with DenseNet, the accuracy is 97.93% (Piecewise) vs. 98.39% (Base), with a 1.18x execution time reduction. These results confirm the proposed method’s capability to efficiently process large-scale datasets and computationally demanding tasks, offering a practical means to accelerate deep learning models, including LSTMs, without compromising accuracy.

Activation function↗

High-Resolution Simulations of Geological CO 2 Injection: Application to the SPE11 Benchmark

Geological carbon sequestration (GCS) will play a critical role in decarbonization and in facilitating the transition to clean energy systems. Because CO 2 is highly mobile, ensuring its safe and permanent injection into subsurface geological formations involves monitoring over larger spatial domains and longer time periods than is typical for hydrocarbon reservoirs. This can benefit from simulation tools capable of modeling key CO 2 trapping mechanisms, particularly those optimized for speed and scalability on high-performance computing systems. Using isothermal versions of the SPE11B and SPE11C benchmark cases, we conduct a mesh refinement study simulating CO 2 injection into kilometer-scale rock formations at centimeter resolution with the GEOS open-source simulation framework. We focus on how mesh refinement improves the accuracy of convective mixing in both 2D and 3D simulations. The computational costs associated with achieving a converged solution highlight the need for predictive upscaling techniques. A systematic performance scaling analysis—including both central processing unit (CPU) and graphics processing unit (GPU) architectures—complements the “Results” section.

Geosciences↗

MS25: Materials Science-Focused Benchmark Data Set for Machine Learning Interatomic Potentials

Here, we present MS25, a benchmark data set for evaluating machine learning interatomic potentials (MLIPs) across diverse materials-relevant systems including MgO surfaces, liquid water, zeolites, a catalytic Pt surface reaction, high-entropy alloys (HEAs), and disordered Zr-oxides. Five MLIP architectures (MACE, NequIP, Allegro, MTP, and Torch-ANI) are trained and tested, focusing not only on traditional metrics (energies, forces, and stresses) but also explicitly validating derived physical observables such as lattice constants, volumes, and reaction barriers. We find that most models reach comparable accuracy on standard error metrics across the simple systems, although equivariant MLIPs offer 1.5–2× improvements over nonequivariant MLIPs in energy and force error for structurally complex or compositionally disordered environments such as HEAs and Zr–O systems. Our analysis highlights that low errors in energy and force predictions do not guarantee reliable observables, emphasizing the necessity of explicit validation. We demonstrate limitations in cross-framework transferability, as models trained on one zeolite framework (CHA) fail to reliably generalize to predictions of structurally distinct frameworks (e.g., MFI). Size-extensive tests show some dependence on system size for MgO, resulting from forced periodicity. The HEA and Zr–O data sets are identified as challenging tests for future benchmarks and MLIP model architecture developments as they show significant differentiation in error between MLIP architectures and are still relatively difficult at 1000 training images. Moving forward, we recommend that benchmarking efforts shift their focus from marginal accuracy improvements in energy and force errors toward identifying and understanding model failure modes, rigorously assessing transferability, and evaluating how their errors affect observable predictions. For researchers looking to choose an MLIP architecture, we suggest selecting equivariant MLIP architectures if the complexity of the system is a challenge. For simple materials problems, auxiliary features such as integration with molecular dynamics engines, trade-offs between computational data set generation cost vs MLIP inference speed, and framework integration may play a more important decision factor than small differences in error metrics that are unlikely to matter for production-level research.

chemical structure↗

Patchy nanoparticles by atomic stencilling

Stencilling, in which patterns are created by painting over masks, has ubiquitous applications in art, architecture and manufacturing. Modern, top-down microfabrication methods have succeeded in reducing mask sizes to under 10 nm, enabling ever smaller microdevices as today’s fastest computer chips. Meanwhile, bottom-up masking using chemical bonds or physical interactions has remained largely unexplored, despite its advantages of low cost, solution-processability, scalability and high compatibility with complex, curved and three-dimensional (3D) surfaces. Here we report atomic stencilling to make patchy nanoparticles (NPs), using surface-adsorbed iodide submonolayers to create the mask and ligand-mediated grafted polymers onto unmasked regions as ‘paint’. We use this approach to synthesize more than 20 different types of NP coated with polymer patches in high yield. Polymer scaling theory and molecular dynamics (MD) simulation show that stencilling, along with the interplay of enthalpic and entropic effects of polymers, generates patchy particle morphologies not reported previously. These polymer-patched NPs self-assemble into extended crystals owing to highly uniform patches, including different non-closely packed superlattices. We propose that atomic stencilling opens new avenues in patterning NPs and other substrates at the nanometre length scale, leading to precise control of their chemistry, reactivity and interactions for a wide range of applications, such as targeted delivery, catalysis, microelectronics, integrated metamaterials and tissue engineering.

36 MATERIALS SCIENCE↗

Repositioning Quantum Cellular Automata for Dependable Quantum-Classical Systems

Quantum Cellular Automata (QCA) provides a structured model of distributed quantum computation with inherent locality and regularity properties that are suited to dependable execution. However, QCA remain largely absent from discussions on reproducibility, fault management, and orchestration in heterogeneous quantum-classical systems. We propose a dual-axis framework that situates QCA within both computation and physical realizability, revealing regions where robust, scalable, and hardware-constrained quantum dynamics may reside. By revisiting prior results through the lens of reproducibility and architecture resilience, we suggest that QCA offers a potential substrate for benchmarking and system-level co-design.

Stapleton, Nicholas [ORNL] (ORCID:0000000335305325↗

Surrogate Neural Architecture Codesign Package (SNAC-Pack)

Neural architecture search (NAS) is a powerful approach for automating model design, but existing methods often optimize for accuracy alone or rely on proxy metrics such as bit operations (BOPs) that correlate poorly with hardware cost. This gap is particularly large for FPGA deployment, where cost is dominated by a multi-dimensional budget of lookup tables, DSPs, flip-flops, BRAM, and latency. We present the Surrogate Neural Architecture Codesign Package (SNAC-Pack), an open-source AutoML framework for hardware-aware neural architecture codesign and end-to-end FPGA deployment. SNAC-Pack runs a multi-objective global search with Optuna and NSGA-II, loading trials to a shared SQLite store that enables parallel workers across compute nodes. A hardware surrogate model outputs per-trial resource and latency estimates, avoiding the synthesis cost that would otherwise dominate the search loop. A local search stage then applies quantization-aware training (QAT) together with iterative magnitude pruning in a combined compression loop, after which the final model is synthesized to FPGA firmware via the hls4ml Python library. A YAML configuration and an optional agentic frontend let users run the pipeline on new datasets without modifying the framework. We demonstrate SNAC-Pack on jet classification at the Large Hadron Collider and superconducting qubit readout, discovering compact architectures that match or exceed strong baselines on the task metric while reducing FPGA resource utilization and, in the qubit readout case, reducing the design space exploration process from months of manual fine-tuning to hours of automated search.

Weitz, Jason [UC, San Diego]↗

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗

Memcomputing the Spectrum of Correlated Quantum Systems

The goals and objectives of the grant DE‐SC0020892 were to apply a new computing paradigm, MemComputing, to efficiently simulate properties of correlated quantum systems. The method has been applied to a wide set of problems ranging from quantum state tomography to finding the ground state of correlated systems. In all cases, substantial advantages compared to state-of-the-art approaches have been obtained. The project has also led to the suggestion of the transformer architecture (used nowadays in large-language models) as an efficient quantum state representation, and a better understanding of the role of memory in the generation of long-range order in neural networks. This grant has supported the work of a PhD student, inspired a new class on unconventional computing taught at the University of California, San Diego and has generated several peer-reviewed papers.

97 MATHEMATICS AND COMPUTING↗

Best of both worlds: Enforcing detailed balance in machine learning models of transition rates

The slow microstructural evolution of materials often plays a key role in determining material properties. When the unit steps of the evolution process are slow, direct simulation approaches such as molecular dynamics become prohibitive and Kinetic Monte-Carlo (kMC) algorithms, where the state-to-state evolution of the system is represented in terms of a continuous-time Markov chain, are instead frequently relied upon to efficiently predict long-time evolution. The accuracy of kMC simulations however relies on the complete and accurate knowledge of reaction pathways and corresponding kinetics. This requirement becomes extremely stringent in complex systems such as concentrated alloys where the astronomical number of local atomic configurations makes the a priori tabulation of all possible transitions impractical. Machine learning models of transition kinetics have been used to mitigate this problem by enabling the efficient on-the-fly prediction of kinetic parameters. While conventional KMC methods based on transition state theory naturally yield reversible dynamics that exactly obey the detailed balance criterion, providing strong guarantees on the properties of the stationary distribution, many recently-proposed ML-based approaches to barrier predictions provide no such guarantees. In this study, we derive conditions under which physics-informed ML architectures exactly enforce the detailed balance condition by construction, even when relying on non-extensive descriptions of states in terms of local environments around mobile defects. In conclusion, using the diffusion of a vacancy in a concentrated alloy as an example, we show that such ML architectures also exhibit superior performance in terms of prediction accuracy, demonstrating that the imposition of physical constraints can facilitate the accurate learning of barriers at no increase in computational cost.

36 MATERIALS SCIENCE↗

PvaPy streaming framework for real-time data processing

User facility upgrades, new measurement techniques, advances in data analysis algorithms as well as advances in detector capabilities result in an increasing amount of data collected at X-ray beamlines. Some of these data must be analyzed and reconstructed on demand to help execute experiments dynamically and modify them in real time. In turn, this requires a computing framework for real-time processing capable of moving data quickly from the detector to local or remote computing resources, processing data, and returning results to users. In this paper, we discuss the streaming framework built on top of PvaPy, a Python API for the EPICS pvAccess protocol. We describe the framework architecture and capabilities, and discuss scientific use cases and applications that benefit from streaming workflows implemented on top of this framework. We also illustrate the framework's performance in terms of achievable data-processing rates for various detector image sizes.

EPICS pvAccess↗

Deep learning for time series forecasting: a survey of recent advances

Time series forecasting plays a critical role in numerous real-world applications, such as finance, healthcare, transportation, and scientific computing. In recent years, deep learning has become a powerful tool for modeling complex temporal patterns and improving forecasting accuracy. This survey provides an overview of recent deep learning approaches for time series forecasting, involving various architectures including RNNs, CNNs, GNNs, transformers, large language models, MLP-based models, and diffusion models. We first identify key challenges in the field, such as temporal dependency, efficiency, and cross-variable dependency, which drive the development of forecasting techniques. Then, the general advantages and limitations of each architecture are discussed to contextualize their adaptation in time series forecasting. Furthermore, we highlight promising design trends like multi-scale modeling, decomposition, and frequency-domain techniques, which are shaping the future of the field. This paper serves as a compact reference for researchers and practitioners seeking to understand the current landscape and future trajectory of deep learning in time series forecasting.

97 MATHEMATICS AND COMPUTING↗