Search NASASearch

SEARCH · Search NASA

Results for “DATA PROCESSOR”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

53 records · Page 3

Status of the Mu2e calorimeter readout electronics

The Mu2e experiment [1] at Fermilab will search for the neutrino-less coherent conversion of a muon into an electron in the field of a nucleus. Mu2e detectors comprise a straw tracker, an electromagnetic calorimeter and a veto for cosmic rays. The calorimeter employs 1348 Cesium Iodide crystals readout by silicon photomultipliers and fast front-end and digitization electronics. The front-end electronics consists of two discrete readout circuits (AMP-HV) for each crystal. These provide the amplification, shaping stage and linear regulation of the SiPM bias voltage and monitoring. The SiPM and front-end control electronics is implemented in a battery of mezzanine boards each equipped with an ARM processor that controls a group of 20 Amp-HV circuits distributing the low voltage and the high-voltage. The electronic is hosted in crates located on the external surface of calorimeter disks. The crates also host the waveform digitizer board (DIRAC) that performs digitization of the front end signals and transmit the digitized data to the Mu2e DAQ. Calorimeter electronic is hosted inside the cryostat and must sustain very high radiation and magnetic field so it was necessary to fully qualify it. The system design and quality assurance procedures will be reviewed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

A Cryogenic Muon Tagging System Integrated with a Superconducting Qubit Device for Radiation-Induced Error Mitigation

Superconducting qubits are highly sensitive to ionizing radiation, which can induce correlated errors and limit scalable fault-tolerant quantum computing. In particular, cosmic-ray muons can deposit energy in the substrate, generating phonon bursts that break Cooper pairs and produce quasiparticles, leading to correlated decoherence events across multiple qubits. We present the development of a cryogenic muon tagging system based on Kinetic Inductance Detectors (KIDs) and its integration with superconducting quantum hardware. Originally developed within the ACE-SuperQ project and validated as a standalone detector, the system demonstrated a muon tagging efficiency of approximately 90% and excellent agreement with Monte Carlo simulations. Building on this validation, the tagging system has been integrated with a multi-qubit superconducting chip operated in a dilution refrigerator. The detector configuration consists of a multi-layer KID stack arranged above and below the quantum device, enabling time-coincident identification of muon-induced events within the same cryogenic environment. The integrated setup has been successfully commissioned, enabling simultaneous operation of the qubit chip and the muon tagging system. A first measurement campaign has been carried out, and preliminary data show time-correlated events between the muon tagging detectors and the qubit readout. A quantitative analysis of radiation-induced effects on qubit performance is currently ongoing. This work represents a step toward the implementation of event-level radiation tagging as a tool for characterizing and potentially mitigating correlated errors in superconducting quantum processors, while establishing a modular platform for future studies at the interface between particle physics and quantum information science.

Roy, Tanay [Fermilab] (ORCID:000000019442862X)

SQMS Quantum R&D in Machine Learning, Optimization and Sensing beyond Fundamental Physics Applications

This newly formed team at SQMS under the Ecosystem Thrust is looking to develop capabilities impacting societal advances outside the core domain of HEP and condensed matter physics. We explicitly leverage the experimental and algorithmic innovations developed across all groups as well as connect to broad-scope external projects of the diverse team of PIs. As the inaugural set of projects, we are studying numerically quantum machine learning models inspired by efficiently trainable echo-state and orthogonal neural networks and developing designs for related experiments to be performed on quantum processors based on SQMS SRF cQED technology and Rigetti s transmon arrays. Investigated models exploit ideas and lessons learned from multiple prior work by SQMS team members in a variety of internal and external activities [R1]. Target initial applications include noisy signal processing, potentially captured by quantum sensors or noisy QPUs, as well as simulation and classification of healthcare data. For instance, image reconstruction of the brain s electrical properties by solving the inverse Maxwell equation problem with uncertainty [R2] through a hybrid quantum-classical physics-informed architecture for time-dependent processes [R3]. The group is also investigating the application and development of novel quantum sensors based on magnetic levitation of a superconducting sphere coupled to a superconducting qubit. This coupling enables high-precision measurements of the position of the sphere, which can be used for sensitive detection of forces, enabling practical applications such as gravimetry for geophysics analysis, or accelerometry for GPS-denied navigation [R4] [R1] Rieffel, Eleanor G., Ata Akbari Asanjan, M. Sohaib Alam, Namit Anand, David E. Bernal Neira, Sophie Block, Lucas T. Brady et al. "Assessing and advancing the potential of quantum computing: A NASA case study." Future Generation Computer Systems (2024). [R2] Yu, X., Serrall s, J.E., Giannakopoulos, I.I., Liu, Z., Daniel, L., Lattanzi, R. and Zhang, Z., 2023. Pifon-ept: Mr-based electrical property tomography using physics-informed fourier networks. IEEE Journal on Multiscale and Multiphysics Computational Techniques. [R3] Wudarski, Filip, Daniel OConnor, Shaun Geaney, Ata Akbari Asanjan, Max Wilson, Elena Strbac, P. Aaron Lott, and Davide Venturelli. "Hybrid quantum-classical reservoir computing for simulating chaotic systems." arXiv preprint arXiv:2311.14105 (2023). [R4] Higgins, Gerard, Saarik Kalia, and Zhen Liu. "Maglev for dark matter: Dark-photon and axion dark matter sensing with levitated superconductors." Physical Review D 109.5 (2024): 055024.

Venturelli, Davide

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES

Internship Work Report

I worked on two projects during my summer internship at Sandia. My official title was “Intern - Mission Tech Electrical Eng./Computer Eng.- R&D Undergraduate Summer.” I worked at the central location, which is Albuquerque, New Mexico. The department you are placed in at Sandia doesn’t always correspond to the people you will be working with. For example, I only directly worked with one person from my department this summer. On one of my projects, I worked with a diverse team of engineers from many different departments. On my other project, I mainly worked with two departments, as the project had two distinct parts. As mentioned earlier, I worked on two projects during my summer at Sandia. The first project focused on a lightweight embedded controller in an advanced FPGA System-on-Chip for radar signal processing applications. The term “controller” refers to a hardware device that directs the flow of data between two entities. An FPGA is a reprogrammable integrated circuit (as opposed to an integrated circuit with one purpose). An FPGA was used on this project so in order to protype various ideas for our System-on-Chip. My role on the project was to implement designs on the fabric of the FPGA and design a state machine (written in C) for the processor. My second project was also heavily involved with embedded systems but had a different application. It focused on using a Newton-Raphson control algorithm to stabilize an inverted pendulum using a novel microcontroller. The pendulum dynamics were derived, and it was successfully simulated in MATLAB. I worked on integrating the microcontroller with the inverted pendulum machinery, and converting the Newton-Raphson control algorithm from MATLAB into C. The inverted pendulum was successfully stabilized using a simple PID controller and industry-standard microcontroller. The project is still ongoing, and the team is gearing up for more tests using the novel Newton-Raphson control algorithm and novel microcontroller

42 ENGINEERING

Low Power, Radiation Resilient Synchronous Edge Processing for Remote Monitoring

Next-generation space remote sensing systems may be equipped with imaging arrays that sense data at a rate that outstrips the processing capability of any computing hardware that can operate within a satellite’s power budget. This project developed novel convolutional and recurrent neural networks to detect and estimate point-like events amid clutter, and investigated their efficient and accurate implementation on analog in-memory computing systems that are 10-1000× more energy-efficient than digital processors. This project leveraged two memory devices at different levels of technological maturity: a large-scale analog computing prototype using commercial SONOS charge-trap memory, and electrochemical memory (ECRAM) with intrinsic radiation hardness. We experimentally demonstrated end-to-end analog processing of our neural networks on SONOS and characterized the radiation response of both SONOS and ECRAM. We advanced the state-of-the-art in ECRAM precision and reliability, and developed co-design methods to enable accurate long-term operation of SONOS analog accelerators in space radiation environments.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

3-center and 4-center 2-particle Gaussian AO integrals on modern accelerated processors

We report an implementation of the McMurchie–Davidson (MD) algorithm for 3-center and 4-center 2-particle integrals over Gaussian atomic orbitals (AOs) with low and high angular momenta l and varying degrees of contraction for graphical processing units (GPUs). This work builds upon our recent implementation of a matrix form of the MD algorithm that is efficient for GPU evaluation of 4-center 2-particle integrals over Gaussian AOs of high angular momenta (l ≥ 4) [A. Asadchev and E. F. Valeev, J. Phys. Chem. A 127, 10889–10895 (2023)]. The use of unconventional data layouts and three variants of the MD algorithm allow for the evaluation of integrals with double precision and sustained performance between 25% and 70% of the theoretical hardware peak. Performance assessment includes integrals over AOs with l ≤ 6 (a higher l is supported). Preliminary implementation of the Hartree–Fock exchange operator is presented and assessed for computations with up to a quadruple-zeta basis and more than 20 000 AOs. The corresponding C++ code is part of the experimental open-source LibintX library available at https://github.com/ValeevGroup/libintx.

Chemistry

Rapid organic carbon spiraling in a headwater stream linked with streamflow, biogeochemistry, and canopy phenology

Headwater streams are abundant worldwide and important to global biogeochemical cycles, serving as critical processors and transporters of C. C spiraling is a useful way to understand the retention and mineralization of organic C (OC) in streams. However, analyses of seasonal and interannual variability in OC spiraling are currently limited. In this study, we aimed to understand the temporal patterns and driving mechanisms of OC spiraling, which will inform our understanding of future OC changes under climate change. We used 7 y of daily data in a small headwater stream (Walker Branch, Tennessee, USA) to assess seasonal and interannual variability in OC spiraling length (S OC ) and mineralization velocity (v fOC ), as well as their potential related variables. On average, S OC in Walker Branch was ~10× shorter than in previously studied small streams, indicating strong connections between the water column and the benthic environment where OC mineralization mostly takes place. OC spiraling was faster during the more biologically active periods of spring and autumn compared with more elongated OC spiraling in summer and winter, when OC retention was lower and downstream transport was higher. Gross primary production (GPP) was most strongly related to S OC and v fOC . Photosynthetically active radiation (PAR) and NO 3 − were also positively and negatively related to v fOC , respectively. Trends toward earlier and longer canopy cover and reduced GPP and PAR may result in longer S OC and slower v fOC , reducing localized instream processing of OC and potentially shunting more OC downstream. However, long-term observations indicate reduced NO 3 − at Walker Branch, suggesting opposing effects to those of GPP and PAR, leading to faster v fOC and greater OC retention. Time-series analyses of OC spiraling in streams can enhance our understanding of current and future responses of OC processing and downstream transport to climate change, as well as implications for downstream OC dynamics.

biological activity

Captan+X Data Converter Integration

Fermi National Accelerator Laboratory's CAPTAN (Compact And Programmable daTa Acquisition Node) series provides a flexible hardware platform for data acquisition across a range of experiments and facilities. The latest iteration, CAPTAN+X, is built around a Kintex-7 FPGA supporting four FPGA Mezzanine Card (FMC) connections. As part of a broader laboratory effort to bring facility systems under a Model-Based Systems Engineering (MBSE) framework, CAPTAN+X is one of several systems slated to be incorporated into this modeling environment in the near term. A necessary step toward that goal is incorporating the platform's core functionality, which centers on integration with the LXD31K4 FMC, a data converter module combining dual AD9652 analog-to-digital converters and dual AD9142A digital-to-analog converters. Achieving compatibility required resolving pin-mapping conflicts between the LXD31K4's High Pin Count connector and the CAPTAN+X's available pin types, adapting a Board Support Project originally written for an UltraScale-class evaluation board to the Kintex-7 architecture, replacing incompatible primitives, restructuring clock distribution, and manually configuring chip initialization in place of an unsupported soft-processor-based approach. Functional verification of the ADC and DAC channels, followed by closed-loop testing combining both converters with real-time filtering, confirmed correct operation of the integrated system. These results establish a working hardware and firmware baseline for the CAPTAN+X platform, positioning it for future inclusion in the laboratory's growing MBSE modeling effort.

Espinoza, David [Illinois U., Urbana (main)]

CAPTAN+X Data Converter Integration

Fermi National Accelerator Laboratory's CAPTAN (Compact And Programmable daTa Acquisition Node) series provides a flexible hardware platform for data acquisition across a range of experiments and facilities. The latest iteration, CAPTAN+X, is built around a Kintex-7 FPGA supporting four FPGA Mezzanine Card (FMC) connections. As part of a broader laboratory effort to bring facility systems under a Model-Based Systems Engineering (MBSE) framework, CAPTAN+X is one of several systems slated to be incorporated into this modeling environment in the near term. A necessary step toward that goal is incorporating the platform's core functionality, which centers on integration with the LXD31K4 FMC, a data converter module combining dual AD9652 analog-to-digital converters and dual AD9142A digital-to-analog converters. Achieving compatibility required resolving pin-mapping conflicts between the LXD31K4's High Pin Count connector and the CAPTAN+X's available pin types, adapting a Board Support Project originally written for an UltraScale-class evaluation board to the Kintex-7 architecture, replacing incompatible primitives, restructuring clock distribution, and manually configuring chip initialization in place of an unsupported soft-processor-based approach. Functional verification of the ADC and DAC channels, followed by closed-loop testing combining both converters with real-time filtering, confirmed correct operation of the integrated system. These results establish a working hardware and firmware baseline for the CAPTAN+X platform, positioning it for future inclusion in the laboratory's growing MBSE modeling effort.

Espinoza, David [Illinois U., Urbana (main)]

Evaluation of New Additions to OLI Software in Predicting Mercuric and Mercurous Species in Liquid Waste Operations

Speciation of mercury during the pretreatment steps of tank waste processing is critical to successful mercury removal prior to vitrification during Liquid Waste Operations (LWO) at SRS. OLI software has been used to predict mercury speciation and activity throughout LWO. The OLI software operates based on a thermodynamic framework called the Mixed Solvent Electrolyte (MSE) framework. The MSE framework allows prediction in theoretically infinitely dilute to concentrated mixtures (e.g., purely solute solutions). Before modification to the MSE framework databanks, certain critical mercury species were missing in the MSE databank, and some thermodynamic data needed to be updated for the OLI software to accurately predict mercury chemical species in SRS waste tanks. To better reflect streams across LWO, new mercury species were integrated into the MSE database. To evaluate the changes to the OLI MSE framework per the Technical Task Request (TTR) and the Task Technical and Quality Assurance Plan (TTQAP), waste stream compositions from Tanks 38, 43, and Tank 50 decontaminated salt solution (DSS) were used as model inputs. Models were developed and executed using both the old and new databases. Compositional analyses from caustic Tank 50 DSS and caustic Tanks 38 and 43 were used as the input streams. These streams represent the most comprehensive chemical data sets where both mercury and tank constituents were measured together. Results for Tank 50 DSS predict HgO as the predominant species in both databases. Both methyl and dimethyl Hg species are present when the new database is ‘on’ and are not predicted with the new database turned ‘off’. The new database predicts a greater amount of HgO and a greater fraction of it in the solid phase. Pourbaix diagrams (potential vs. pH) generated for each Tank 50 DSS were identical regardless of which database was used. Elemental Hg and HgO were predicted in the water stable region under basic conditions. Tanks 38 and 43 follow similar trends as the Tank 50 DSS models. Unlike Tanks 38 and 50 DSS, the Tank 43 Pourbaix plot shows a region of stability for an aqueous HgOHCO3 - species between approximately pH 7-11. In all streams, when MeHg+ is included in the inputs, the new database predicts aqueous MeHgOH as the dominant species. If elemental or dimethyl mercury is in the waste stream, the new database model predicts they are unchanged and remain in those states and quantities. Additionally, the total mercury values are reported for both the measured input data and the OLI output data for all considered tanks. The summary indicates that the percentage error between the measured and calculated values is less than 1% in all cases The reconciliations and generation of the Pourbaix diagrams for Tank 50 DSS took approximately ten times longer with the new database ‘on’. In addition, over the course of that time, models with the new database ‘on’ were more likely to crash or display an error. Some modest performance improvements were noted when modeling with an i7 processor versus an i5. An example error is found in Appendix A. Furthermore, Appendix B provides V&V for two chemical systems analyzed with the OLI software, results were satisfactory. It is recommended to utilize the new databases (i.e., HCO.ddb and SR-Hg.ddb) in future Savannah River Mission Completion applications of OLI to represent pseudo steady-state. Furthermore, the integration and utilization of the new databases (i.e., HCO.ddb and SR-Hg.ddb) in modeling applications (e.g., Aspen) is also recommended.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING

Analog In-Memory Computing for the Synthetic Aperture Radar Polar Format Algorithm

As the utility of synthetic aperture radar (SAR) systems increases in autonomous vehicles, satellites, and other power- and space-constrained edge applications, there is a growing need for processors that can form SAR images at low power. In recent years, analog in-memory compute (AIMC) has shown immense promise for accelerating neural networks and other matrix-vector multiplication (MVM) heavy workloads at the edge. Here, in this work, we examine how the polar format algorithm (PFA), a popular SAR image formation algorithm, can be mapped to these AIMC systems. The PFA maps readily onto analog MVMs because it primarily consists of two linear operations: interpolation of frequency-domain data to a Cartesian grid, followed by a 2-D Fourier transform. This work presents two approaches to map the interpolation operation onto MVMs in analog hardware: a chirp transform and a modified form of sinc interpolation. These mappings introduce algorithmic errors, and their effect on the quality of SAR image formation is examined, both quantitatively and qualitatively. In addition, the impact of errors introduced by the analog hardware is explored to determine which approach is optimal under varying assumptions about the underlying analog memory devices and circuits.

Analog computing

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING

Refactoring the elastic–viscous–plastic solver from the sea ice model CICE v6.5.1 for improved performance

This study focuses on the performance of the elastic–viscous–plastic (EVP) dynamical solver within the sea ice model, CICE v6.5.1. The study has been conducted in two steps. First, the standard EVP solver was extracted from CICE for experiments with refactored versions, which are used for performance testing. Second, one refactored version was integrated and tested in the full CICE model to demonstrate that the new algorithms do not significantly impact the physical results. The study reveals two dominant bottlenecks, namely (1) the number of Message Parsing Interface (MPI) and Open Multi-Processing (OpenMP) synchronization points required for halo exchanges during each time step combined with the irregular domain of active sea ice points and (2) the lack of single-instruction, multiple-data (SIMD) code generation. The standard EVP solver has been refactored based on two generic patterns. The first pattern exposes how general finite differences on masked multi-dimensional arrays can be expressed in order to produce significantly better code generation by changing the memory access pattern from random access to direct access. The second pattern takes an alternative approach to handle static grid properties. The measured single-core performance improvement is more than a factor of 5 compared to the standard implementation. The refactored implementation of strong scales on the Intel® Xeon® Scalable Processors series node until the available bandwidth of the node is used. For the Intel® Xeon® CPU Max series, there is sufficient bandwidth to allow the strong scaling to continue for all the cores on the node, resulting in a single-node improvement factor of 35 over the standard implementation. This study also demonstrates improved performance on GPU processors.

58 GEOSCIENCES

Coupling Noah-Multiparameterization land-surface Model with Energy Research and Forecasting Model

The Energy Research and Forecasting (ERF) model is a high-performance atmospheric model built on the AMReX adaptive mesh refinement (AMR) framework, enabling efficient simulations on heterogeneous computing platforms that combine multicore processors with hardware accelerators. To support land–atmosphere interactions within ERF’s AMR-based environment, a land-surface model must be capable of operating directly on hierarchically refined meshes. In this work, we present a methodology for coupling the Fortran-based Noah-Multiparameterization (Noah-MP) land-surface model with ERF’s C++ codebase. Rather than rewriting Noah-MP, we construct a Fortran–C interoperability layer using CodeScribe, a tool that leverages large language models (LLMs) to automate the generation of interface code. CodeScribe applies structured prompting techniques to generate bindings that support efficient data exchange and function calls between ERF and Noah-MP. The coupling framework also incorporates AMR-aware data handling strategies, allowing NoahMP to operate seamlessly within ERF’s hierarchical mesh structure. This work provides a structured approach for integrating legacy Fortran models into modern C++-based modeling systems using LLM-assisted code generation.

54 ENVIRONMENTAL SCIENCES

An Integrated Framework for Memory-Centric Analysis: From Trace Collection to Co-Design

The memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems—poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing certain behaviors that emerge at larger scales. These limitations reflect a processor-centric design philosophy increasingly misaligned with memory-bound workloads where detailed understanding of memory access patterns, cache hierarchy interactions, and contention is critical for effective optimization. This paper presents an integrated framework for memory-centric analysis that enables effective hardware-software co-design. We describe practical trace collection techniques, including hardware-assisted processor tracing with minimal overhead and portable software-based instrumentation with statistical sampling. We present multi-perspective analysis methods that examine memory behavior from temporal, sequential, spatial, and relational viewpoints, revealing distinct optimization opportunities invisible in aggregate metrics. We detail an architectural modeling framework that uses sampled traces with temporal interpolation and confidence-based filtering to evaluate cache and memory configurations. Evaluation on representative benchmarks demonstrates that this framework achieves practical accuracy (L2 cache errors of 2.64\%, confidence-filtered L3 errors of 9.92\%, bandwidth errors of 7.33\%) while providing substantial speedup (26.8×) over cycle-accurate simulation, enabling rapid design space exploration. We demonstrate how this integrated framework enables systematic identification of both hardware optimizations (memory controller tuning, bank partitioning, NUMA configuration) and software optimizations (data layout restructuring, prefetching strategies, memory-aware scheduling). Through this comprehensive treatment of the memory-centric analysis pipeline—from trace collection through architectural modeling to co-design application—we provide researchers and practitioners with practical techniques for addressing memory bottlenecks in contemporary computing systems.

Gajaria, Dhruv Mayur