Search NASA⌕ Search

SEARCH · Search NASA

Results for “processor”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Streaming Matching and Edge Cover in Practice

Graph algorithms with polynomial space and time requirements often become infeasible for massive graphs with billions of edges or more. State-of-the-art approaches therefore employ approximate serial, parallel, and distributed algorithms to tackle these challenges. However, such approaches require storing the entire graph in memory and thus need access to costly computing resources such as clusters and supercomputers. In this paper, we present practical streaming approaches for solving massive graph problems using limited memory for two prototypical graph problems: maximum weighted matching and minimum weighted edge cover. For matching, we conduct a thorough computational study on two of the semi-streaming algorithms including a recent breakthrough result that achieves a $1/(2+\varepsilon)$-approximation of the weight while using $O( n \log W /\epsilon)$ memory (here $n$ is the number of vertices and $W$ is the maximum edge weight), designed by Paz and Schwartzman [SODA, 2017]. Empirically, we show that the semi-streaming algorithms produce matchings whose weight is close to the best $1/2$-approximate offline algorithm while requiring less time and an order-of-magnitude less memory. For minimum weighted edge cover, we develop three novel semi-streaming algorithms. Two of these algorithms require a single pass through the input graph, require $O(n \log n)$ memory, and provide a 2-approximation guarantee on the objective. We also leverage a relationship between approximate maximum weighted matching and approximate minimum weighted edge cover to develop a two-pass $3/2+\epsilon$-approximate algorithm with the memory requirement of Paz and Schwartzman's semi-streaming matching algorithm. These streaming approaches are compared against the state-of-the-art 3/2-approximate offline algorithm. The semi-streaming matching and the novel edge cover algorithms proposed in this paper can process graphs with several billions of edges in under 30 minutes using 6 GB of memory, which is at least an order of magnitude improvement from the offline (non-streaming) algorithms. For the largest graph, the best alternative offline parallel approximation algorithm (GPA+ROMA) could not finish in three hours even while employing hundreds of processors and 1 TB of memory. We also demonstrate an application of the semi-streaming algorithm by computing a matching using linearly bounded memory on item intersection graphs derived from three machine learning datasets, whereas the existing offline algorithms could not complete on one of these datasets since their memory requirements exceeded 1TB.

Ferdous, S M.↗

Deriving Effective Coupling Strength with Born-Oppenheimer Approximation

The development of high-fidelity quantum gates is paramount to the scalability of quantum processors. Tunable couplers have played a key role in the realization of these low error rate quantum gates in superconducting circuits. However, the derivation of the effective coupling strength between the qubits is usually enabled by a Schrieffer-Wolff transformation, which is complex and requires prior-knowledge of the proper generator for the transformation. We propose a simpler method of obtaining this effective coupling using an approach similar to the Born-Oppenheimer approximation.

Wichmann, Conrad↗

Calibration of the analog beam-signal hardware for the credited engineered beam power limit system at the Proton Power Upgrade Project at the Spallation Neutron Source

A programmable signal processor-based credited safety control that calculates pulsed beam power based on beam kinetic energy and charge was designed as part of the Proton Power Upgrade (PPU) project at the Spallation Neutron Source (SNS). The system must reliably shut off the beam if the average power exceeds 2.145 MW averaging over 60 seconds. System calibration requires pedigree in measurements, calibration setup, and calculations. This paper discusses the calibration of the analog beam signal components up to and including the Analog Digital Convertors (ADCs) for implementation into the Safety Programmable Logic Controllers (PLCs) and Field Programmable Gate Arrays (FPGAs).

Bobrek, Miljko↗

autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm Architectures

This paper presents an open-source library that pushes the limits of performance portability for irregular General Matrix Multiplication (GEMM) on the widely-used Arm architectures. Our library, autoGEMM, is designed to support a wide range of Arm processors: from edge devices to HPC-grade CPUs. autoGEMM generates optimized kernels for various hardware configurations by auto-combining fragments of autogenerated micro-kernels that employ hand-written optimizations to maximize computational efficiency. We optimize the kernel pipeline by tuning the register reuse and the data load/store overlapping. In addition, we use a dynamic tiling scheme to generate balanced tile shapes. Finally, we position autoGEMM on top of the TVM framework where our dynamic tiling scheme prunes the search space for TVM to identify the optimal combination of parameters for code optimization. Evaluations on five different classes of Arm chips demonstrate the advantages of autoGEMM. For small matrices, autoGEMM achieves 98% of peak and up to 2.0x speedup over state-of-the-art libraries such as LIBXSMM and LibShalom. For irregular matrices (i.e. tall skinny and long rectangles), autoGEMM is 1.3-2.0x faster than widely-used libraries such as OpenBLAS and Eigen. autoGEMM is publicly available at: https://github.com/wudu98/autoGEMM.

Wu, Du↗

Window obscuration sensors for mobile gas and chemical imaging cameras

An infrared (IR) imaging system for determining a concentration of a target species in an object is disclosed. The imaging system can include an optical system including a focal plane array (FPA) unit behind an optical window. The optical system can have components defining at least two optical channels thereof, said at least two optical channels being spatially and spectrally different from one another. Each of the at least two optical channels can be positioned to transfer IR radiation incident on the optical system towards the optical FPA. The system can include a processing unit containing a processor that can be configured to acquire multispectral optical data representing said target species from the IR radiation received at the optical FPA. One or more of the optical channels may be used in detecting objects on or near the optical window, to avoid false detections of said target species.

Mallery, Ryan↗

Secure hierarchical processing using a secure ledger

Disclosed is a system and method for processing data using blockchain technology. The system includes a memory having programmable instructions stored thereon that, when executed by a processor, cause the system to: authenticate one or more sensors in anticipation of receiving component data; receive component data, upon successful authentication; store the component data locally or to a cloud-based server and/or calculate a root value for the component data; store or embed the root value with the stored component data; condense the component data and link the condensed component data to the stored component data via the root value. The system further includes instructions to log the condensed data, including the root value, to a ledger, and to identify a tag or transaction id corresponding to the logging event for subsequent retrieval of the condensed data using the tag or transaction id.

Zhao, Wenbing↗

Systems, methods, and devices for failure detection of one or more energy storage devices

An energy storage device management system can include a management portion for charging/discharging an energy storage device and an ultrasound interrogation portion for passing ultrasound energy through the energy storage device during charge/discharge cycles. A memory stores a stream of capture data instances derived from ultrasound energy exiting the energy storage device and baseline ultrasound data instances corresponding with the energy storage device during normal charging/discharging thereof. A processor can compare each capture data instance with the baseline ultrasound data and detect abnormal operating states of the energy storage device. A warning system can issue a notification when abnormal operating states are detected.

Kowalski, Jeffrey A.↗

System and method for high power diode based additive manufacturing

The present disclosure relates to a system for performing an Additive Manufacturing (AM) fabrication process on a powdered material, deposited as a powder bed and forming a substrate. The system makes use of a laser for generating a laser beam, and an optical subsystem. The optical subsystem is configured to receive the laser beam and to generate an optical signal comprised of electromagnetic radiation sufficient to melt or sinter the powdered material. The optical subsystem uses a digitally controlled mask configured to pattern the optical signal as needed to melt select portions of a layer of the powdered material to form a layer of a 3D part. A power supply and at least one processor are also included for generating a plurality of different power density levels selectable based on a specific material composition, absorptivity and diameter of the powder particles, and a known thickness of the powder bed. The powdered material is used to form the 3D part in a sequential layer-by-layer process.

El-Dasher, Bassem S.↗

Status of the Mu2e calorimeter readout electronics

The Mu2e experiment [1] at Fermilab will search for the neutrino-less coherent conversion of a muon into an electron in the field of a nucleus. Mu2e detectors comprise a straw tracker, an electromagnetic calorimeter and a veto for cosmic rays. The calorimeter employs 1348 Cesium Iodide crystals readout by silicon photomultipliers and fast front-end and digitization electronics. The front-end electronics consists of two discrete readout circuits (AMP-HV) for each crystal. These provide the amplification, shaping stage and linear regulation of the SiPM bias voltage and monitoring. The SiPM and front-end control electronics is implemented in a battery of mezzanine boards each equipped with an ARM processor that controls a group of 20 Amp-HV circuits distributing the low voltage and the high-voltage. The electronic is hosted in crates located on the external surface of calorimeter disks. The crates also host the waveform digitizer board (DIRAC) that performs digitization of the front end signals and transmit the digitized data to the Mu2e DAQ. Calorimeter electronic is hosted inside the cryostat and must sustain very high radiation and magnetic field so it was necessary to fully qualify it. The system design and quality assurance procedures will be reviewed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Field Programmable Gate Array Data Capture for Control Systems

Some Industrial Control Systems (ICS) networks are based on protocols such as Serial and Industrial Ethernet. These protocols currently have no existing cybersecurity monitoring tools, leaving a large gap in the cyber defense of critical infrastructure. In order to analyze such ICS traffic, it is first necessary to implement methods of capturing the ICS data. Whereas traditional methods of analyzing data would use microprocessors, the nature of high-speed analog data can be difficult to implement on such a versatile processor, as they are rather inefficient for doing a single task. Whereas Field Programmable Gate Arrays (FPGAs) provide an adequate tool in analyzing high speed data, as despite the lack of program versatility, Programmable Logic can implement a solution with minimal clock cycles, allowing time for each new packet of data to be captured before a new data sample is taken.

42 ENGINEERING↗

Smartphone application for visualizing building air leakage

Annually, unwanted air leakage through building envelopes accounts for 4 quads of energy consumption in the United States, which translates to about 10% of total building energy consumption. Locating and sealing leakage sites is crucial fordecreasing building energy consumption. Smartphones are ubiquitous and contain sophisticated cameras and highperformance processors that could be employed to visualize air leakage using the background-oriented schlieren imaging technique, making leak detection cheaper and easier. This technique requires a textured and high-contrast background such as a brick or concrete masonry unit wall, a building air leak that has a temperature difference compared with the ambient air, and an imaging system. This work focuses on using smartphones as the imaging system to visualize air leakages. The paper discusses application development and the results of testing to determine leakage visualization performance as a function of leak temperature. The results show that leaks with a temperature difference greater than 16°C compared with the ambient air temperature were visualized using existing smartphones.

Boudreaux, Philip [ORNL] (ORCID:0000000229564665)↗

Risk-informed Graded Approach for Reliability and Performance Assessment of Sensor and Instrumentation Systems within Advanced Condition Monitoring Technologies

Advanced condition monitoring (ACM) technologies, such as digital twins, are innovative strategies designed to provide real-time health insights, including the remaining useful life of components. The primary goal of ACM is to predict and alert operators to potential functional failures before they occur. ACM systems achieve this by integrating predictive models with various sensor instrumentation, analog-to-digital converters, data warehouses, and data pre-processors. These sensor and instrumentation systems (SIS) are essential for forming a comprehensive understanding of component conditions and ensuring the predictive success of ACM programs. Introducing new technologies like ACM involves varying degrees of risk that can impact plant reliability. Therefore, risk mitigation should be commensurate with the performance and reliability of the developed technology, following a risk-informed graded approach (RIGA). Establishing a RIGA process requires a clear understanding of the hazards and reliability of all subsystems, including their interdependencies and potential impacts on the overall system. Given the critical role of SIS in ACM, this work reviews hazard identification and reliability quantification methods for SIS. It also considers these methods' implications when developing a RIGA process for ACM.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Mechanical and biochemical recovery of landfill waste in an underserved community

Historically in the United States, waste collected for recycling has been sold and shipped to processors in China. In 2013 and 2018, China introduced the Green Fence and National Sword policies which restricts the import of contaminated materials and banned the import of many recyclables. The cost of recycling in the United States has increased following these policy changes, which has led to many communities reducing their recycling programs or halting them altogether. Rural and underserved communities that don’t have resources to afford sophisticated recycling programs have been heavily impacted. Previous work at INL demonstrated that MSW is a potentially viable feedstock for both biochemical and thermochemical conversion. The goal of this project is to assess preprocessing tools that can produce consistent feedstocks that meet conversion specifications, remove problematic contaminants, and reduce the amount of waste that is landfilled. Municipal solid waste was collected from an underserved community in southeast Idaho, contaminants were characterized, and mechanically separated into two discrete fractions. The unit operations identified during mechanical separation trials will be mobilized to on-site with a goal of 50% recovery of paper and plastic waste.

09 - BIOMASS FUELS↗

Distributed Quantum-Enhanced Optimization: A Topographical Preconditioning Approach for High-Dimensional Search

Optimization problems become fundamentally challenging as the number of variables increases. Because the volume of the search space grows exponentially, classical algorithms frequently fail to locate the global minimum of non-convex functions. While quantum optimization offers a potential alternative, mapping continuous problems onto near-term quantum hardware introduces severe scaling limits and barren plateaus. To bridge this gap, we propose the Distributed Quantum-Enhanced Optimization (D-QEO) framework. Instead of forcing the quantum processor to find the exact minimum, we use it simply as a topographical preconditioner. The QPU maps the landscape to locate the most promising basin of attraction, generating high-quality seed points for a classical GPU-accelerated solver to refine. To make this approach viable for utility-scale problems, we exploit the mathematical structure of separable functions. This allows us to cut a 50-qubit (i.e., $2^{50}$) global search space into independent and manageable sub-spaces using 5-qubit subcircuits. By executing these fragments concurrently with CUDA-Q, we completely bypass the overhead of cross-register entanglement and classical tensor knitting for separable functions. Benchmarks on the 10-dimensional Rastrigin and Ackley functions show that D-QEO prevents the exponential failure rates observed in purely classical algorithms. Furthermore, this quantum warm-start significantly reduces the number of classical BFGS iterations required to converge, providing a highly practical blueprint for utilizing near-term quantum resources in complex global search.

Soos, Dominik [Old Dominion U.]↗

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Exact and Fixed-Point Grover Search with Qudits

Grover's algorithm provides a quadratic speedup for searching unstructured databases and is traditionally implemented with qubits in Hilbert spaces whose dimensions are powers of two. With the advent of quantum platforms utilizing qudits---quantum systems with more than two levels---there is a need to generalize Grover search to these architectures, including heterogeneous systems with qudits of varying dimensions. Here, we present a unified framework for qudit-based Grover search, detailing the construction of oracles and diffusion operators with and without ancilla qubits and generalizing deterministic and fixed-point search variants that ensure exact or bounded success probabilities. We analyze phase-matching techniques and provide explicit circuit decompositions suitable for diverse hardware platforms. We also compare the corresponding trajectories on the Bloch sphere to provide an intuitive visualization of how the different phase choices amplify the target state. These results facilitate flexible, hardware-oriented protocols for implementing Grover search on qudit processors, potentially reducing circuit depth and enhancing success probabilities, thereby offering a practical toolkit for quantum computation and sensing applications leveraging multilevel quantum systems.

Roy, Tanay [Fermilab] (ORCID:000000019442862X)↗