Search NASA⌕ Search

SEARCH · Search NASA

Results for “fpga”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Complexity Agnostic Clustering Engine for Time Projection Chambers and its Implementation in FPGA

A clustering functional block implemented in field-programable-gate-array (FPGA) for time projection chambers (TPC) operating with predictable time regardless the complexity of the event is described in this paper. The clustering functional block reorganizes input data and the hits data belonging to the same clusters are output together for further process in the later stages. The clustering operation consists of two phases, data filling phase and data outputting phase, and the later uses the same number of clock cycles as the data filling phase. The clustering block can accommodate events with arbitrary number of clusters and number of hits per cluster as long as the total number of hits is within a predesigned limit. The operation time is exactly twice of the data filling time with no residual O(n2) term. The clustering block has been implemented with operating frequency of 200 MHz in a low-cost FPGA evaluation module and test results confirm the expected performance.

Wu, Jinyuan [Fermilab] (ORCID:0000000344329521)↗

Entropy Analysis of FPGA Interconnect and Switch Matrices for Physical Unclonable Functions

Random variations in microelectronic circuit structures represent the source of entropy for physical unclonable functions (PUFs). In this paper, we investigate delay variations that occur through the routing network and switch matrices of a field-programmable gate array (FPGA). The delay variations are isolated from other components of the programmable logic, e.g., look-up tables (LUTs), flip-flops (FFs), etc., using a feature of Xilinx FPGAs called dynamic partial reconfiguration (DPR). A set of partial designs is created to fix the placement of a time-to-digital converter (TDC) and supporting infrastructure to enable the path delays through the target interconnect and switch matrices to be extracted by subtracting out common-mode delay components. Delay variations are analyzed in the different levels of routing resources available within FPGAs, i.e., local routing and across-chip routing. Data are collected from a set of Xilinx Zynq 7010 devices, and a statistical analysis of within-die variations in delay through a set of the randomly-generated and hand-crafted interconnects is presented.

97 MATHEMATICS AND COMPUTING↗

Embedded FPGA developments in 130 nm and 28 nm CMOS for machine learning in particle detector readout

Embedded field programmable gate array (eFPGA) technology allows the implementation of reconfigurable logic within the design of an application-specific integrated circuit (ASIC). This approach offers the low power and efficiency of an ASIC along with the ease of FPGA configuration, particularly beneficial for the use case of machine learning in the data pipeline of next-generation collider experiments. An open-source framework called "FABulous" was used to design eFPGAs using 130 nm and 28 nm CMOS technology nodes, which were subsequently fabricated and verified through testing. The capability of an eFPGA to act as a front-end readout chip was assessed using simulation of high energy particles passing through a silicon pixel sensor. A machine learning-based classifier, designed for reduction of sensor data at the source, was synthesized and configured onto the eFPGA. A successful proof-of-concept was demonstrated through reproduction of the expected algorithm result on the eFPGA with perfect accuracy. Finally, further development of the eFPGA technology and its application to collider detector readout is discussed.

47 OTHER INSTRUMENTATION↗

ESnet-JLab FPGA Accelerated Transport (control plane) [EJFAT (udplbd2)] v2.0

The ESnet-JLab FPGA Accelerated Transport system is a solution for streaming high-speed scientific measurement data from Data Acquisition Systems (DAQs) to high-performance computing facilties. It is generally compatible with many science workflows, and makes no assumptions about the specifics of any particular experiment. This program (udplbd version 2) implements the control plane for the system. It is responsible for programming network forwarding rules into the data plane (implemented by the hardware designed named udplb, described separately). It also implements the control loop necessary to match up offered workload with available capacity on high-performance compute nodes.

Howard, Derek [Lawrence Berkeley National Laborato↗

Extinction Monitoring of Pulsed Proton Beams Using FPGA-Based Peak Detection

The Mu2e experiment at Fermilab imposes stringent requirements on the elimination of out-of-time beam in its pulsed proton beam - a requirement known as "extinction". We present a method to measure the out-of-time particle rates to calculate the level of extinction in the inter-pulse gaps. The proposed method utilizes an array of quartz Cherenkov radiators and photomultiplier tubes to detect particles scattered from a vacuum chamber in the M4 transfer beamline at Fermilab.The measurement will employ a new μTCA-based FPGA system for data acquisition and signal processing, utilizing real-time peak detection algorithms to count scattered beam particles. By integrating data over many transfers, the time profile of the out-of-time beam will be resolved to fractional levels relative to that of the in-time beam. These results are compared with G4beamline simulations to validate models of beam transport, dynamics, and extinction, providing critical input for optimizing beam delivery to Mu2e.

Hensley, Ryan [UC, Davis]↗

Calculating beam extinction in a pulsed proton beam using FPGA-based peak detection

The Mu2e experiment at Fermilab imposes stringent requirements on the elimination of out-of-time beam in its pulsed proton beam, a requirement known as “extinction”. Utilizing a new μTCA-based FPGA data acquisition system, we recorded live particle data from scattered particles incident on an array of quartz Cherenkov radiators and photomultiplier tubes to measure the extinction in the inter-pulse gaps in the pulsed proton beam. Minuscule errors in the derived signal period can make a measurement of the extinction impossible, so after taking a Fourier transform, further optimizations on the period were done based on the assumption that the signal period is stable over the full time of the beam spill while it is being resonantly extracted. After these optimizations, the beam extinction was shown to be on the level of 10^3.

Hensley, Ryan [UC, Davis]↗

FPGA-Based Spill Regulation System for the Muon Delivery Ring at Fermilab

The Muon to Electron Experiment (Mu2e) requires a uniform beam profile from the Muon Delivery Ring to meet their experimental needs. A specialized Spill Regulation System (SRS) has been developed to help achieve consistent spill uniformity. The system is based on a custom-designed carrier board featuring an Arria 10 SoC, capable of executing real-time feedback control. The FPGA processes beam pulses of approximately 200 ns every 1.695 $μ$s, allowing for continuous monitoring of the extracted spill intensity through fast bunch integration. The system directly controls three quadrupole magnets, which work in conjunction with sextupole magnets to achieve third-order resonant extraction. Furthermore, the board interfaces with Fermilab's Accelerator Control Network (ACNET), enabling operators to modify spill regulation settings in real-time via the control network while providing diagnostic waveforms. These waveforms help operators monitor the process and fine-tune the feedback mechanisms. This paper presents an overview of the board's architecture and its initial progress toward regulating beam extraction. This initial version of the regulation system aims to evaluate baseline performance to inform future system improvements.

Berlioz, J. R. [Fermilab]↗

Lightweight Embedded Controller in Advanced FPGA SoC for Radar Signal Processing [Poster]

The objective of the project is to demonstrate that critical control functions can be implemented using little resources in modern microelectronics. A finite state machine (Figure 1) is implemented onto a field programmable gate array (FPGA). The functionality of the system is demonstrated by sending binary instructions to the controller. The controller transmits patterns through an LED, controls an electromechanical device, and uses pulse-width modulation (PWM) for radar functions.

42 ENGINEERING↗

An Arbitrary Time Interval Generator Base on Vernier Clocks with 0.67 ps Adjustable Steps Implemented in FPGA

In TDC testing or timing system implementation tasks, it is often desirable to generate signal pulses with fine adjustable time intervals. In delay cell-based schemes, the time adjustment steps are limited by the propagation delays of the cells, which are typically 15 to 20 picoseconds per step and are sensitive to temperature and operating voltage. In this document, a purely digital scheme based on two vernier clocks with small frequency difference generated using cascaded PLL is reported. The scheme is tested in two families of low-cost FPGA and 0.67 and 0.97 picoseconds adjustable steps of the time intervals are achieved.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

FPGA-Based Spill Regulation System for the Muon Delivery Ring at Fermilab

The Muon to Electron Experiment (\acrshort{Mu2e}) requires a uniform beam profile from the Muon Delivery Ring to meet their experimental needs. A specialized Spill Regulation System (\acrshort{SRS}) has been developed to help achieve consistent spill uniformity. The system is based on a custom-designed carrier board featuring an Arria 10 SoC, capable of executing real-time feedback control. The FPGA processes beam pulses of approximately 200 ns every 1.695 microseconds, allowing for continuous monitoring of the extracted spill intensity through fast bunch integration. The system directly controls three quadrupole magnets, which work in conjunction with sextupole magnets to achieve third-order resonant extraction. Furthermore, the board interfaces with Fermilab s Accelerator Control Network (ACNET), enabling operators to modify spill regulation settings in real-time via the control network while providing diagnostic waveforms. These waveforms help operators monitor the process and fine-tune the feedback mechanisms. This paper presents an overview of the board's architecture and its initial progress toward regulating beam extraction. This initial version of the regulation system aims to evaluate baseline performance to inform future system improvements.

Berlioz, Jose Rene [Fermilab]↗

The Tiny Median Filter: A Small Size, Flexible Arbitrary Percentile Finder Scheme Suitable for FPGA Implementation

This document reports the design, implementation and testing of a small silicon resource usage, very flexible arbitrary percentile finding scheme called the Tiny Median Filter. It can be used not only as a median filter in image processing with square filtering windows, but also for applications of any percentile filter or maximum or minimum finder with any size of data set as long as the number of bits of the data is finite. It opens possibilities for image processing tasks with non-square or irregular filter windows. In this scheme, data swapping or data bit manipulating are avoided and high functional efficiency of the logic components is applied to save silicon resources. Some logic functions are absorbed into other functions to further reduce the complexity. The combinational logic paths are designed to be sufficiently short so that the firmware can be compiled to the maximum operating frequency allowed by the block memories of the FPGA devices. The Tiny Median Filter receives, processes and output data in non-stop manner with no irregular timing which helps to simplify design of surrounding stages.

Wu, Jinyuan [Fermilab] (ORCID:0000000344329521)↗

An Arbitrary Time Interval Generator Base on Vernier Clocks with 0.67 ps Adjustable Steps Implemented in FPGA

In TDC testing and timing system implementations, it is often necessary to generate signal pulses with finely adjustable time intervals. In delay cell–based schemes, the adjustment resolution is constrained by the propagation delay of the cells—typically 15–20 ps per step—and is sensitive to temperature and supply voltage variations. This document presents a fully digital approach that uses two vernier clocks, generated by two stages of cascaded phase-locked loops (PLLs) with a slight frequency difference, to achieve adjustable timing intervals through accumulated phase differences. The scheme was validated on two families of low-cost FPGA devices, achieving adjustable step sizes of 0.67 ps and 0.97 ps.

Wu, Jin-yuan [Fermilab] (ORCID:0000000344329521)↗

White-Rabbit-Disciplined FPGA Readout for Fermilab Timing Events

Fermilab's accelerator timing links broadcast short event codes to thousands of devices at once, but the links themselves carry no absolute notion of time; that comes separately from a White Rabbit reference. This work builds the piece that ties the two together on a single board. On a Xilinx Kria KR260 (Zynq UltraScale+), the programmable logic decodes a real Fermilab TCLK link, stamps every event with an absolute White-Rabbit \{sec, ns\} UTC time, and reads the timestamped stream out over AXI4-Lite; a thin Linux process on the same die publishes each event into a Redis stream on the control network. To exercise the full chain on one board, the decoded events are re-encoded as gigabit ACLK, transmitted out an SFP+ optical port, looped back over a short fiber jumper, and decoded again on the same timeline, and are additionally mirrored as an ACLK-Lite Manchester waveform for benchtop probing. Across sustained, multi-day testing against real Fermilab TCLK, the pipeline has decoded, timestamped, and published hundreds of millions of events with practically zero loss, and folding the timestamped stream on the 60-second accelerator supercycle recovers the machine's periodic structure directly from the published data.

Rossel, Jacob [Fermilab; UC, Berkeley (main)] (ORC↗

Neuro-Spark: A Submicrosecond Spiking Neural Networks Architecture for In-Sensor Filtering

Neuro-Spark, which is a new neuromorphic architecture with a field-programmable gate array (FPGA) implementation for ultrafast spiking neural network (SNN) inference at the edge, facilitates smart-pixel in-sensor filtering for high-energy physics experiments at the Large Hadron Collider (LHC). Utilizing the evolutionary optimization for neuromorphic systems (EONS) training method, we generate compact SNN models with 91% signal efficiency, akin to convolutional neural networks but with half the parameters. However, deploying near the detector poses a challenge because the SNN must handle a sustained input data rate exceeding 1013 GB/s. To overcome this, we propose a novel hardware architecture that uses high-level synthesis to construct a tuned architecture for the EONS-trained SNN. In addition to the analysis and validation with an AMD Xilinx Artix-A7 FPGA, our solution consumes only ç24% of FPGA LUT and flipflops. We also introduce an innovative quantization method that reduces FPGA resource utilization by ç15% without compromising accuracy. Our FPGA implementation achieves computing latency of ç10 ns for smart-pixel application inference on an edge FPGA.

Miniskar, Narasinga Rao↗

Hls4ml Synthesis Testing

HLS4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where HLS4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗