Search NASA⌕ Search

SEARCH · Search NASA

Results for “PROGRAMMER”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Investigating resource-efficient neutron/gamma classification ML models targeting eFPGAs

There has been considerable interest and resulting progress in implementing machine learning (ML) models in hardware over the last several years from the particle and nuclear physics communities. A big driver has been the release of the Python package, hls4ml, which has enabled porting models specified and trained using Python ML libraries to register transfer level (RTL) code. So far, the primary end targets have been commercial field-programmable gate arrays (FPGAs) or synthesized custom blocks on application specific integrated circuits (ASICs). However, recent developments in open-source embedded FPGA (eFPGA) frameworks now provide an alternate, more flexible pathway for implementing ML models in hardware. These customized eFPGA fabrics can be integrated as part of an overall chip design. In general, the decision between a fully custom, eFPGA, or commercial FPGA ML implementation will depend on the details of the end-use application. In this work, we explored the parameter space for eFPGA implementations of fully-connected neural network (fcNN) and boosted decision tree (BDT) models using the task of neutron/gamma classification with a specific focus on resource efficiency. We used data collected using an AmBe sealed source incident on Stilbene, which was optically coupled to an OnSemi J-series silicon photomultiplier (SiPM) to generate training and test data for this study. We investigated relevant input features and the effects of bit-resolution and sampling rate as well as trade-offs in hyperparameters for both ML architectures while tracking total resource usage. The performance metric used to track model performance was the calculated neutron efficiency at a gamma leakage of 10 -3 . The results of the study will be used to aid the specification of an eFPGA fabric, which will be integrated as part of a test chip.

47 OTHER INSTRUMENTATION↗

Embedded FPGA developments in 130 nm and 28 nm CMOS for machine learning in particle detector readout

Embedded field programmable gate array (eFPGA) technology allows the implementation of reconfigurable logic within the design of an application-specific integrated circuit (ASIC). This approach offers the low power and efficiency of an ASIC along with the ease of FPGA configuration, particularly beneficial for the use case of machine learning in the data pipeline of next-generation collider experiments. An open-source framework called "FABulous" was used to design eFPGAs using 130 nm and 28 nm CMOS technology nodes, which were subsequently fabricated and verified through testing. The capability of an eFPGA to act as a front-end readout chip was assessed using simulation of high energy particles passing through a silicon pixel sensor. A machine learning-based classifier, designed for reduction of sensor data at the source, was synthesized and configured onto the eFPGA. A successful proof-of-concept was demonstrated through reproduction of the expected algorithm result on the eFPGA with perfect accuracy. Finally, further development of the eFPGA technology and its application to collider detector readout is discussed.

47 OTHER INSTRUMENTATION↗

Sensor response and radiation damage effects for 3D pixels in the ATLAS IBL Detector

Pixel sensors in 3D technology equip the outer ends of the staves of the Insertable B Layer (IBL), the innermost layer of the ATLAS Pixel Detector, which was installed before the start of LHC Run 2 in 2015. 3D pixel sensors are expected to exhibit more tolerance to radiation damage and are the technology of choice for the innermost layer in the ATLAS tracker upgrade for the HL-LHC programme. While the LHC has delivered an integrated luminosity of ≃ 235 fb -1 since the start of Run 2, the 3D sensors have received a non-ionising energy deposition corresponding to a fluence of ≃ 8.5 × 10 14 1 MeV neutron-equivalent cm -2 averaged over the sensor area. This paper presents results of measurements of the 3D pixel sensors' response during Run 2 and the first two years of Run 3, with predictions of its evolution until the end of Run 3 in 2025. Data are compared with radiation damage simulations, based on detailed maps of the electric field in the Si substrate, at various fluence levels and bias voltage values. These results illustrate the potential of 3D technology for pixel applications in high-radiation

47 OTHER INSTRUMENTATION↗

An integer- N frequency synthesizer for flexible on-chip clock generation

A low-power integer-N frequency synthesizer for flexible on-chip clock generation has been designed in a 65 nm CMOS process. The circuit can be programmed to generate two independent low-jitter clocks between 30 MHz and 3 GHz that are locked to a 10–50 MHz reference input. The design uses a phase-locked loop (PLL) with a dual-tuned LC voltage-controlled oscillator (VCO), programmable feedback divider, and dual output dividers. The total power consumption from 1.2 V and 0.8 V supplies is 4.0 mW. In conclusion, experimental results confirm the functionality of the proposed synthesizer over a wide range of output frequencies.

47 OTHER INSTRUMENTATION↗

Silicon wafer fracture stress for tracking sensors in particle physics experiments

For the construction of the ATLAS Inner Tracker strip detector, silicon strip sensor modules are glued directly onto carbon fibre support structures using a soft silicone gel. During tests at temperatures below -35°C, several of the sensors were found to crack due to a mismatch in coefficients of thermal expansion between polyimide circuit boards with copper metal layers (glued onto the sensor) and the silicon sensor itself. While module assembly procedures were developed to minimise variations between modules, cold tests showed a wide range of temperatures at which supposedly comparable modules failed. The observed variance (fracture temperatures between -35°C and -70°C) for supposedly comparable modules suggests an undetected variation between modules suspected to be intrinsic to the silicon wafer itself. Therefore, a test programme was developed to investigate the fracture stress of representative sensor wafer cutoffs. This paper presents results for the fracture stress of silicon sensors used in detector modules.

Detector design and construction technologies and ↗

High-bandwidth frequency domain multiplexed readout of transition-edge sensors for neutrinoless double beta decay searches

The next-generation of cryogenic neutrinoless double-beta decay experiments require increasingly fast readout in order to improve background discrimination. These experiments, operated as cryogenic calorimeters at ∼ 10 mK, are usually read out by high-impedance neutron transmutation doped (NTD) thermistors, which provide good energy resolution, but are limited by ∼ 1 ms response times. Superconducting detectors, such as transition-edge sensors (TESs) with a time resolution of ∼ 100 μs, offer superior timing performance over NTD semiconductor bolometers. To make this technology viable for an application to a thousand or more channels, multiplexed readout is necessary in order to minimize the thermal load and radioactive contamination induced by the readout. Frequency-domain multiplexing readout (fMUX) for TESs, previously developed at Berkeley Lab and McGill University, is currently in use for mm-wave telescopes with detector sampling rates in the order of 100 Hz. We demonstrate a new readout system, based on the McGill/Berkeley digital fMux readout, to satisfy the higher bandwidth and noise requirements of the next generation of TES-instrumented cryogenic calorimeters. Each multiplexing readout module comprises 10 superconducting resonators in the 1–5 MHz range and a DC superconducting quantum interference device (DC-SQUID), interfaced to high-speed field programmable gate array (FPGA)-based electronics for digital signal processing and low-latency SQUID feedback. The new readout samples detectors at 156 kHz, three orders of magnitude faster than its cosmology-oriented predecessor, and demonstrates a stable feedback bandwidth of 3 kHz in a real TES-based system.

47 OTHER INSTRUMENTATION↗

Heat load measurements for the PIP-II pHB650 cryomodule

Phase-3 testing of the pHB650 cryomodule at the PIP-II Injector Test Facility was conducted to evaluate the effectiveness of heat load mitigations performed after earlier phases of testing and to continue pinpointing any sources of unexpectedly high heat loads.. The programme measured HTTS, LTTS, and 2 K isothermal/non-isothermal loads under "standard", "linac", and "simulated dynamic" operating modes, recording data both inside the cryomodule and across the bayonet can circuits. Thermal-acoustic oscillations were eliminated by replacing the original G10 cooldown-valve stem with a stainless-steel stem fitted with wipers. A newly developed Python script automated acquisition of ACNET data, performed real-time heat-load calculations, and generated plots and tables that were posted to the electronic logbook within minutes, vastly reducing manual effort and accelerating feedback between SRF and cryogenics teams. Analysis showed that JT heat-exchanger effectiveness and temperature stratification in the two-phase and relief piping strongly influence the observed loads and helped isolate sources of excess heat. The campaign demonstrates that rigorous pre-test planning, real-time diagnostics, and automated reporting can improve both accuracy and efficiency, providing a template for future PIP-II cryomodule tests and for implementing targeted heat-load mitigations.

Porwisiak, D. [Fermilab; Wroclaw Tech. U.] (ORCID:↗

Distilling particle knowledge for fast reconstruction at high-energy physics experiments

Knowledge distillation is a form of model compression that allows artificial neural networks of different sizes to learn from one another. Its main application is the compactification of large deep neural networks to free up computational resources, in particular on edge devices. In this article, we consider proton-proton collisions at the High-Luminosity Large Hadron Collider (HL-LHC) and demonstrate a successful knowledge transfer from an event-level graph neural network (GNN) to a particle-level small deep neural network (DNN). Our algorithm, DistillNet, is a DNN that is trained to learn about the provenance of particles, as provided by the soft labels that are the GNN outputs, to predict whether or not a particle originates from the primary interaction vertex. The results indicate that for this problem, which is one of the main challenges at the HL-LHC, there is minimal loss during the transfer of knowledge to the small student network, while improving significantly the computational resource needs compared to the teacher. This is demonstrated for the distilled student network on a CPU, as well as for a quantized and pruned student network deployed on a field programmable gate array. Our study proves that knowledge transfer between networks of different complexity can be used for fast artificial intelligence (AI) in high-energy physics that improves the expressiveness of observables over non-AI-based reconstruction algorithms. Such an approach can become essential at the HL-LHC experiments, e.g. to comply with the resource budget of their trigger stages.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Accelerating data acquisition with FPGA-based edge machine learning: a case study with LCLS-II

New scientific experiments and instruments generate vast amounts of data that need to be transferred for storage or further processing, often overwhelming traditional systems. Edge machine learning (EdgeML) addresses this challenge by integrating machine learning (ML) algorithms with edge computing, enabling real-time data processing directly at the point of data generation. EdgeML is particularly beneficial for environments where immediate decisions are required, or where bandwidth and storage are limited. In this paper, we demonstrate a high-speed configurable ML model in a fully customizable EdgeML system using a field programmable gate array (FPGA). Our demonstration focuses on an angular array of electron spectrometers, referred to as the ‘CookieBox,’ developed for the Linac Coherent Light Source II project. The EdgeML system captures 51.2 Gbps from a 6.4 GS s −1 analog to digital converter and is designed to integrate data pre-processing and ML inside an FPGA. Our implementation achieves an inference latency of 0.2 µs for the ML model, and a total latency of 0.4 µs for the complete EdgeML system, which includes pre-processing, data transmission, digitization, and ML inference. The modular design of the system allows it to be adapted for other instrumentation applications requiring low-latency data processing.

97 MATHEMATICS AND COMPUTING↗

SymbolNet: neural symbolic regression with adaptive dynamic pruning for compression

Abstract Compact symbolic expressions have been shown to be more efficient than neural network (NN) models in terms of resource consumption and inference speed when implemented on custom hardware such as field-programmable gate arrays (FPGAs), while maintaining comparable accuracy (Tsoi et al 2024 EPJ Web Conf. 295 09036). These capabilities are highly valuable in environments with stringent computational resource constraints, such as high-energy physics experiments at the CERN Large Hadron Collider. However, finding compact expressions for high-dimensional datasets remains challenging due to the inherent limitations of genetic programming (GP), the search algorithm of most symbolic regression (SR) methods. Contrary to GP, the NN approach to SR offers scalability to high-dimensional inputs and leverages gradient methods for faster equation searching. Common ways of constraining expression complexity often involve multistage pruning with fine-tuning, which can result in significant performance loss. In this work, we propose S y m b o l N e t , a NN approach to SR specifically designed as a model compression technique, aimed at enabling low-latency inference for high-dimensional inputs on custom hardware such as FPGAs. This framework allows dynamic pruning of model weights, input features, and mathematical operators in a single training process, where both training loss and expression complexity are optimized simultaneously. We introduce a sparsity regularization term for each pruning type, which can adaptively adjust its strength, leading to convergence at a target sparsity ratio. Unlike most existing SR methods that struggle with datasets containing more than O ( 10 ) inputs, we demonstrate the effectiveness of our model on the LHC jet tagging task (16 inputs), MNIST (784 inputs), and SVHN (3072 inputs).

Tsoi, Ho Fung (ORCID:0000000225502184)↗

Geometric GNNs for charged particle tracking at GlueX

Nuclear physics experiments are aimed at uncovering the fundamental building blocks of matter. The experiments involve high-energy collisions that produce complex events with many particle trajectories. Tracking charged particles resulting from collisions in the presence of a strong magnetic field is critical to enable the reconstruction of particle trajectories and precise determination of interactions. It is traditionally achieved through combinatorial approaches that scale worse than linearly as the number of hits grows. Since particle hit data naturally form a point cloud and can be structured as graphs, graph neural networks (GNNs) emerge as an intuitive and effective choice for this task. In this study, we evaluate the GNN model for track finding on the data from the GlueX experiment at Jefferson Lab. We use simulation data to train the model and test on both simulation and real GlueX measurements. We demonstrate that GNN-based track finding outperforms the currently used traditional method at GlueX in terms of segment-based efficiency at a fixed purity while providing faster inferences. We show that the GNN model can achieve significant speedup by processing multiple events in batches, which exploits the parallel computation capability of graphical processing units (GPUs). Finally, we compare the GNN implementation on GPU and field-programmable gate array and describe the trade-off.

batched GNN pipeline↗

Leveraging dendritic complexity for neuromorphic computing

Abstract Beyond-von Neumann computing approaches are necessary to sustain the growth of microelectronics and the increasing appetite for artificial intelligence/machine learning algorithms. Neuromorphic computing is an emerging paradigm that takes inspiration from the brain to provide a path forward to improve the computational efficiency and computational density of next-generation computing architectures. In nature, we observe brains performing complex computations with a much smaller energy footprint than conventional computing approaches. Current neuromorphic systems are focused primarily on scalability, namely, increasing the number of computational units (neurons) and connections between units (synapses). However, for brain-like cognition and efficiency in next-generation computing hardware, we need increased complexity in function, as well as improved connection density for scalability. Here, we present our work that aims to incorporate dendrites for ‘compute-on-wire’ in neuromorphic architectures to increase the computational complexity (e.g. number of programmable parameters, nonlinear dynamics) as well as computational efficiency (energy/compute) of artificial neural networks (ANNs). We do this by showcasing neuromorphic dendrite elements that can be leveraged for various applications. We will present examples of neuroscience-inspired direction-selective circuits and an ANN with active dendrites leveraging shunting inhibition. We also demonstrate the benefits of using dendrites in deep neural networks. To conclude, we discuss how we can utilize emerging hardware devices in these systems and design next-generation neuromorphic architectures with dendrites.

Cardwell, Suma G. (ORCID:0000000226575545)↗

Developmentally-specific physiological and metabolic responses support drought resilience in switchgrass and constrains biofuel yield

Switchgrass (Panicum virgatum) is a promising bioenergy crop due in part to its resilience to drought stress. However, the significance of drought timing remains poorly understood, both from a plant biology perspective and its impact on downstream biofuel production. This study determines the developmental stage-specific physiological and metabolic responses of switchgrass to drought stress and its implications for biofuel production using a custom-built programmable irrigation system. Vegetative, flowering, and senescence-stage drought significantly reduced carbon dioxide assimilation, and stomatal conductance without affecting biomass yield. Metabolic profiling revealed significant accumulation of glucose, fructose, quinic acid, shikimate and GABA during vegetative-stage drought, while flowering and senescence stages exhibited limited metabolic changes. Similarly, specialized metabolites also displayed distinct developmental patterns, with vegetative-stage drought driving the most pronounced metabolic alterations. Thermochemically-treated and hydrolyzed switchgrass biomass from vegetative-stage drought showed elevated lignocellulose-derived compounds and saponins with the latter most positively correlating with fermentation lag times. Conversely, senescence-stage drought enhanced ethanol yields while lowering saponin levels in the hydrolysates. While vegetative-stage drought enhanced physiological resilience, it compromises downstream biofuel production by introducing fermentation inhibitors, particularly saponins.

biofuel↗

Discovery of a z ∼ 0.8 ultra steep spectrum radio halo in the MeerKAT-South Pole Telescope Survey

Radio haloes are diffuse synchrotron sources that trace the turbulent intracluster medium (ICM) of galaxy clusters. However, their origin remains unknown. Two main formation models have been proposed: the hadronic model, in which relativistic electrons are continuously injected by cosmic-ray protons; and the leptonic turbulent re-acceleration model, where cluster mergers re-energize electrons in situ. A key discriminant between the two models would be the existence of ultra-steep spectrum radio haloes (USSRHs), which can only be produced through turbulent re-acceleration. Here, we report the discovery of an USSRH in the galaxy cluster SPT-CLJ2337–5942 at redshift $z = 0.78$ in the MeerKAT-South Pole Telescope 100 deg$^2$ UHF (0.58–1.09 GHz) survey. This discovery is noteworthy for two primary reasons: it is the highest redshift USSRH system to date; and the close correspondence of the radio emission with the thermal ICM as traced by Chandra X-ray observations, further supporting the leptonic re-acceleration model. The halo is underluminous for its mass, consistent with a minor merger origin, which produces steep-spectrum, lower luminosity haloes. This result demonstrates the power of wide-field, high-fidelity, low-frequency ($\lesssim 1$ GHz) surveys like the MeerKAT-SPT 100 deg$^2$ programme to probe the origin and evolution of radio haloes over cosmic time, ahead of the Square Kilometre Array.

X-rays: galaxies↗

Floquet Engineering of Interactions and Entanglement in Periodically Driven Rydberg Chains

Neutral atom arrays driven into Rydberg states constitute a promising approach for realizing programmable quantum systems. Enabled by strong interactions associated with Rydberg blockade, they allow for simulation of complex spin models and quantum dynamics. We introduce a new Floquet engineering technique for systems in the blockade regime that provides control over novel forms of interactions and entanglement dynamics in such systems. Our approach is based on time-dependent control of Rydberg laser detuning and leverages perturbations around periodic many-body trajectories as resources for operator spreading. These time-evolved operators are utilized as a basis for engineering interactions in the effective Hamiltonian describing the stroboscopic evolution. As an example, we show how our method can be used to engineer strong spin exchange, consistent with the blockade, in a one-dimensional chain, enabling the exploration of gapless Luttinger liquid phases. In addition, we demonstrate that combining gapless excitations with Rydberg blockade can lead to dynamic generation of large-scale multipartite entanglement. Experimental feasibility and possible generalizations are discussed.

Floquet systems↗

Control of Dipolar Dynamics by Geometrical Programming

We propose and theoretically analyze methods for quantum many-body control through geometric reshaping of molecular tweezer arrays. Dynamic rearrangement during entanglement is readily available due to the extended coherence times of molecular rotational qubits. We show how motional dephasing can be suppressed and enhanced spin squeezing can be achieved in an actively rearranged short-range XY model. We also analyze in detail a specific static geometry that significantly suppresses decoherence. These general methods as applied to programmable quantum systems offer robust control modalities that are well suited to molecules.

optical tweezers↗

On-demand magnon resonance isolation in cavity magnonics

Cavity magnonics is a promising field focusing on the interaction between spin waves (magnons) and other types of signal. In cavity magnonics, isolation of magnons from the cavity to allow signal storage and processing fully in the magnonic domain is highly desired, but its realization is often hindered by the lack of necessary tunability of the interaction. This work shows that by using the collective mode of two yttrium iron garnet spheres and applying Floquet engineering, magnonic signals can be switched on demand to a magnon dark mode that is protected from the environment, enabling a variety of manipulation over the magnon dynamics. Furthermore, our demonstration can be scaled up to systems with an array of magnonic resonators, paving the way for large-scale programmable hybrid magnonic circuits.

42 ENGINEERING↗

Spatiotemporal quenches for efficient critical ground state preparation in the two-dimensional transverse field Ising model

Quantum simulators have the potential to shed light on the study of quantum many-body systems and materials, offering unique insights into various quantum phenomena. Although adiabatic evolution has been conventionally employed for state preparation, it faces challenges when the system evolves too quickly or the coherence time is limited. In such cases, shortcuts to adiabaticity, such as spatiotemporal quenches, provide a promising alternative. This paper numerically investigates the application of spatiotemporal quenches in the two-dimensional transverse field Ising model with ferromagnetic interactions, focusing on the emergence of the ground state and its correlation properties at criticality when the gap vanishes. We demonstrate the effectiveness of these quenches in rapidly preparing ground states in critical systems. Our simulations reveal the existence of an optimal quench front velocity at the emergent speed of light, leading to minimal excitation energy density and correlation lengths of the order of finite system sizes we can simulate. These findings emphasize the potential of spatiotemporal quenches for efficient ground state preparation in quantum systems, with implications for the exploration of strongly correlated phases and programmable quantum computing.

2-dimensional systems↗