Search NASASearch

SEARCH · Search NASA

Results for “serialization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Readout optimization of multi-amplifier sensing charge-coupled devices for single-quantum measurement

The non-destructive readout capability of the Skipper Charge Coupled Device (CCD) has been demonstrated to reduce the noise limitation of conventional silicon devices to levels that allow single-photon or single-electron counting. The noise reduction is achieved by taking multiple measurements of the charge in each pixel. These multiple measurements come at the cost of extra readout time, which has been a limitation for the broader adoption of this technology in particle physics, quantum imaging, and astronomy applications. This work presents recent results of a novel sensor architecture that uses multiple non-destructive floating-gate amplifiers in series to achieve sub-electron readout noise in a thick, fully-depleted silicon detector to overcome the readout time overhead of the Skipper-CCD. This sensor is called the Multiple-Amplifier Sensing Charge-Coupled Device (MAS-CCD) can perform multiple independent charge measurements with each amplifier, and the measurements from multiple amplifiers can be combined to further reduce the readout noise. We will show results obtained for sensors with 8 and 16 amplifiers per readout stage in new readout operations modes to optimize its readout speed. The noise reduction capability of the new techniques will be demonstrated in terms of its ability to reduce the noise by combining the information from the different amplifiers, and to resolve signals in the order of a single photon per pixel. The first readout operation explored here avoids the extra readout time needed in the MAS-CCD to read a line of the sensor associated with the extra extent of the serial register. The second technique explore the capability of the MAS-CCD device to perform a region of interest readout increasing the number of multiple samples per amplifier in a targeted region of the active area of the device.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel

FAIRLinked: Data FAIRification Tools for Materials Data Science

FAIRLinked is a software package created to support the FAIRification of materials science data, ensuring proper alignment with FAIR principles: Findable, Accessible, Interoperable, and Reusable. It is built to be compatible with MDS-Onto, an ontology designed to capture the semantics of various types of materials data, enabling integration and sharing across different research workflows. The package is subdivided into three subpackages: InterfaceMDS, RDFTableConversion, and QBWorkflow. The first subpackage, InterfaceMDS allows users to search for terms using either string search or various filters, explore different domains and subdomains, and add terms to MDS-Onto. RDFTableConversion is used for serialization and deserialization of data from CSV into JSONLDs and vice versa in a way that captures the semantics of the data using MDS-Onto. Lastly, QBWorkflow is a serialization and deserialization workflow that incorporates RDF Data Cube vocabulary, useful for working with multidimensional datasets. By offering these packages, FAIRLinked lowers the barrier of creating FAIR, machine-actionable data for researchers in the materials science community.

FAIR

Xyce™ Parallel Electronic Simulator Users' Guide (V.7.9)

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: • Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. • A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. • Device models that are specifically tailored to meet Sandia’s needs, including some radiation-aware devices (for Sandia users only). • Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase — a message passing parallel implementation — which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

42 ENGINEERING

Xyce™ Parallel Electronic Simulator Users’ Guide, Version 7.10

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: • Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. • A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. • Device models that are specifically tailored to meet Sandia’s needs, including some radiation-aware devices (for Sandia users only). • Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase — a message passing parallel implementation — which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

97 MATHEMATICS AND COMPUTING

Field Validation of Dynamic Mechanical Torque Measurements for Geared Wind Turbines

Accurate knowledge of the mechanical loads of wind turbine gearboxes has become essential in modern, highly loaded gearbox designs, as maintaining or even improving gearbox reliability with increasing torque density demands is proving to be challenging. Unfortunately, the traditional method of measuring dynamic mechanical torque using strain gauges placed on the outer surface of a rotating shaft and transmitting the resulting signal is unsuitable for serial deployment due to technical and economic constraints. An alternative method based on fiber-optic strain sensors placed on the stationary outer surface of the gearbox ring gear has been proposed. Like shaft torsion, the radial deformation of the ring gear is proportionate to the rotor torque. Placing the sensors on a stationary component is a cost-effective alternative for serial implementation because the need for complex and expensive data transfer via wireless transmission or a slip ring is eliminated. In this paper, we present the results of an extensive field experiment conducted to evaluate the torque measurement accuracy of this novel sensing solution installed on the gearbox of a Gamesa G97 2-MW wind turbine at the National Renewable Energy Laboratory's Flatirons Campus. Torque measurements derived from fiber-optic strain sensors placed on the ring gear of the planetary stage are compared to conventional torque measurements from strain gauges placed on the main shaft. Two different torque estimation data processing methods were evaluated, with the method based on operational deflection shapes providing the most accurate results with an average normalized root mean square error below 0.7% for a load revolution distribution analysis. The effect of operating conditions on the torque estimate was also investigated, and the third planet-passing operational deflection shape was found to be the least sensitive to nontorque load-related effects. The fiber-optic strain sensors' successful operation during the complete test campaign has demonstrated a robust and accurate solution for fleet-wide enhanced gearbox remaining useful life estimation.

17 WIND ENERGY

Defect And Damage Characterization Of Additively Manufactured Titanium Alloy Ti-5553 Using Traditional Computed Tomography Volume Segmentation And Machine Learning Algorithms

The mechanical response of a component is affected by defects, such as porosity, arising from the laser powder bed fusion (LPBF) fabrication process. Thus, it is important to develop accurate and efficient inspection methods for identifying porosity. In this work, porosity identified in an X-ray computed tomography (XCT) volume of a Ti-5553 coupon was compared to pores identified in a serial sectioned volume that represented the ground truth. The porosity of the XCT scan was identified using contrast-based, ISO-based, and machine learning (ML) methods for segmentation. Large inherent porosity was easy to identify, but the ISO thresholding still struggled due to the intensity gradient resulting from both the beam hardening in XCT and the uneven lighting of the serial sectioning panels. Further, the results show that ML-based methods were better suited for identifying small pores and reducing the amount of false positives. Additionally, high strain-rate impact testing was done on some of the XCT samples as well as post-mortem XCT inspection, and the same suite of segmentation and quantification tools were used to identify the large spallation cavities. The comparison of porosity pre- and post-mortem provides insight on the influence of the LPBF porosity on the formation of spall cavities.

36 MATERIALS SCIENCE

Exponential Backoff and Its Security Implications for Safety-Critical OT Protocols over TCP/IP Networks

The convergence of Operational Technology (OT) and Information Technology (IT) networks has become increasingly prevalent with the growth of Industrial Internet of Things (IIoT) applications. This shift, while enabling enhanced automation, remote monitoring, and data sharing, also introduces new challenges related to communication latency and cybersecurity. Oftentimes, legacy OT protocols were adapted to the TCP/IP stack without an extensive review of the ramifications to their robustness, performance, or safety objectives. To further accommodate the IT/OT convergence, protocol gateways were introduced to facilitate the migration from serial protocols to TCP/IP protocol stacks within modern IT/OT infrastructure. However, they often introduce additional vulnerabilities by exposing traditionally isolated protocols to external threats. This study investigates the security and reliability implications of migrating serial protocols to TCP/IP stacks and the impact of protocol gateways, utilizing two widely used OT protocols: Modbus TCP and DNP3. Our protocol analysis finds a significant safety-critical vulnerability resulting from this migration, and our subsequent tests clearly demonstrate its presence and impact. A multi-tiered testbed, consisting of both physical and emulated components, is used to evaluate protocol performance and the effects of device-specific implementation flaws. Through this analysis of specifications and behaviors during communication interruptions, we identify critical differences in fault handling and the impact on time-sensitive data delivery. The findings highlight how reliance on lower-level IT protocols can undermine OT system resilience, and they inform the development of mitigation strategies to enhance the robustness of industrial communication networks.

DNP3

Structural Mechanism of an Efficacy Photoswitch Targeting the β 2 ‐adrenergic Receptor

The field of photopharmacology develops light-responsive drugs that can modulate protein activity, enabling precise and dynamic investigations of their roles in health and disease. Adrenergic receptors are prominent targets for this approach because they are prototypical G protein-coupled receptors with high clinical relevance in bronchial and cardiovascular diseases. Here, we employed the azobenzene-based compound photoazolol-1 in combination with time-resolved serial crystallography at X-ray free-electron lasers to resolve the molecular mechanisms by which photoswitchable β-blockers modulate activity of the β 2 -adrenoceptor (β 2 AR). Time-resolved structures of the receptor bound to trans-photoazolol-1 (pre-photoconversion), a strained intermediate in the nanosecond range, and the fully photoisomerized cis-photoazolol-1 reveal how isomerization of the azobenzene moiety induces distinct conformational changes within the orthosteric ligand binding pocket. Within seconds, light-excited photoazolol-1 adopts a new binding pose, altering interactions with extracellular loop 2 and shifting the positions of transmembrane helices 5, 6, and 7. Functional assays of β 2 AR in cellular membranes show that photoazolol-1 acts as an efficacy photoswitch, changing from an inverse agonist to a neutral antagonist upon isomerization without leaving the binding pocket. In combination, these findings suggest a molecular mechanism for activity modulation via efficacy photoswitches and provide a framework for designing ligands that exploit light-driven transitions within the binding pocket to achieve spatiotemporal control of receptor function.

G protein-coupled receptors

Comparison of three measurement modalities for 3D characterization of manufactured features and process-induced porosity in titanium alloy additively manufactured parts

Nondestructive characterization of internal features and defects within complex components is vital for many industrial applications, particularly with the advent of additive manufacturing (AM) technologies. However, community understanding of the limitations of nondestructive methods such as X-ray Computed Tomography (CT) can be limited in certain industrial sectors as these may be emergent applications. In this paper, we investigate the limits of X-ray CT measurements and compare extracted data with mechanical polishing serial sectioning (MPSS) and confocal laser scanning microscopy (CLSM). The test object is an additively manufactured titanium alloy disk that contains both process-induced porosity and machined features, including focused ion beam milled features designed to probe the resolution limits of X-ray CT. Results show that each of these characterization techniques has advantages and disadvantages. We compare data acquisition times, spatial resolution, geometric measurement accuracy and defect visualization fidelity across these modalities to establish a practical framework.

Additive manufacturing

Macromolecular crystallography and biology at the Linac Coherent Light Source

The Linac Coherent Light Source (LCLS) has significantly impacted the field of biology by providing advanced capabilities for probing the structure and dynamics of biological molecules with high precision. The ultrashort coherent X-ray pulses from the LCLS have enabled ultrafast, time-resolved, serial femtosecond crystallography that is inaccessible at conventional synchrotron light sources. Since the facility's founding, scientists have captured detailed insights into biological processes at atomic resolution and fundamental timescales. The ability to observe these processes in real time and under conditions closely resembling their natural state is transforming our approach to studying biochemical mechanisms and developing new medical and energy applications. This work recounts some of the history of the LCLS, advances in biological research enabled by the LCLS, key biological areas that have been impacted and how the LCLS has helped to unravel complex biological phenomena in these fields.

59 BASIC BIOLOGICAL SCIENCES

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES

Ligand‐Mediated Quantum Yield Enhancement in 1‐D Silver Organothiolate Metal–Organic Chalcogenolates

X-ray free electron laser (XFEL) microcrystallography and synchrotron single-crystal crystallography are used to evaluate the role of organic substituent position on the optoelectronic properties of metal–organic chalcogenolates (MOChas). MOChas are crystalline 1D and 2D semiconducting hybrid materials that have varying optoelectronic properties depending on composition, topology, and structure. While MOChas have attracted much interest, small crystal sizes impede routine crystal structure determination. A series of constitutional isomers where the aryl thiol is functionalized by either methoxy or methyl ester are solved by small molecule serial femtosecond X-ray crystallography (smSFX) and single crystal rotational crystallography. While all the methoxy examples have a low quantum yield (0-1%), the methyl ester in the ortho position yields a high quantum yield of 22%. Here, the proximity of the oxygen atoms to the silver inorganic core correlates to a considerable enhancement of quantum yield. Four crystal structures are solved at a resolution range of 0.8–1.0 Å revealing a collapse of the 2D topology for functional groups in the 2- and 3- positions, resulting in needle-like crystals. Further analysis using density functional theory (DFT) and many-body perturbation theory (MBPT) enables the exploration of complex excitonic phenomena within easily prepared material systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Probing substrate water access through the O1 channel of Photosystem II by single site mutations and membrane inlet mass spectrometry

Abstract Light-driven water oxidation by photosystem II sustains life on Earth by providing the electrons and protons for the reduction of CO 2 to carbohydrates and the molecular oxygen we breathe. The inorganic core of the oxygen evolving complex is made of the earth-abundant elements manganese, calcium and oxygen (Mn 4 CaO 5 cluster), and is situated in a binding pocket that is connected to the aqueous surrounding via water-filled channels that allow water intake and proton egress. Recent serial crystallography and infrared spectroscopy studies performed with PSII isolated fromThermosynechococcus vestitus(T. vestitus) support that one of these channels, the O1 channel, facilitates water access to the Mn 4 CaO 5 cluster during its S 2 →S 3 and S 3 →S 4 →S 0 state transitions, while a subsequent CryoEM study concluded that this channel is blocked in the cyanobacteriumSynechocystis sp.PCC 6803, questioning the role of the O1 channel in water delivery. Employing site-directed mutagenesis we modified the two O1 channel bottleneck residues D1-E329 and CP43-V410 (T. vestitusnumbering) and probed water access and substrate exchange via time resolved membrane inlet mass spectrometry. Our data demonstrates that water reaches the Mn 4 CaO 5 cluster via the O1 channel in both wildtype and mutant PSII. In addition, the detailed analysis provides functional insight into the intricate protein-water-cofactor network near the Mn 4 CaO 5 cluster that includes the pentameric, near planar ‘water wheel’ of the O1 channel.

Plant Sciences

Effect of stress on Laves phase precipitation in creep ruptured Grade 92 ferritic martensitic steel characterized by a novel accessible method

At high temperature conditions relevant to fossil and nuclear energy plants, Laves phase (Fe 2 X, X=Mo, W) precipitation is observed in common ferritic martensitic (FM) structural steels, with various reported effects on creep behavior. Despite being valuable metrics to correlate with mechanical properties and other precipitate phases, the volume fraction and number density of Laves phase precipitates has been difficult to quantify accurately using common techniques such as transmission electron microscopy (TEM) due to the relatively large size (∼0.25 μm) and low number density (∼10 11 cm −3 ) of Laves precipitates. Here, to address this characterization challenge, we developed and demonstrated a high-throughput and widely accessible method to quantify the volume fraction and number density of the Laves phase based on scanning electron microscope (SEM) images with a backscattered electron signal and the information depth (ID) of backscattered electrons. We applied this new technique in creep ruptured Grade 92 FM steel to study the effect of Laves phase on creep properties and determine the influence of stress on Laves phase precipitation. The quantitative accuracy of the SEM-based volume fraction and number density values was verified using synchrotron high energy X-ray diffraction and serial sectioning tomography. Stress did not significantly affect the Laves phase size or volume fraction during creep testing at 550 – 650°C and stress levels of 90 – 260 MPa (vs. unstressed conditions). Conversely, a moderate but statistically significant stress-enhanced increase in Laves phase number density, corresponding to an increase in nucleation rate, occurred during creep exposure above 110 MPa.

Creep

Noise-aware optimization in nominally identical manufacturing and measuring systems for high-throughput parallel workflows

Device-to-device variability in experimental noise critically impacts reproducibility, especially in automated, high-throughput systems like additive manufacturing farms. While manageable in small labs, such variability can escalate into serious risks at larger scales, such as architectural 3D printing, where noise may cause structural or economic failures. This contribution presents a noise-aware decision-making algorithm that quantifies and models device-specific noise profiles to manage variability adaptively. It uses distributional analysis and pairwise divergence metrics with clustering to choose between single-device and robust multi-device Bayesian optimization strategies. Unlike conventional methods that assume homogeneous devices or enforce generic robustness, the proposed framework explicitly determines whether shared optimization across devices is appropriate based on the degree of inter-device noise heterogeneity. This enables improved performance, reproducibility, and efficiency. An experimental case study involving three nominally identical 3D printers (same brand, model, and close serial numbers) demonstrates reduced redundancy, lower resource usage, and improved reliability, along with improved convergence stability and solution quality through the selection of the appropriate optimization strategy based on the degree of inter-device noise heterogeneity. Overall, this framework establishes a general approach for precision- and resource-aware optimization in scalable, automated experimental platforms, demonstrated here on a representative multi-device 3D printing case study.

Schenk, Christina