Search NASA⌕ Search

SEARCH · Search NASA

Results for “layout”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Evaluation of the Room Temperature Field Quality Measurements of HL-LHC MQXFA Magnet Assemblies

As a member of the multi-lab U.S. High Luminosity LHC Accelerator Upgrade Project (HL-LHC AUP), Lawrence Berkeley National Laboratory (LBNL) is assembling high-field Nb3Sn low-beta MQXFA quadrupole magnets for eventual installation at CERN as part of the HL-LHC upgrade. Each magnet undergoes room temperature magnetic measurements at two points during the assembly process: once after completion of the coil pack sub-assembly and once after the magnet is fully assembled and pre-loaded. Beyond its use to verify that each magnet assembly can meet operational requirements, the data from this two-stage measurement process allow for investigation into the impact of the pre-loading operation on field quality. Recent measurements on magnet rebuilds as well as exploratory coil pack reconfigurations also provide an opportunity to see in isolation the effects of changes to a magnet’s coil selection or coil layout. In this work, we present the magnetic measurement results of MQXFA magnets assembled or currently in process, focusing particularly on magnets with data corresponding to multiple builds, and discuss trends and insights which may be beneficial for the remaining MQXFA magnet production or for future Nb3Sn accelerator magnets.

Doyle, Jennifer [LBNL, Berkeley] (ORCID:0009000079↗

Status Quo and Gap Analysis of Heliostat Field Deployment Processes for Concentrating Solar Tower Plants

In this study, deployment of the solar field of a concentrating solar power plant is one of many factors that are integral to the success of a project. Knowledge transfer from outside the industry is limited due to the unique nature of heliostats, which redirect sunlight to a receiver with high precision while maintaining a high level of reflectivity. Moreover, learning from project to project can be limited due to the site-specific nature of projects, as the market includes several developers, each with their own unique design. In this paper, we discuss the state of the art in heliostat field deployment. We cover all the key aspects of deployment from project assessment to a fully functioning system, which include site selection, layout development, supply chain, assembly, site preparation and construction, calibration, and operations and maintenance. We then perform a gap analysis on field deployment and recommend priorities for future research.

14 SOLAR ENERGY↗

Measurement of the electric potential and the magnetic field in the shifted analysing plane of the KATRIN experiment

The projected sensitivity of the effective electron neutrino-mass measurement with the KATRIN experiment is below 0.3 eV (90 % CL) after 5 years of data acquisition. The sensitivity is affected by the increased rate of the background electrons from KATRIN’s main spectrometer. A special shifted-analysing-plane (SAP) configuration was developed to reduce this background by a factor of two. The complex layout of electromagnetic fields in the SAP configuration requires a robust method of estimating these fields. We present in this paper a dedicated calibration measurement of the fields using conversion electrons of gaseous 83m Kr, which enables the neutrino-mass measurements in the SAP configuration.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Exploring Architectural-Aware Affinity Policies in Modern HPC Runtimes

Modern commodity and High-Performance Computing (HPC) systems are evolving with complex CPU architectures. These architectures now feature higher core and NUMA domain counts and implement features such as hyperthreading. When considering significant differences in hardware configurations, library availability, and hardware-tailored system/software stacks, which could substantially vary from one system to another, performance portability is hard to achieve. Throughout the years, this trend resulted in an increasingly high burden on application developers to fine-tune their workloads for each architecture. This work explores how hardware-dependent aspects such as locality/process/thread affinity affect performance in modern CPU architectures. We focus our study on the Global Memory and Threading (GMT) distributed runtime system as a representative of Partitioned Global Address Space (PGAS) software stacks commonly adopted for productivity. In particular, to appreciate performance implications, we evaluate GMT’s thread affinity policies, and, introduce two new ones which exploit architectural awareness. Finally, we explore alternative NUMA configurations via different process bindings and perform a scalability study on three HPC clusters with varying CPU architectures and NUMA layouts. Our analysis indicates that more complex architectures are more affected by affinity and binding policies and highlights the importance of setting proper runtime configurations to achieve superior performance.

Di Dio Lavore, Ian↗

PyOED: An Extensible Suite for Data Assimilation and Model-Constrained Optimal Design of Experiments

This article describes PyOED, a highly extensible scientific package that enables developing and testing model-constrained optimal experimental design (OED) for inverse problems. Specifically, PyOED aims to be a comprehensive Python toolkit for model-constrained OED. The package targets scientists and researchers interested in understanding the details of OED formulations and approaches. It is also meant to enable researchers to experiment with standard and innovative OED technologies with a wide range of test problems (e.g., simulation models). OED, inverse problems (e.g., Bayesian inversion), and data assimilation (DA) are closely related research fields, and their formulations overlap significantly. Thus, PyOED is continuously being expanded with a plethora of Bayesian inversion, DA, and OED methods as well as new scientific simulation models, observation error models, and observation operators. These pieces are added such that they can be permuted to enable testing OED methods in various settings of varying complexities. The PyOED core is completely written in Python and utilizes the inherent object-oriented capabilities; however, the current version of PyOED is meant to be extensible rather than scalable. Specifically, PyOED is developed to “enable rapid development and benchmarking of OED methods with minimal coding effort and to maximize code reutilization.” This article provides a brief description of the PyOED layout and philosophy and provides a set of exemplary test cases and tutorials to demonstrate the potential of the package.

97 MATHEMATICS AND COMPUTING↗

FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving

Production LLM serving must simultaneously deliver high throughput, low latency, and sufficient context capacity under non-stationary traffic and mixed request requirements. Data parallelism (DP) maximizes throughput by running independent replicas, while tensor parallelism (TP) reduces per-request latency and pools memory for long-context inference. However, existing serving stacks typically commit to a static parallelism configuration at deployment; adapting to bursts, priorities, or long-context requests is often disruptive and slow. We present Flying Serving, a vLLM-based system that enables online DP-TP switching without restarting engine workers. Flying Serving makes reconfiguration practical by virtualizing the state that would otherwise force data movement: (i) a zero-copy Model Weights Manager that exposes TP shard views on demand, (ii) a KV Cache Adaptor that preserves request KV state across DP/TP layouts, (iii) an eagerly initialized Communicator Pool to amortize collective setup, and (iv) a deadlock-free scheduler that coordinates safe transitions under execution skew. Across three popular LLMs and realistic serving scenarios, Flying Serving improves performance by up to 4.79 × under high load and 3.47 × under low load while supporting latency- and memory-driven requests.

Gao, Shouwei [ORNL]↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

FastrSHWFS Analysis

This code analyzes the reflection geometry of modified Shack-Hartmann Wavefront Sensor (SHWFS) masks, with a focus on the fastrSHWFS designs. It takes a height map of a reflective mask, divides it into sub apertures, removes the global focus term by referencing a central block, and fits local surface planes to each sub aperture. From these fits it computes reflection angles and propagates the corresponding rays downstream to determine spot positions and the overall reflected beam footprint as a function of distance. The implementation is parameterized in sub aperture size, number of sub apertures, focal length, and mask dimensions, so it can be adapted to different mask designs beyond the two current fastrSHWFS masks. The code: Converts mask bitmaps into physical height in microns and reconstructs a focus subtracted mask, divides the mask into a grid of sub apertures and assigns pixel and micron coordinates to each block, excludes user specified unused sub apertures from the analysis, fits a plane to each sub aperture to obtain local surface normals, propagates reflected rays to a range of z positions to compute ray spots and the full reflected beam width, computes the effective mask tilt angle with respect to the incoming beam as a function of propagation distance, and provides helper routines for plotting the grid over the mask and for visualizing the beam geometry. These tools are intended for iterating on fastrSHWFS mask designs and for planning downstream optical layouts, for example choosing lens positions and apertures that capture the reflected beam given known focal plane distances and beam widths.

Gerard, BenjaminL [Lawrence Livermore National Lab↗

Polarized Deep-Inelastic Scattering with Spin Correlations in Herwig 7

This repository is the research software and reproducibility companion for the HerwigPol polarized deep-inelastic scattering implementation developed for Herwig 7. It brings together the modified Herwig and ThePEG source snapshots, the curated POLDIS fixed-order reference code, the custom Rivet analyses, the DIS validation workflow, and the paper source in a single formal repository layout. The repository is intended to preserve the source-level ingredients needed to rebuild and re-run the validated DIS studies. It therefore tracks code, input cards, workflow drivers, and technical notes, while intentionally excluding generated artifacts such as build products, campaign outputs, merged YODA files, plots, and rendered paper outputs.

Papaefstathioou, Andreas [Kennesaw State Universit↗

Designing and prototyping extensions to the Message Passing Interface in MPICH

As HPC system architectures and the applications running on them continue to evolve, the MPI standard itself must evolve. The trend in current and future HPC systems toward powerful nodes with multiple CPU cores and multiple GPU accelerators makes efficient support for hybrid programming critical for applications to achieve high performance. However, the support for hybrid programming in the MPI standard has not kept up with recent trends. The MPICH implementation of MPI provides a platform for implementing and experimenting with new proposals and extensions to fill this gap and to gain valuable experience and feedback before the MPI Forum can consider them for standardization. Here, in this work, we detail six extensions implemented in MPICH to increase MPI interoperability with other runtimes, with a specific focus on heterogeneous architectures. First, the extension to MPI generalized requests lets applications integrate asynchronous tasks into MPI’s progress engine. Second, the iovec extension to datatypes lets applications use MPI datatypes as a general-purpose data layout API beyond just MPI communications. Third, a new MPI object, MPIX_Stream, can be used by applications to identify execution contexts beyond MPI processes, including threads and GPU streams. MPIX stream communicators can be created to make existing MPI functions thread-aware and GPU-aware, thus providing applications with explicit ways to achieve higher performance. Fourth, MPIX Streams are extended to support the enqueue semantics for offloading MPI communications onto a GPU stream context. Fifth, thread communicators allow MPI communicators to be constructed with individual threads, thus providing a new level of interoperability between MPI and on-node runtimes such as OpenMP. Lastly, we present an extension to invoke MPI progress, which lets users spawn progress threads with fine-grained control to adapt the communication performance to their application designs. We describe the design and implementation of these extensions, provide usage examples, and highlight their expected benefits with performance results.

97 MATHEMATICS AND COMPUTING↗

ORCHA: A performance portability system for extreme heterogeneity

Heterogeneity is the prevalent trend in the rapidly evolving high-performance computing (HPC) landscape in both hardware and application software. The diversity in hardware platforms, currently comprising various accelerators and a future possibility of specializable chiplets, poses a significant challenge for scientific software developers aiming to harness optimal performance across different computing platforms while maintaining the quality of solutions when their applications are simultaneously growing more complex. Code synthesis and code generation can provide mechanisms to mitigate this challenge. We have developed a divide and conquer approach where different aspects of performance are handled by different stand-alone tools that are interfaced with the application through generated code. This portability system, ORCHA, enables users to configure and orchestrate their computations among available resources on a platform by specifying a high-level recipe, thereby permitting a many-to-many paradigm where each recipe results in a different variant of the application. The core design goal is to let users decide the application’s hardware mapping and orchestration by editing only the high-level recipe—without modifying the maintained source code or binding the application to a particular runtime system. Tools in ORCHA distribution are: CG-Kit for translating the recipe into an execution graph; Milhoja to execute the graph by orchestrating data and task movement among hardware resources; and Macroprocessor that enables users to define their own code-shorthand for higher composability and easier management of code variants. Additionally, the design of ORCHA permits tools to work in a plug-and-play mode where the application can build and run without CG-Kit and Milhoja, and either tool can be swapped out for other tools with similar capabilities by modifying the code generation portion of ORCHA. In this paper, we describe the design of ORCHA and the role that code-generation plays in isolating applications from tools. We demonstrate the breadth of configurations ORCHA enables with a case study in which an application configuration is realized on three distinct hardware mappings—a GPU-centric, a CPU/GPU balanced, and a CPU/GPU concurrent layouts by using different recipes.

Lee, Youngjun↗

Community Geothermal: Soil Conductivity, Borehole Design, Energy Models, and Load Data for a Residential System Development - Hinesburg, VT

This dataset contains materials from the Coalition for Community-Supported Affordable Geothermal Energy Systems (C2SAGES) project, which evaluated the techno-economic feasibility of a community geothermal system for a residential development in Hinesburg, VT. The dataset includes detailed soil conductivity test reports, energy models, borehole design reports, hourly energy loads for heating, cooling, and hot water, and design layouts. EnergyPlus was used to model building energy loads, and Modelica software was applied for geothermal loop sizing based on these loads and soil conductivity results. Python scripts for network design further refined the models. Key files include PDF reports on borehole design (with projections for 1-year, 15-year, and 30-year systems), soil conductivity test results, EnergyPlus modeling outputs, and 2D/3D design drawings in PDF, DWG, and DXF formats. Python notebooks for network design and OnePipe model files are also provided, with Modelica required for viewing certain files. Outputs and modeling data are in various formats including CSV, JPG, HTML, and IDF, with units and data clearly labeled to support understanding of system design and performance for the proposed geothermal solution.

15 GEOTHERMAL ENERGY↗

Geomatchd

Geomatchd is a software tool to assist non-experts in the rapid conceptual design of geothermal district heating systems. This tool comprises three main parts: user-side heating and cooling demand profiles, district heating and cooling (DHC) network models for sizing and optimal layout selection, and geothermal system models for heat extraction calculation and cost comparison. The architecture of Geomatchd is opensource and is intended to simplify efforts by other coders to improve on existing functions, to add functions and to leverage available existing code and data – to improve the tool for public use. Geomatchd and the material supporting its development may provide a more intuitive window on district heating and the use of local shallow geothermal resources.

AS↗

STATUS OF THE SECOND INTERACTION REGION DESIGN FOR ELECTRON-ION COLLIDER

Provisions are being made in the Electron Ion Collider (EIC) design for future installation of a second Interaction Region (IR), in addition to the day-one primary IR. The envisioned location for the second IR is the existing experi- mental hall at RHIC IP8. It is designed to work with the same beam energy combinations as the first IR, covering a full range of the center-of-mass energy of ?20 GeV to ?140 GeV. The goal of the second IR is to complement the first IR, and to improve the detection of scattered particles with magnetic rigidities similar to those of the ion beam. To achieve this, the second IR hadron beamline features a secondary focus in the forward ion direction. The design of the second IR is still evolving. This paper reports the current status of its pa- rameters, magnet layout, and beam dynamics and discusses the ongoing improvements being made to ensure its optimal performance.

Gamage, B.↗

Progress in the design of the future circular collider FCC-ee interaction region

In this paper we discuss the latest developments for the FCC-ee interaction region layout, which represents one of the key ingredients to establish the feasibility of the FCC-ee. The collider has to achieve extremely high luminosities over a wide range of center-of-mass energies with two or four interaction points. The complex final focus hosted in the detector region has to be carefully designed, and the impact of beam losses and of any type of synchrotron radiation generated in the interaction region, including beamstrahlung, have to be evaluated in detail with simulations. We give an overview of the progress of the whole machine-detector-interface-related studies, among which are the updated mechanical model of the interaction region, the plans for a novel R&D activity of a IR mockup which is just starting, the collimation scheme and evaluation of beam induced backgrounds in the detectors, evaluation of radiation dose in the experimental area, and MDI integration with the detector.

43 PARTICLE ACCELERATORS↗

Selected advances in the accelerator design of the Future Circular Electron-Positron Collider (FCC-ee)

In autumn 2023, the FCC Feasibility Study underwent a crucial “mid-term review”. We describe some accelerator performance risks for the proposed future circular electron-positron collider, FCC-ee, identified for, and during, the mid-term review. For the collider rings, these are the collective effects when running on the Z resonance – especially resistive wall, beam-beam, and electron cloud –, the beam lifetime, dynamic aperture, alignment tolerances, and beam-based alignment. For the booster, the primary concern is the vacuum system, with regard to impedance and effects of the residual gas. For the injector, the layout and the linac repetition rate are primary considerations. We discuss the various issues and report the planned mitigations.

43 PARTICLE ACCELERATORS↗

FFA BEAM TRANSPORT DEMONSTRATION DEVELOPMENT FOR THE CEBAF 22 GeV UPGRADE

Jefferson Lab is planning an upgrade of the Continuous Electron Beam Accelerator Facility (CEBAF) to deliver highly polarized electron beams up to 22 GeV using Fixed- Field Alternating-gradient (FFA) magnets. As the application of FFA technology in the 10–22 GeV energy range is unprecedented, experimental validation is required prior to full-scale implementation. To support this effort, a dedicated FFA test insert is proposed within the existing CEBAF infrastructure, with candidate locations in the Beam Switchyard (BSY) dump line or the Hall C beamline. The testbed will consist of a half or full FFA cell using combined-function permanent magnets and will enable systematic studies of beam transport, field quality, alignment, and magnet performance under realistic conditions. Operation with polarized beams in the 5–11 GeV range will closely replicate the energy scaling of the full upgrade. This paper presents the current design status and layout options for the proposed FFA beam transport test line.

Ogur, S. [Thomas Jefferson National Accelerator Fa↗

Extraction and Injection in the Electron Injector for the Electron-Ion Collider

The electron injector for the Electron-Ion Collider (EIC) consists of a linear accelerator, a beam accumulation ring, and the Rapid Cycling Synchrotron (RCS) before the electrons are injected into the Electron Storage Ring (ESR) and collided. Extraction out of the RCS is complicated by limited space and the nominal beam pipe aperture, while injection into the ESR is complicated due to the limitation of kicker strength, so that the kickers will not impact the proton beam in the adjacent Hadron Storage Ring (HSR); additionally, the ESR kickers must also provide enough kick to the stored bunch for the swap-out scheme. This paper covers the injection into and extraction out of the RCS, as well as injection into the ESR, detailing layout, optics, and anticipated parameters of the septa and different kickers.

Deitrick, K. [Thomas Jefferson National Accelerato↗