Precision Cascade: A novel algorithm for multi-precision extreme compression
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The explosive demand for artificial intelligence (AI) workloads has led to a significant increase in silicon area dedicated to lower-precision computations on recent high-performance computing hardware designs. However, mixed-precision capabilities, which can achieve performance improvements of up to 8x compared to double-precision in extreme compute-intensive workloads, remain largely untapped in most scientific applications. A growing number of efforts have shown that mixed-precision algorithmic innovations can deliver superior performance without sacrificing accuracy. These developments should prompt computational scientists to seriously consider whether their scientific modeling and simulation applications could benefit from the acceleration offered by new hardware and mixed-precision algorithms. In this survey, we (1) review progress across diverse scientific domains—fluid dynamics, weather and climate, quantum chemistry, and computational genomics—that have begun adopting mixed-precision strategies; (2) examine state-of-the-art algorithmic techniques such as iterative refinement, splitting and emulation schemes, and adaptive precision solvers; (3) assess their implications for accuracy, performance, and resource utilization; and (4) survey the emerging software ecosystem that enables mixed-precision methods at scale. We conclude with perspectives and recommendations on cross-cutting opportunities, domain-specific challenges, and the role of co-design between application scientists, numerical analysts, and computer scientists. Collectively, this survey underscores that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving landscape of hardware capabilities.
Mixed-precision algorithms have been proposed as a way for scientific computing to benefit from some of the gains seen for AI on recent high performance computing (HPC) platforms. A few applications dominated by dense matrix operations have seen substantial speedups by utilizing low precision formats such as FP16. However, a majority of scientific simulation applications are memory bandwidth limited. Beyond preliminary studies, the practical gain from using mixed-precision algorithms on a given high-performance computing (HPC) system is largely unclear. The High Performance GMRES Mixed Precision (HPG-MxP) benchmark has been proposed to measure the useful performance of a HPC system on sparse matrix-based mixed-precision applications. In this work, we present an implementation of the HPG-MxP benchmark for an exascale system and describe our algorithm enhancements. We show for the first time a speedup of 1.6x using a combination of double- and single-precision keeping the same residual level on modern GPU-based supercomputers.
In nature, many complex multi‐physics coupling problems exhibit significant diffusivity inhomogeneity, where one process occurs several orders of magnitude faster than others temporally. Simulating rapid diffusion alongside slower processes demands intensive computational resources due to the necessity for small time steps. To address these computational challenges, we have developed an efficient numerical solver named Finite Difference informed Random Walker (FDiRW). In this study, we propose a GPU‐accelerated, mixed‐precision configuration for the FDiRW solver to maximize efficiency through GPU multi‐threaded parallel computation and lower precision computation. Numerical evaluation results reveal that the proposed GPU‐accelerated mixed‐precision FDiRW solver can achieve a 117× speedup over the CPU baseline, while an additional 1.75× speedup is achieved by employing lower precision GPU computation. Notably, for large model sizes, the GPU‐accelerated mixed‐precision FDiRW solver demonstrates strong scaling with the number of nodes used in simulation. When simulating radionuclide absorption processes by porous wasteform particles with a medium‐sized model of 192 × 192 × 192, this approach reduces the total computational time to 10 min, enabling the simulation of larger systems with strongly inhomogeneous diffusivity.
Absolute gas counting (AGC) was applied to two gas blends of 85Kr in argon-methane (P10) counting gas to establish a high-precision specific activity (Bq/cm3) reference value for characterizing 85Kr detection efficiency for groundwater age dating measurements. The AGC or length-compensated technique has been utilized by the metrology community for decades and is an accepted method for developing radioactive gas standards. The AGC capability at Pacific Northwest National Laboratory (PNNL) uses a set of nine unequal-length proportional counters with precisely-measured internal volumes, and a gas loading system with high-precision pressure and temperature sensors. A series of AGC measurements were collected at multiple pressures to determine the inverse pressure relationship (1/P) for 85Kr and define a wall-effect correction that accounts for events decaying into the detector wall and not depositing sufficient energy in the gas to be detected. In addition to the wall-effect, two additional corrections were evaluated and are discussed in detail. Specifically, the threshold effect which accounts for events deposited below the analysis threshold and a detection efficiency as a function of detector volume effect that was observed during analysis. A robust uncertainty model was developed using the Guide to the expression of Uncertainty in Measurements (GUM) approach. The combination of carefully scrutinized correction factors, precise measurements of pressure, temperature and detector volume, and robust counting statistics resulted in the determination of high-precision specific activity values with 0.50% or less total combined uncertainty for two Kr-in-P10 reference gas standards (KP10) that will enable new groundwater age-dating measurements at PNNL.
In the standard model of particle physics, the masses of the W and Z bosons, the carriers of the weak interaction, are uniquely related. A precise determination of their masses is important because quantum loops of heavy, undiscovered particles could modify this relationship. Although the Z mass is known to the remarkable precision of 22 parts per million (2.0 MeV), the W mass is known much less precisely. A global fit to measured electroweak observables predicts the W mass with 6 MeV uncertainty [1$-$3]. Reaching a comparable experimental precision would be a sensitive and fundamental test of the standard model, made even more urgent by a recent challenge to the global fit prediction by a measurement from the CDF Collaboration at the Fermilab Tevatron collider [4]. Here we report the measurement of the W mass by the CMS Collaboration at the CERN LHC, based on a large data sample of $W \to \mu \nu$ events collected in 2016 at the proton-proton collision energy of 13 TeV. The measurement exploits a high-granularity maximum likelihood fit to the kinematic properties of muons produced in W decays. By combining an accurate determination of experimental effects with marked in situ constraints of theoretical inputs, we reach a precise measurement of the W mass, of 80 360.2 $\pm$ 9.9 MeV, in agreement with the standard model prediction.
The advent of high-intensity, high-polarization electron beams led to significantly improved measurements of the ratio of the proton’s charge to electric form factors, 𝐺 𝐸 𝑝 /𝐺 𝑀 𝑝 . However, high-𝑄 2 measurements of this ratio yielded significant disagreement with extractions based on unpolarized scattering measurements, raising questions about the reliability of the measurements and consistency of the techniques. Jefferson Lab experiment E01-001 was designed to provide a high precision extraction of 𝐺 𝐸 𝑝 /𝐺 𝑀 𝑝 from unpolarized cross-section measurements using a modified version of the Rosenbluth separation technique to allow for a more precise comparison with polarization data. Rosenbluth separations involve precise measurements of the angular dependence of the elastic 𝑒−𝑝 cross section at fixed momentum transfer, 𝑄 2 . Conventional Rosenbluth separations detect the scattered electron, requiring the comparisons of measurements with very different detected electron energy and rate for electrons at different angles. Our ‘‘super-Rosenbluth’’ measurement detected the struck proton, rather than the scattered electron to extract the elastic 𝑒−𝑝 cross section. This yielded a fixed momentum for the detected particle and dramatically reduced variation of the cross section with angle, significantly reducing rate- and momentum-dependent corrections and uncertainties. We measure the cross section vs angle with high relative precision, allowing for extremely high precision extractions of 𝐺 𝐸 𝑝 /𝐺 𝑀 𝑝 at 𝑄 2 = 2.64, 3.20, and 4.10 GeV 2 . Our results are consistent with traditional Rosenbluth extractions, but with much smaller corrections and systematic uncertainties, comparable to the uncertainties from polarization measurements. Our data confirm the discrepancy between Rosenbluth and polarization extractions of the proton form factor ratio using an improved Rosenbluth extraction that yields smaller and less-correlated uncertainties than those typical of previous Rosenbluth extractions. Here, we compare our results to calculations of two-photon exchange effects and find that the observed discrepancy can be relatively well explained by such effects.
Optical atomic clocks with unrivaled precision and accuracy have advanced the frontier of precision measurement science and opened new avenues for exploring fundamental physics. A fundamental limitation on clock precision is the standard quantum limit (SQL), which stems from the uncorrelated projection noise of each atom. State-of-the-art optical lattice clocks interrogate large ensembles to minimize the SQL, but density-dependent frequency shifts pose challenges to scaling the atom number. The SQL can be surpassed, however, by leveraging entanglement, though it remains an open problem to achieve quantum advantage from spin squeezing at state-of-the-art stability levels. Here, we demonstrate clock performance beyond the SQL, achieving a fractional frequency precision of 1.1 × 10 −18 for a single spin-squeezed clock. With cavity-based quantum nondemolition measurements, we prepare two spin-squeezed ensembles of ∼30 000 strontium atoms confined in a two-dimensional optical lattice. A synchronous clock comparison with an interrogation time of 61 ms achieves a metrological improvement of 2.0(2) dB beyond the SQL, after correcting for state preparation and measurement errors. These results establish the most precise entanglement-enhanced clock to date and offer a powerful platform for exploring the interplay of gravity and quantum entanglement.
The quantum approximate optimization algorithm (QAOA) is a hybrid quantum-classical algorithm that seeks to achieve approximate solutions to optimization problems by iteratively alternating between intervals of controlled quantum evolution. Here, we examine the effect of analog precision errors on QAOA performance from the perspective of both algorithmic training and performance guarantees. Leveraging cumulant expansions, we recast the faulty QAOA as a control problem in which precision errors are expressed as multiplicative control noise and derive bounds on the performance of QAOA. We show using both analytical techniques and numerical simulations that fixed precision implementations of QAOA circuits are subject to an exponential degradation in performance dependent upon the number of optimal QAOA layers and magnitude of the precision error. Despite this significant reduction, we show that it is possible to mitigate precision errors in QAOA via digitization of the variational parameters at the cost of increasing circuit depth.
Precision agriculture, where sensing of soil, environment and crop conditions are used to precisely synchronize inputs (such as water and fertilizer) to crop needs enhances input use efficiency. This can improve yields and farm profitability while mitigating environmental losses, improving soil carbon content and substantially decreasing energy use for food, feed and fuel crops. Unfortunately, farmers are not yet able to harness the full potential of these management technologies as there is a lack of available management information, and there is therefore a need for sensors that are able to economically measure spatio-temporal variability in soil and crop properties of extremely heterogeneous farm fields precisely at high resolution and at low cost. Real-time, in-situ monitoring of agricultural soil conditions is today carried out using devices that limit the total number of nodes that can be used economically to typically one per acre or less. Higher spatio-temporal resolution sensing would enable more precise agricultural input optimization, with significant benefits to the farmer and the environment. In order to address this issue, this project focused on developing additively manufactured, biodegradable, soil sensors with predicted costs of < $\$$1 per unit to monitor crop inputs (such as water and fertilizer) that predictably, harmlessly degrade away into the soil when no longer needed. These sensor nodes should be easy to place, accurately and continuously monitor soil and crop conditions for an entire season, be read remotely using existing farm equipment, require no ongoing maintenance, not impede farm operations and produce no persistent waste. This approach could enable a >100× increase in information density over current solutions for precision farming of row and other crops, and lead to significant reductions in input energy use and provide increased yield for biofuel crops. Over the course of this project the team at the University of Colorado Boulder, University of California Berkeley, and Colorado State University/Kansas State University investigated a wide range of printable biodegradable electronic materials and sensor designs for determining soil moisture and soil nitrate concentration. These efforts expanded the available materials set for printed soil degradable electronic materials, particularly for conductors, enabling high conductivity and stability. Printed soil moisture and nitrate sensors with suitable sensitivity and selectivity were developed and characterized. Low power and passive wireless electronic systems were integrated with the soil sensors, and testing was carried out with completed sensors to understand their functionality under agricultural conditions. Additionally, other sensor types enabled by the biodegradable materials set created during this project, such as soil microbial activity sensors, were also developed and demonstrated. Project outputs include 10 peer reviewed publications, 4 patent applications, 21 technical presentations, 3 PhD thesis, 10 media reports, 8 additional grants worth over $\$$6M, and the formation of 3 start-up companies.
The purpose is to ensure that today’s processing codes produced output to meet today’s accuracy needs. Since 1958 ENDL and about 1965 ENDF have each used a text format to define nuclear and atomic data in 11 columns for each data field. When these formats originated this was judged to be adequate to reproduce the accuracy of data at the time and to meet the needs of our applications. When these formats originated the dominant computer language of the day was FORTRAN and if written using an E11.4 format it would include only 4 or 5 digits of precision, e.g., 0.1234E-03 or 1.2345E-02, varying from one computer/system to another the result was not even unique. In the case of ENDF the 4 digit precision was not even adequate to uniquely define the atomic weight of the target, e.g., U238 = 92238 = 0.9224E+5 = WRONG! From its inceptions the ENDF format had a precision problem. One of my first tasks when in 1967 fresh out of graduate school I joined what later became the National Nuclear Data Center (NNDC), was to address this precision problem. By working with ENDF producers and users throughout the U.S. we verified, 1) E or D is not required to define FORTRAN readable numbers, e.g., E+4 or +4 are both o.k. 2) With ENDF energy eV and cross section in barns, 2 digit exponents are almost never required. 3) Since energy is never negative we could use the first of the 11 columns for a digit. Knowing this allowed us to produce ENDF/B-II to 6 or 7 digit precision, e.g., blank, decimal point, 2 or 3 digit exponent, e.g., ^1.23456-12 or ^1.234567-3. Below is an example of the actual ENDF/B-II data released. Note, the date 1970 and the atomic weight, ZA, uniquely defined to 6-digit accuracy, ^9.22350+ 4.
The 2HDM+S is the singlet extension of the Two-Higgs-Doublets Model (2HDM). The singlet field and its mixing with the 2HDM Higgs sector lead to new contributions to the electroweak precision observables, in particular, the oblique parameters. In this paper, we performed a systematic study of the impacts of each mixing angle to the oblique parameters. We adopted the mixing angles and physical Higgs masses as our parameters, which allows a mapping when specific symmetry structure of the Higgs potential and various theoretical considerations are taken into account. We identify five benchmark cases, where at most one mixing angle is nonzero and analyze the 95\% C.L. allowed parameter space by the oblique parameters. In the alignment limit of the 2HDM, we find that other than the usual mass relations of $m_H\sim m_{H^\pm}$ or $m_A\sim m_{H^\pm}$, electroweak precision measurements also impose an upper limit on the neutral Higgs masses. In the cases with nonzero singlet mixing with the 2HDM Higgses $H$ or $A$, we find approximate mass relations of $c^2_{\alpha_{HS}} m_{H} + s^2_{\alpha_{HS}}m_{h_S} = m_{H^\pm}$ or $c^2_{\alpha_{AS}} m_{A} + s^2_{\alpha_{AS}}m_{A_S} = m_{H^\pm}$. Those relations are universal to the 2HDM+S models, with or without further symmetry assumption. We also study the non-alignment limit of the 2HDM+S, which typically has tighter constraints on the masses and mixing angles. At the end, we examine the complementarity between the electroweak precision analyses and the Higgs coupling precision measurements.
The PICOSEC Micromegas detector is a precise-timing gaseous detector based on a Cherenkov radiator coupled with a semi-transparent photocathode and a Micromegas amplifying structure, targeting a time resolution of tens of picoseconds for minimum ionising particles. Initial single-pad prototypes have demonstrated a time resolution below σ = 25 ps, prompting ongoing developments to adapt the concept for High Energy Physics applications, where sub-nanosecond precision is essential for event separation, improved track reconstruction and particle identification. The achieved performance is being transferred to robust multi-channel detector modules suitable for large-area detection systems requiring excellent timing precision. To enhance the robustness and stability of the PICOSEC Micromegas detector, research on robust carbon-based photocathodes, including Diamond-Like Carbon (DLC) and Boron Carbide (B 4 C), is pursued. Results from prototypes equipped with DLC and B 4 C photocathodes exhibited a time resolution of σ ≈ 32 ps and σ ≈ 34.5 ps, respectively. Efforts dedicated to improve detector robustness and stability enhance the feasibility of the PICOSEC Micromegas concept for large experiments, ensuring sustained performance while maintaining excellent timing precision.
High-performance optical interference coatings have transformed precision interferometry and spectroscopy by enabling unparalleled control over light–matter interactions. This review explores recent innovations in ion-beam sputtered amorphous dielectric, as well as substrate-transferred crystalline coatings, and their impact on systems at the forefront of precision metrology. These state-of-the-art coating techniques generate multilayers with ultralow optical losses, yielding mirrors with exceptional reflectivity. Refinements in their noise performance push the ultimate limits of sensitivity, resolution, and stability in demanding laser-based metrology applications. These technologies underpin the most advanced timekeeping and spatial measurement tools, enabling high-finesse reference cavities for the world’s most precise optical atomic clocks and low-noise reflective test masses for km-baseline gravitational-wave detectors. Emerging hybrid designs combining these techniques expand access to the mid-infrared spectral region, enabling the first ultralow-optical-loss coatings in the 3000–5000 nm wavelength range for enhanced spectroscopy and trace-gas detection. We highlight how these technologies redefine coating performance metrics and set new benchmarks in quantum science, fundamental physics, and precision optical sensing.
Future high-precision X-ray and gravitational-wave observations of neutron stars (NSs) are expected to constrain NS radii with uncertainties as small as σ ≃ 0.1 km. Such unprecedented precision offers a unique opportunity to extract new information about the nature and equation of state (EOS) of supradense matter in NS cores. Using mock radius data with uncertainties ranging from σ = 1.0 to 0.1 km, together with a flexible meta-model NS EOS that allows for a first-order hadron–quark phase transition, we perform a Bayesian statistical analysis to assess the impact of radius measurements on EOS constraints. We find that high-precision radius measurements, particularly for massive NSs, significantly tighten constraints on the hadron–quark transition density ρ t , the quark matter mass fraction in NS cores, and several parameters characterizing the EOS of supranuclear hadronic matter, although the degree of improvement depends on the assumed prior range of ρ t . In contrast, even with the highest precision considered, NS radii—including those of massive stars—remain largely insensitive to the stiffness of quark matter, independent of the measurement accuracy or the prior range adopted for ρ t .
Precise insertion of DNA sequences at targeted locations in plant genomes is pivotal for synthetic biology, genetics, and crop improvement. Construct design plays a critical role in achieving precise insertions, yet practical guidance remains limited. This review provides an in-depth overview of construct design principles and targeted DNA insertion (knock-in) strategies in plants. We assess the strengths, limitations, and construct requirements of current knock-in methods for specific applications, including short, large, and multifragment insertions. Additionally, we explore the potential of adopting advanced nonplant technologies to enhance knock-in efficiency and precision in plants. This review provides a valuable resource for facilitating the effective application of knock-in technologies to genetically improve crops with minimal off-target effects.
The ability to uniquely identify a compound requires highly precise and orthogonal measurements. Here we describe a newly developed analytical platform that integrates high resolution ion mobility and cryogenic vibrational ion spectroscopy for high-precision structural characterizations. This platform allows for the temporal separation of isomeric/isobaric ions and provides a highly sensitive description of the ion’s adopted geometry in the gas phase. The combination of these orthogonal structural measurements yields precise descriptors that can be used to resolve between and confidently identify highly similar ions. The unique benefits of our instrument, which integrates a structures for lossless ion manipulations ion mobility (SLIM IM) device with messenger tagging infrared spectroscopy, include the ability to perform high-resolution ion mobility separations and to record the IR spectra of all ions simultaneously. The SLIM IM device, with its 13 m separation path length, allows for multipass experiments to be performed for increased resolution as needed. It is integrated with an Agilent qTOF MS where the collision cell was replaced with a cryogenically held (30 K) TW-SLIM module. The cryo-SLIM is operated in a novel manner that allows ions to be streamed through the device and collisionally cooled to a temperature where they can form noncovalently bound N 2 complexes that are maintained as they exit the device and are detected by the TOF mass analyzer. The instrument can be operated in two modes: IMS+IR where the IR spectra for mobility-selected ions can be recorded and IR-only mode where the IR spectra for all mass-resolved ions can be recorded. In IR-only mode, IR spectra (400 cm –1 spectral range) can be recorded in as short as 2 s for high throughput measurements. Further, this work details the construction of the instrument and modes of operation. It provides initial benchmarking of CCS and IR measurements to demonstrate the utility of this instrument for targeted and untargeted approaches.
Here, we present a high-precision mass measurement of the proton-rich nucleus 23 Si, performed with the LEBIT Penning trap at the Facility for Rare Isotope Beams (FRIB) utilizing the time-of-flight ion cyclotron resonance (TOF-ICR) technique. We determined a mass excess of 23362.9(5.8) keV, which agrees with a recent storage-ring measurement from the experimental Cooler-Storage Ring (CSRe) in Lanzhou but has a factor of 20 improved precision 23 Si is hence the nucleus with the most precisely known mass among all nuclei with an isospin projection of 𝑇 𝑧 = −5/2. We performed shell-model calculations with the USDC and USDCm Hamiltonians to study binding energy differences and Thomas-Ehrmann shifts in mirror systems with an isospin up to 𝑇 = 5/2. Our experimental result and other recently reported masses of neutron-deficient sd-shell nuclei agree well with the theoretical predictions, demonstrating that isospin symmetry breaking in sd-shell nuclei—even at high isospin values—is well described by modern shell-model calculations.