Search NASA⌕ Search

SEARCH · Search NASA

Results for “Mixed precision”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes

Modern exascale GPU- and APU-based systems provide multiple power and energy sensors, but differences in scope, update rate, timing, and filtering complicate the attribution of short-lived accelerator activity. This paper presents a methodology to characterize and correct these effects on Cray EX systems with AMD Instinct MI250X GPUs (Frontier) and MI300A APUs (Portage). Using controlled square-wave workloads, we quantify update intervals, delay, aliasing, and variability across up to 512 GPUs and 480 APUs with on-chip (rocm-smi/amd-smi) and off-chip Cray Power Management sensors. We reconstruct power from cumulative energy counters to achieve faster response times, validate it against on-chip, off-chip, and node-level sensors, and integrate the resulting streams into a Score-P/PAPI-based tool for time-aligned, phase-level attribution. Applied to rocHPL, rocHPL-MxP, and HPG-MxP, the method separates energy savings due to reduced runtime from changes in power. Mixed precision reduces node energy on Frontier by 79% for rocHPL-MxP and 31% for HPG-MxP, with similar trends on Portage. These results provide portable guidance for sensor validation and power-aware optimization on current and future exascale systems.

Mcdaniel, Adam [ORNL] (ORCID:000000016926028X)↗

The First Swift Intensive AGN Accretion Disk Reverberation Mapping Survey

Swift intensive accretion disk reverberation mapping of four AGN yielded light curves sampled ∼200–350 times in 0.3–10 keV X-ray and six UV/optical bands. Uniform reduction and cross-correlation analysis of these data sets yields three main results: (1) The X-ray/UV correlations are much weaker than those within the UV/optical, posing severe problems for the lamp-post reprocessing model in which variations in a central X-ray corona drive and power those in the surrounding accretion disk. (2) The UV/optical interband lags are generally consistent with t μ l4 3 as predicted by the centrally illuminated thin accretion disk model. While the average interband lags are somewhat larger than predicted, these results alone are not inconsistent with the thin disk model given the large systematic uncertainties involved. (3) The one exception is the U band lags, which are on average a factor of ∼2.2 larger than predicted from the surrounding band data and fits. This excess appears to be due to diffuse continuum emission from the broad-line region (BLR). The precise mixing of disk and BLR components cannot be determined from these data alone. The lags in different AGN appear to scale with mass or luminosity. We also find that there are systematic differences between the uncertainties derived by JAVELIN versus more standard lag measurement techniques, with JAVELIN reporting smaller uncertainties by a factor of 2.5 on average. In order to be conservative only standard techniques were used in the analyses reported herein.

R. Edelson↗

Performance Optimization Methods for a Memory-Bound, Unstructured-Grid CFD Application on Massively Parallel GPU Platforms

Computational performance of the FUN3D unstructured-grid computational fluid dynamics (CFD) application on massively parallel GPU environments is memory-bound and highly dependent upon efficient reads from and atomic updates to the irregular cell-, edge-, and node-based data structures. In this talk, we present recent efforts into optimizing select performance-critical kernels on NVIDIA Tesla V100 and A100 GPUs and AMD CDNA MI100 GPUs. A novel use of L2 cache residency controls and asynchronous loads into on-chip shared memory are explored on the A100 GPU for the sparse iterative solver, which is dominated by mixed-precision, sparse matrix vector multiplication. Demonstrations show that these methods improve global memory bandwidth utilization by 13.5% on the A100 GPU. Several techniques are also presented that use registers and/or shared memory to facilitate array transposition and aggregation which combine to reduce the frequency and increase the cache efficiency of floating-point atomic updates to the irregular data structures. These methods are demonstrated to improve the kernel throughput by nearly 500% on select kernels on the AMD MI100 over atomic updates directly to global memory. Overall, both V100 and A100 GPUs outperformed the MI100 GPU on kernels dominated by double-precision atomic updates; however, the techniques demonstrated here reduced the performance gap and improved the MI100 performance.

GPU CPU unstructured CFD memory↗

Results of Propellant Mixing Variable Study Using Precise Pressure-Based Burn Rate Calculations

A designed experiment was conducted in which three mix processing variables (pre-curative addition mix temperature, pre-curative addition mixing time, and mixer speed) were varied to estimate their effects on within-mix propellant burn rate variability. The chosen discriminator for the experiment was the 2-inch diameter by 4-inch long (2x4) Center-Perforated (CP) ballistic evaluation motor. Motor nozzle throat diameters were sized to produce a common targeted chamber pressure. Initial data analysis did not show a statistically significant effect. Because propellant burn rate must be directly related to chamber pressure, a method was developed that showed statistically significant effects on chamber pressure (either maximum or average) by adjustments to the process settings. Burn rates were calculated from chamber pressures and these were then normalized to a common pressure for comparative purposes. The pressure-based method of burn rate determination showed significant reduction in error when compared to results obtained from the Brooks' modification of the propellant web-bisector burn rate determination method. Analysis of effects using burn rates calculated by the pressure-based method showed a significant correlation of within-mix burn rate dispersion to mixing duration and the quadratic of mixing duration. The findings were confirmed in a series of mixes that examined the effects of mixing time on burn rate variation, which yielded the same results.

Stefanski, Philip L.↗

Energy-efficient, Large-scale Molecular Dynamics Simulations via Hardware- and Algorithm-level Optimization

This work aims to develop a framework for energy-efficient computing that will enable molecular dynamics (MD) simulations of large-scale phenomena with atomic precision and simultaneously remove computational bottlenecks limiting the speed of MD simulations. We seek to implement such an approach through the development of surrogate models for the interatomic force calculation combined with the use of mixed numerical precision formats. For a model system of neutral atoms (only pairwise interactions), significant force calculation efficiency improvements were achieved, without detrimental effects on atomic structures or average energies, using single precision, by developing a surrogate model (deep neural network), and by quantizing this surrogate model. For a model system of charged atoms, the reciprocal-space calculation of electrostatic interactions was identified as the main bottleneck, and the development of a surrogate model should be pursued to achieve an estimated one-order-of-magnitude additional speedup.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Studies of long-life pulsed CO2 laser with Pt/SnO2 catalyst

Closed-cycle CO2 laser testing with and without a catalyst and with and without CO addition indicate that a catalyst is necessary for long-term operation. Initial results indicate that CO addition with a catalyst may prove optimal, but a precise gas mix has not yet been determined. A long-term run of 10 to the 6th power pulses using 1.3% added CO and a 2% Pt on SnO2 catalyst yields an efficiency of about 95% of open-cycle steady-state power. A simple mathematical analysis yields results which may be sufficient for determining optimum running conditions. Future plans call for testing various catalysts in the laser and longer tests, 10 to the 7th power pulses. A Gas Chromatograph will be installed to measure gas species concentration and the analysis will be slightly modified to include neglected but possibly important parameters.

Sidney, Barry D.↗

GEER Status Update

History: The Glenn Extreme Environments Rig (GEER) first became operational in the early part of 2015. Since that time GEER has completed a number of scientific tests and has undergone improvements in the chemical delivery system and analytics following a year of operations experience. Recent Updates: In June 2016, the GEER process system was rebuilt to provide a more robust system, higher accuracy and new capabilities. New insulation was installed on the exterior of the pressure vessel and gas lines. The newly revamped GEER plumbing system can provide extremely precise custom gas mixtures using any gas desired by the investigator in any combination. GEER can heat the resulting mixture up to 500 deg C and 1500 psia. The 304 stainless steel vessel walls were polished to reduce corrosion rate and reduce unwanted chemical reactions. The process lines were replaced with high purity Sulfinert coated tubing. The GEER team added the ability to individually boost specialty gases to GEER thus allowing operators to make very precise changes to the gas chemistry inside of GEER during a test while at high temperature and pressure. High accuracy mass flow meters were added to further improve gas mixing accuracy and precision. An in-line, integrated Inficon MicroGC Fusion was added for real time gas analysis along with a high purity gas sampling system, providing fully automated, real time analysis of the gas chemistry inside of GEER in minutes. This complements a co-located mass spectrometer and both are used for regular monitoring of the vessel chemistry. All internal vessel components were replaced with polished 304SS equivalents. Hot vent down capability was increased. Finally, an automated liquid injection system was added and is rated for max vessel operating conditions of (1500 psia, 500 C). Recent Results and Publications: In May 2016, GEER completed a test that exposed high temperature electronics to Venus surface conditions for 21.5 days. This demonstrated the potential for operating robotic spacecraft in the Venus environment without the need for thermal or environmental protection. Results from this test were published in December 2016 and received national media attention. In April 2017 GEER implemented an 80 day test at Venus surface conditions to simulate chemical weathering of expected Venus minerals. This test supported a ROSES award to a team led by Prof. Ralph Harvey of Case Western Reserve University. The test concluded in July 2017 and nearly doubled previous operation record of 42 days at Venus surface conditions. Preliminary results of these and previous experiments were presented at the recent Venus Modeling Workshop. In June 2017, NASA TM2017-219437 "Chemical and Microstructural Changes in Metallic and Ceramic Materials Exposed to Venusian Surface Conditions" was published. This report provides an extensive and valuable resource detailing the behavior of a variety of engineering materials at Venus surface conditions. Community Involvement: An external science advisory panel has been formed.

Kremic, Tibor↗

Quick mixing of epoxy components

Two materials are mixed quickly, thoroughly, and in precise proportion by disposable cartridge. Cartridge mixes components of fast-curing epoxy resins, with no mess, just before they are used. It could also be used in industry and home for caulking, sealing, and patching. Materials to be mixed are initially isolated by cylinder wall within cartridge. Cylinder has vanes, with holes in them, at one end and handle at opposite end. When handle is pulled, grooves on shaft rotate cylinder so that vanes rotate to extrude material A uniformly into material B.

Dunlap, D. E., Jr.↗

Atmospheric CO2 Column Measurements with an Airborne Intensity-Modulated Continuous-Wave 1.57-micron Fiber Laser Lidar

The 2007 National Research Council (NRC) Decadal Survey on Earth Science and Applications from Space recommended Active Sensing of CO2 Emissions over Nights, Days, and Seasons (ASCENDS) as a mid-term, Tier II, NASA space mission. ITT Exelis, formerly ITT Corp., and NASA Langley Research Center have been working together since 2004 to develop and demonstrate a prototype Laser Absorption Spectrometer for making high-precision, column CO2 mixing ratio measurements needed for the ASCENDS mission. This instrument, called the Multifunctional Fiber Laser Lidar (MFLL), operates in an intensity-modulated, continuous-wave mode in the 1.57- micron CO2 absorption band. Flight experiments have been conducted with the MFLL on a Lear-25, UC-12, and DC-8 aircraft over a variety of different surfaces and under a wide range of atmospheric conditions. Very high-precision CO2 column measurements resulting from high signal-to-noise (great than 1300) column optical depth measurements for a 10-s (approximately 1 km) averaging interval have been achieved. In situ measurements of atmospheric CO2 profiles were used to derive the expected CO2 column values, and when compared to the MFLL measurements over desert and vegetated surfaces, the MFLL measurements were found to agree with the in situ-derived CO2 columns to within an average of 0.17% or approximately 0.65 ppmv with a standard deviation of 0.44% or approximately 1.7 ppmv. Initial results demonstrating ranging capability using a swept modulation technique are also presented.

Dobler, Jeremy T.↗

Efficient Quantum Gibbs Samplers with Kubo–Martin–Schwinger Detailed Balance Condition

Lindblad dynamics and other open-system dynamics provide a promising path towards efficient Gibbs sampling on quantum computers. In these proposals, the Lindbladian is obtained via an algorithmic construction akin to designing an artificial thermostat in classical Monte Carlo or molecular dynamics methods, rather than being treated as an approximation to weakly coupled system-bath unitary dynamics. Recently, Chen, Kastoryano, and Gilyén (arXiv:2311.09207) introduced the first efficiently implementable Lindbladian satisfying the Kubo–Martin–Schwinger (KMS) detailed balance condition, which ensures that the Gibbs state is a fixed point of the dynamics and is applicable to non-commuting Hamiltonians. This Gibbs sampler uses a continuously parameterized set of jump operators, and the energy resolution required for implementing each jump operator depends only logarithmically on the precision and the mixing time. In this work, we build upon the structural characterization of KMS detailed balanced Lindbladians by Fagnola and Umanità, and develop a family of efficient quantum Gibbs samplers using a finite set of jump operators (the number can be as few as one), akin to the classical Markov chain-based sampling algorithm. Compared to the existing works, our quantum Gibbs samplers have a comparable quantum simulation cost but with greater design flexibility and a much simpler implementation and error analysis. Moreover, it encompasses the construction of Chen, Kastoryano, and Gilyén as a special instance.

97 MATHEMATICS AND COMPUTING↗

Participation in Intensity Frontier Neutrino Physics (closeout report)

The flagship currently-running experiment at Fermilab is the NuMI Off-axis νe Appearance (NOvA) experiment. Together with the (complementary) T2K experiment in Japan, it provides the opportunity for improved measurement precision on neutrino mixing parameters and addresses topics of strong interest: whether neutrino masses follow a normal (NH) or inverted (IH) hierarchy; whether the ν 3 state contains a symmetric mixture of ν µ and ν τ (“maximal mixing,” pointing to a possible new symmetry of nature), or, if not, what is the octant of θ 23 ; and whether neutrino mixing violates CP symmetry. Recent results from both experiments suggest we may be on the verge of important discoveries. Towards the end of the decade, these experiments will be superseded by DUNE in the U.S. and HyperK in Japan.

2x2↗

Addressing Low-Cost Methane Sensor Calibration Shortcomings with Machine Learning

Quantifying methane emissions is essential for meeting near-term climate goals and is typically carried out using methane concentrations measured downwind of the source. One major source of methane that is important to observe and promptly remediate is fugitive emissions from oil and gas production sites but installing methane sensors at the thousands of sites within a production basin is expensive. In recent years, relatively inexpensive metal oxide sensors have been used to measure methane concentrations at production sites. Current methods used to calibrate metal oxide sensors have been shown to have significant shortcomings, resulting in limited confidence in methane concentrations generated by these sensors. To address this, we investigate using machine learning (ML) to generate a model that converts metal oxide sensor output to methane mixing ratios. To generate test data, two metal oxide sensors, TGS2600 and TGS2611, were collocated with a trace methane analyzer downwind of controlled methane releases. Over the duration of the measurements, the trace gas analyzer’s average methane mixing ratio was 2.40 ppm with a maximum of 147.6 ppm. The average calculated methane mixing ratios for the TGS2600 and TGS2611 using the ML algorithm were 2.42 ppm and 2.40 ppm, with maximum values of 117.5 ppm and 106.3 ppm, respectively. A comparison of histograms generated using the analyzer and metal oxide sensors mixing ratios shows overlap coefficients of 0.95 and 0.94 for the TGS2600 and TGS2611, respectively. Overall, our results showed there was a good agreement between the ML-derived metal oxide sensors’ mixing ratios and those generated using the more accurate trace gas analyzer. This suggests that the response of lower-cost sensors calibrated using ML could be used to generate mixing ratios with precision and accuracy comparable to higher priced trace methane analyzers. This would improve confidence in low-cost sensors’ response, reduce the cost of sensor deployment, and allow for timely and accurate tracking of methane emissions.

03 NATURAL GAS↗

The Modular Combustion Facility for the Space Station Laboratory - A Requirements and Capabilities Study

This paper describes a modular combustion facility for the Space Station, designed to provide facility-level services to interchangeable experiment modules, each of which designed specifically for the needs of a particular combustion experiment. The facility-level services are to include computer devices for the data acquisition, experimental control, and data reduction and analysis; the electrical power conversion and control; video cameras and recordings; the cooling-loop supply; waste management; gas supply; precision gas-mixing; and the combustion diagnostics support. Summarized categories of the data base are provided, which were developed to assimilate and to process the responses from the investigators.

Sacksteder, K. R.↗

Algorithms and Libraries

This exploratory study initiated our inquiry into algorithms and applications that would benefit by latency tolerant approach to algorithm building, including the construction of new algorithms where appropriate. In a multithreaded execution, when a processor reaches a point where remote memory access is necessary, the request is sent out on the network and a context--switch occurs to a new thread of computation. This effectively masks a long and unpredictable latency due to remote loads, thereby providing tolerance to remote access latency. We began to develop standards to profile various algorithm and application parameters, such as the degree of parallelism, granularity, precision, instruction set mix, interprocessor communication, latency etc. These tools will continue to develop and evolve as the Information Power Grid environment matures. To provide a richer context for this research, the project also focused on issues of fault-tolerance and computation migration of numerical algorithms and software. During the initial phase we tried to increase our understanding of the bottlenecks in single processor performance. Our work began by developing an approach for the automatic generation and optimization of numerical software for processors with deep memory hierarchies and pipelined functional units. Based on the results we achieved in this study we are planning to study other architectures of interest, including development of cost models, and developing code generators appropriate to these architectures.

Dongarra, Jack↗

Nd-142/Nd-144 in bulk planetary reservoirs, the problem of incomplete mixing of interstellar components and significance of very high precision Nd-145/Nd-144 measurements

Apart from the challenge of very high precision Nd-142/Nd-144 ratio measurement, accurate applications of the coupled Sm-(146,147)-Nd-(142,143) systematics in planetary differentiation studies require very precise knowledge of the present-day (post-Sm-146 decay) Nd-142/Nd-144 ratios of bulk planetary objects (BP). The coupled systematics yield model ages for the time of formation of Sm/Nd-fractionated reservoirs by differentiation of Sm/Nd-unfractionated bulk planetary reservoirs. Estimates of (Nd-142/Nd-144)(sub BP) and (Nd-143/Nd-144)(sub BP) therefore provide the critical baseline relative to which these model ages are referenced. In the Sm-147-Nd-143 systematics, Nd-143/Nd-144 variations are mostly large; therefore, small variations in initial Nd-143/Nd-144 ratios generally can be ignored. However, in the case of Sm-146-Nd-142, the range of Nd-142/Nd-144 divergence for differentiated planetary reservoirs is much smaller. Consequently Sm-(146,147)-Nd-(142,143) model ages are sensitive to small variations in bulk planetary Nd-142/Nd-144 (both present-day and initial). One major unanswered question is whether or not Nd shelf standards (CIT Nd beta/Ames metal, La Jolla, NASA-JSC/Ames metal) have Nd-142/Nd-144 identical to the bulk Earth or otherwise might record some degree of radiogenic evolution in an early-fractionated reservoir. Our discussions of earth Earth differentiation based on Nd-142/Nd-144 in Isua and Acasta samples have employed a working assumption: (Nd-142/Nd-144)(sub Nd beta) = (Nd-142/Nd-144)(sub Bulk Earth). This requires experimental justification and is apparently contradicted by chondrite Nd-142/Nd-144 measurements, which have been interpreted to indicate: (Nd-142/Nd-144)(sub JSC/Ames metal) = ((Nd-142/Nd-144)(sub CHUR) = 35 plus or minus 8 ppm). At present, interpretations of the early Earth and Moon hinge largely on this issue. Because Ba in bulk chondrite samples exhibit similar magnitude nuclear anomalies, attributable to incomplete mixing of interstellar components, a critical question is whether or not nuclear effects are also present in Nd-142/Nd-144, both in bulk chondrites and between planetary objects.

Harper, C. L., Jr.↗

Simulations, Modeling and Data Analysis of Parity Violating Electron Scattering Experiments

In the Standard Model (SM) of nuclear and particle physics, parity violation is incorporated through the representation of the weak interaction as a chiral gauge interaction. Only the left-handed components of particles and right-handed components of antiparticles participate in weak interactions in the Standard Model. This implies that parity is asymmetric for the weak interaction. Parity violating electron scattering (PVES) experiments are designed to probe the physics parameters related to the SM, with the possibility to discover physics beyond the SM (BSM) by measuring the parity violating asymmetry ¿¿¿ of longitudinally polarized electrons scattered off unpolarized targets with high precision. This dissertation will be focused on two PVES experiments, the next 208Pb Lead Radius Experiment (PREX-II), and the Measurement of a Lepton-Lepton Electroweak Reaction (MOLLER) experiment, as well as in some small sections, the Calcium Radius Experiment (CREX) and P2 experiment which are also PVES experiments). PREX-II and CREX experiments, performed in Hall A at the Thomas Jefferson National Accelerator Facility (Jefferson Lab), measured ¿¿¿ in the elastic scattering of longitudinally polarized electrons from 208Pb and 48Ca targets to provide a precise model independent determination of the neutron skin thickness of 208Pb and 48Ca nuclei, respectively. The MOLLER experiment, proposed to start in 2027 and also to be performed in Hall A at Jefferson Lab, is to measure ¿¿¿ of longitudinally polarized electrons scattered off unpolarized electrons (Møller scattering) to determine the weak charge of electrons ¿¿¿ and the weak mixing angle ¿¿ with high precision. As for the P2 experiment, which will be performed at the upcoming MESA accelerator in Mainz Germany, it is to measure the weak charge of proton ¿¿¿ using ¿¿¿ in the elastic electron-proton scattering of polarized electrons off unpolarized protons. The final results from the PREX-II experiment are presented as ¿¿¿=550±16 (¿¿¿¿)±8 (¿¿¿¿) parts-per-billion (ppb). Combining the PREX-I and PREX-II results, the neutron skin thickness from PREX experiments is determined as ¿¿-¿¿=0.283±0.071 ¿¿ in 208Pb. This thesis lists the software and computational contribution of the author to these PVES experiments, including writing scripts and software to help with the PREX-II/CREX experiments, analyzing data to provide useful information and systematic uncertainty for the PREX-II experiment, modeling and simulations for the MOLLER, and providing an alternative design of an electronic equipment for the P2 experiment.

Chen, Yufan↗

Simulations, Modeling and Data Analysis of Parity Violating Electron Scattering Experiments

In the Standard Model (SM) of nuclear and particle physics, parity violation is incorporated through the representation of the weak interaction as a chiral gauge interaction. Only the left-handed components of particles and right-handed components of antiparticles participate in weak interactions in the Standard Model. This implies that parity is asymmetric for the weak interaction. Parity violating electron scattering (PVES) experiments are designed to probe the physics parameters related to the SM, with the possibility to discover physics beyond the SM (BSM) by measuring the parity violating asymmetry ¿¿¿ of longitudinally polarized electrons scattered off unpolarized targets with high precision. This dissertation will be focused on two PVES experiments, the next 208Pb Lead Radius Experiment (PREX-II), and the Measurement of a Lepton-Lepton Electroweak Reaction (MOLLER) experiment, as well as in some small sections, the Calcium Radius Experiment (CREX) and P2 experiment which are also PVES experiments). PREX-II and CREX experiments, performed in Hall A at the Thomas Jefferson National Accelerator Facility (Jefferson Lab), measured ¿¿¿ in the elastic scattering of longitudinally polarized electrons from 208Pb and 48Ca targets to provide a precise model independent determination of the neutron skin thickness of 208Pb and 48Ca nuclei, respectively. The MOLLER experiment, proposed to start in 2027 and also to be performed in Hall A at Jefferson Lab, is to measure ¿¿¿ of longitudinally polarized electrons scattered off unpolarized electrons (Møller scattering) to determine the weak charge of electrons ¿¿¿ and the weak mixing angle ¿¿ with high precision. As for the P2 experiment, which will be performed at the upcoming MESA accelerator in Mainz Germany, it is to measure the weak charge of proton ¿¿¿ using ¿¿¿ in the elastic electron-proton scattering of polarized electrons off unpolarized protons. The final results from the PREX-II experiment are presented as ¿¿¿=550±16 (¿¿¿¿)±8 (¿¿¿¿) parts-per-billion (ppb). Combining the PREX-I and PREX-II results, the neutron skin thickness from PREX experiments is determined as ¿¿-¿¿=0.283±0.071 ¿¿ in 208Pb. This thesis lists the software and computational contribution of the author to these PVES experiments, including writing scripts and software to help with the PREX-II/CREX experiments, analyzing data to provide useful information and systematic uncertainty for the PREX-II experiment, modeling and simulations for the MOLLER, and providing an alternative design of an electronic equipment for the P2 experiment.

Chen, Yufan↗