Search NASA⌕ Search

SEARCH · Search NASA

Results for “floating”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Breaking the Million-Electron and 1 EFLOP/s Barriers: Biomolecular-Scale Ab Initio Molecular Dynamics Using MP2 Potentials

The accurate simulation of complex biochemical phenomena has historically been hampered by the computational requirements of high-fidelity molecular-modeling techniques. Quantum mechanical methods, such as ab initio wave-function (WF) theory, deliver the desired accuracy, but have impractical scaling for modeling biosystems with thousands of atoms. Combining molecular fragmentation with MP2 perturbation theory, this study presents an innovative approach that enables biomolecular-scale ab initio molecular dynamics (AIMD) simulations at WF theory level. Leveraging the resolution-of-the-identity approximation for Hartree-Fock and MP2 gradients, our approach eliminates computationally intensive four-center integrals and their gradients, while achieving near-peak performance on modern GPU architectures. The introduction of asynchronous time steps minimizes time step latency, overlapping computational phases and effectively mitigating load imbalances. Utilizing up to 9,400 nodes of Frontier and achieving 59% (1006.7 PFLOP/s) of its double-precision floating-point peak, our method enables us to break the million-electron and 1EFLOP/s barriers for AIMD simulations with quantum accuracy.

Kurzak, Jakub↗

Shifting Between Compute and Memory Bounds: A Compression-Enabled Roofline Model

In the evolving landscape of high-performance computing, especially to fight the end of Moore’s Law and Dennard’s Scaling, the ability to shift between compute-bound and memory-bound states is critical for enhancing adaptability and flexibility to diverse system and domain-specific architectures. Such capability is vital for optimizing performance across distinguished hardware configurations, such as accelerators, memory hierarchies, and cache systems. Despite that ad hoc optimization techniques, such as compressed/approximate computation, have been enabled for compute-/data-intensive computing for improved performance in distinct hardware settings, there lacks an understanding of 1) the rational behind performance improvement; 2) capability of different optimizations; 3) what optimization to respond to specific computational and memory demands. This work proposes a compression-enabled roofline model to facilitate this adaptability with data compression techniques to balance and transform between computational and memory demands. This model enables applications to adjust in response to the specific strengths and limitations of the underlying hardware and system to optimize resource utilization. The effectiveness of this approach is demonstrated with matrix multiplication kernels on different input sizes, with turning on/off various compression techniques, including 1) low-precision floating point; 2) sparse matrix formulation; and 3) compressed arrays with ZFP. By reducing memory transfer volumes and cache misses and increasing data locality and computational intensity through compression, the specific roofline model can transform between compute and memory bounds to align more efficiently with system capabilities. This advancement not only improves overall performance but also maximizes adaptability in diverse computing environments.

Naraparaju, Ramasoumya [University of Washington]↗

An Efficient Checkpointing System for Large Machine Learning Model Training

As machine learning models increase in size and complexity rapidly, the cost of checkpointing in ML training became a bottleneck in storage and performance (time). For example, the latest GPT-4 model has massive parameters at the scale of 1.76 trillion. It is highly time and storage consuming to frequently writes the model to checkpoints with more than 1 trillion floating point values to storage. This work aims to understand and attempt to mitigate this problem. First, we characterize the checkpointing interface in a collection of representative large machine learning/language models with respect to storage consumption and performance overhead. Second, we propose the two optimizations: i) A periodic cleaning strategy that periodically cleans up outdated checkpoints to reduce the storage burden; ii) A data staging optimization that coordinates checkpoints between local and shared file systems for performance improvement.

machine learning, artificial intelligence↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Rising Water Levels and Vegetation Shifts Drive Substantial Reductions in Methane Emissions and Carbon Dioxide Uptake in a Great Lakes Coastal Freshwater Wetland

ABSTRACT Coastal freshwater wetlands are critical ecosystems for both local and global carbon cycles, sequestering substantial carbon while also emitting methane (CH 4 ) due to anoxic conditions. Estuarine freshwater wetlands face unique challenges from fluctuating water levels, which influence water quality, vegetation, and carbon cycling. However, the response of CH 4 fluxes and their drivers to altered hydrology and vegetation remains unclear, hindering mechanistic modeling. To address these knowledge gaps, we studied an estuarine freshwater wetland in the Great Lakes region, where rising water levels led to a vegetation shift from emergent Typha dominance in 2015–2016 to floating‐leaved species in 2020–2022. Using eddy covariance flux measurements during the peak growing season (June–September) of both periods, we observed a 60% decrease in CH 4 emissions, from 81 ± 4 g C m −2 in 2015–2016 to 31 ± 3 g C m −2 in 2020–2022. This decline was driven by two main factors: (1) higher water levels, which suppressed ebullitive fluxes via increased hydrostatic pressure and extended CH 4 residence time, enhancing oxidation potential in the water column; and (2) reduced CH 4 conductance through plants. Net carbon dioxide (CO 2 ) uptake decreased by 90%, from −267 ± 26 g C m −2 in 2015–2016 to −27 ± 49 g C m −2 in 2020–2022. Additionally, diel CH 4 flux patterns shifted, with a distinct morning peak observed in 2015–2016 but absent in 2020–2022, suggesting changes in plant‐mediated transport and a potential decoupling from photosynthesis. The dominant factors influencing CH 4 fluxes shifted from water temperature and gross primary productivity in 2015–2016 to atmospheric pressure in 2020–2022, suggesting an increased role of ebullition as a primary transport pathway. Our results demonstrate that changes in water levels and vegetation can substantially alter CH 4 and CO 2 fluxes in coastal freshwater wetlands, underscoring the critical role of hydrological shifts in driving carbon dynamics in these ecosystems.

54 ENVIRONMENTAL SCIENCES↗

Thermo-rheological snapshot of melter feed conversion to glass

Slurry feed charged into an electric melter creates a layer of reacting and melting material (termed cold cap) that floats on the surface of molten glass. The rheological behavior of heated melter feed affects the spreading of slurry at the top of the cold cap and the stability of the primary foam, affecting cold-cap coverage and melter plenum temperatures. The apparent viscosity of a high-alumina high-level waste melter feed was assessed by thermomechanical analysis, high-temperature viscometer, and the hot stage microscopy method, yielding viscosity estimates from ≈10 7.5 Pa s at 550°C to ≈10 2.5 Pa s at 1050°C. As the temperature of feed materials increased, their state changed from rigid solid to dilatant fluid, to pseudoplastic bubbly liquid with dissolving solids, to fully developed foam, and finally to Newtonian glass melt.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Static metrology of the meter-scale deformable heliostat

This paper describes a heliostat metrology system which is developed based on deflectometry, utilizing a static perforated panel instead of conventional dynamic monitor displays to provide incident rays. The developed method is named static screen deflectometry (SSD). This robust and scalable method is especially valuable for outdoor tests of large reflectors used in concentrating solar-thermal power (CSP) systems. The developed method has been successfully demonstrated on a 2.4 𝑚 × 3.3 𝑚 float glass deformable reflector bent to focus sunlight at 113 𝑚 distance throughout a day. From images obtained from a camera at 50𝑚 distance, the reflector surface was measured to an accuracy of less than 1 𝑚𝑟𝑎𝑑 rms slope error in the full test scope.

14 SOLAR ENERGY↗

Topographic effects on reflected acoustic waves from the OSIRIS-REx reentry observed from stratospheric balloons

During long-distance sound propagation in planetary atmospheres, acoustic waves may reflect off the air/surface interface one or more times. For low sound frequencies and flat interfaces, the incident and reflected wave tend to be nearly identical. However, this may not be the case when the downgoing acoustic wave encounters topography. Here, we describe a set of direct and reflected acoustic signals recorded on free-flying balloons during the hypersonic entry of the OSIRIS-REx sample return capsule (SRC). In two of the three cases presented here, an impulsive reflected arrival similar in form to the direct sonic boom of the SRC is observed, followed by a diffuse coda. In contrast, one of the floating stations lacked an impulsive reflection entirely, with only the coda present. We use full waveform modeling to show how reflection in the presence of complex topography can explain coda seen in all three examples as well as the lack of impulsive arrival on the third. Our results indicate that the complex signals often observed in long range acoustic propagation could be due, in part, to interactions with topography during transmission.

Lees, Jonathan M. [University of North Carolina, C↗

Deploying and Tracking Software with NCCS Software Provisioning

The National Center for Computational Sciences (NCCS) at Oak Ridge National Laboratory has a long history of deploying ground-breaking leadership-class supercomputers for the U.S. Department of Energy. The latest in this line of supercomputers is Frontier, the first supercomputer to break the exascale barrier (1018 floating-point operations per second) on the TOP500 list. Frontier serves a wide array of scientific domains, from traditional simulation-based workloads to newer AI and Machine Learning workloads. To best serve the NCCS user community, NCCS uses Spack to deploy a comprehensive software stack of scientific software packages, providing straightforward access to these packages through Lmod Environment Modules. Maintaining a large software stack while also including multiple new compiler releases each year is a very time-consuming task. Additionally, it is not straightforward to provide a software stack alongside existing vendor-provided software such as the HPE/Cray Programming Environment (CPE), and existing CPE, Spack, and Lmod integration does not allow for multiple versions of GPU libraries such as AMD’s ROCm to be used. To address these challenges and shortcomings, NCCS has developed the NCCS Software Provisioning tool (NSP)1, a tool for deploying and monitoring software stacks on HPC systems. NSP allows NCCS to quickly and effectively provision software stacks from the ground up using template-driven recipes and configuration files. NSP is successfully deployed on Frontier and several other NCCS clusters, enabling the NCCS software team to quickly deploy software stacks for newly-released compilers, expand current software offerings, better support GPU-based software, and monitor Lmod module usage to identify unused software packages that can be removed from the software stack. In this work, we discuss the shortcomings of the previous CPE, Spack, and Lmod usage at NCCS, provide further details on the implementation and structure of NSP, then discuss the benefits that NSP provides.

Rentschler, Asa [ORNL] (ORCID:0009000597694743)↗

Generic Multi-Layer Perceptron Inference Accelerator on FPGA (vneuron) v1.0

We have designed and implemented a neural network inference compute engine (vneuron) that can be deployed in the fabric of any FPGA without using special hardware accelerator primitive. The "vneuron" is purely written in verilog, and supports scalable neural network structure with fully connected layers and ReLU activation ( Multi-Layer Perceptron architecture) with 16 bits of precision. We have demonstrated it on an Xilinx Artix 7 FPGA for a 16-input, 8-output MLP with 3 layer, 1600 parameters. It takes 40 DSP48E and 40 BRAM18, and takes 131 clock cycles for computing (1048 ns when clocked at 125MHz). We include PyTorch quantization from a given floating point model, and provide behavioral verification simulation in the disclosed software package.

Du, Qiang↗

Offshore Wind ENergy Simulation Toolkit (OWENS)

SAND2021-2751 O The Offshore Wind ENergy Simulation Toolkit (OWENS) is a collection of aerodynamic, structural, hydrodynamic, drivetrain, controls, composite structure and mesh preprocessing, and data postprocessing. OWENS is primarily an ontology, or glue code, pulling together many open-source and Sandia-developed libraries to model the aero-servo-hydro-elastic physics of wind and marine energy turbines. The toolkit’s intended use is for arbitrary aeroelastic rotor configurations analysis including vertical-axis wind turbines, horizontal-axis wind turbines, and analogous marine energy applications for fixed-bottom and floating configurations. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525

Owens, Brian↗

Virtreg

Virtual register for performing addition and subtraction of floating-point numbers without loss of precision

Hakel, Peter↗

QCUncertainty/sigma

Sigma is a header-only C++ library for uncertainty propagation throughout mathematical operations on floating point values.

Waldrop, Jonathan M.↗

CODARcode/MGARD

MGARD is a software providing error-controlled lossy compression and data refactoring based on multi-grid theories. It transforms floating-point scientific data into a multilevel representation, followed by quantization and lossless encoding processes, resulting in a self-describing compressed buffer. It supports diverse data topologies, error control norms, and computing architectures.

Chen, Jieyang [University of Oregon]↗

ROSE Castor

ROSE Castor is a tool enabling automated verification of C++, built off of the ROSE compiler framework and the Why3 framework. Castor defines a verification language for providing specifications of C++ code, letting users perform automated functional formal verification of their C++ code. Castor is designed to target C++17, and supports a subset of the language, including classes, functions, templates, integers and booleans, pointers and references, and single inheritance. Castor currently does not support multiple or virtual inheritance, virtual functions, floating-point, threading, lambda functions, or the C++ STL, though some of these are planned in future updates. Castor ships with an in-house parser for parsing verification conditions.

Lane, PhillipA [Lawrence Livermore National Labora↗

SIRENOpt.jl

SAND2026-22945O SIRENOpt.jl is a Julia software package for prototype hybrid power, storage, and platform dynamics. It integrates solar, wind, wave, hydrokinetic, diesel, generator, converter, battery, hydrogen, desalination, mooring, and floating-platform model interfaces in an automatic-differentiation-friendly simulation framework. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Michelen Strofer, Carlos [Sandia National Lab. (SN↗

Definition of the IEA Wind 22-Megawatt Offshore Reference Wind Turbine

This technical report describes the design of a new 22 Megawatt reference wind turbine (RWT). The turbine model was designed collaboratively by two teams at the Denmark Technical University and at the National Renewable Energy Laboratory within the International Energy Agency (IEA) Wind Technology Commercialization Programme (TCP) Task 55 on Reference Wind Turbines and Farms. RWTs serve an important purpose in the wind energy community, since they provide openly available data for models representative of current wind turbine technology, which can be used by practitioners for a variety of modeling purposes, ranging from aerodynamic, structural, and aeroelastic turbine modeling to wind farm flow modeling, across a range of fidelities. The IEA 22 RWT aims to model machines with projected installation in the 2025-2030 time frame. The turbine has a rotor diameter of 284 meters and a hub height of 170 meters. It is a class 1-B machine with a rotor specific power nearing 350 W m -2 and it is mounted on either a fixed-bottom offshore foundation or a semi-submersible floating platform.

17 WIND ENERGY↗

Compatibilization Strategy and Mechanism for Co-stabilizing Commingled Plastics and Pyrolyzed Rubber in Asphalt

Hot mix asphalt mixture is considered the ideal approach to reuse waste plastics in high-value applications because of its very high amount of usage in highway construction. However, the differences in polarity and density between polymers and asphalt lead to polymer coalescence and therefore the poor storage stability of modified asphalt. These challenges are exalted when recycling commingled plastics. This study introduced an innovative compatibilization strategy and mechanism for co-stabilizing commingled plastics and pyrolyzed rubber in asphalt. Commingled plastics were first grafted with maleic anhydride for surface activation, followed by reactive kneading with pyrolyzed rubber and crosslinking agent to form an integrated thermoplastic elastomer (ITPE) for asphalt modification. The mechanical, thermal, and interfacial behaviors of the ITPE were evaluated through tensile testing, thermogravimetric analysis, and scanning electron microscopy. The storage stability and rheological properties of the modified binder blends were evaluated through the cigar tube test and dynamic shear rheometer testing. Results demonstrated a successful formation of imide bonds in the ITPE, which can improve the strength, ductility, and thermal stability of rubber–plastic composites. Appropriate utilization of crosslinking agents can improve both rutting and fatigue resistance of ITPE-modified asphalt with good storage stability because of the co-existence of rigid plastic and soft rubbery regimes and the formation of a crosslink network. Furthermore, excessive content of crosslinker led to severe phase separation and reduced storage stability of modified binder blends. Extra crosslinker tended to float in asphalt because of its low density and caused an excessive formation of the crosslink network in the top section of the asphalt.

asphalt binder modifiers↗