Search NASA⌕ Search

SEARCH · Search NASA

Results for “Intel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Evaluating Application Characteristics for GPU Portability Layer Selection

GPUs have become the dominant source of computing power for high performance computing and are increasingly being used across the High Energy Physics computing landscape for a wide variety of tasks. Though NVIDIA is currently the main provider of GPUs, AMD and Intel are rapidly increasing their market share. As a result, programming using a vendor-specific language such as CUDA can significantly reduce deployment choices. There are a number of portability layers such as Kokkos, Alpaka, SYCL, OpenMP and std::par that permit execution on a broad range of GPU and CPU architectures, significantly increasing the flexibility of application programmers. However, each of these portability layers has its own characteristics, performing better at some tasks and worse at others, or placing limitations on aspects of the application. In this presentation, we report on a study of application and kernel characteristics that can influence the choice of a portability layer and show how each layer handles these characteristics. We have analyzed representative heterogeneous applications from CMS (patatrack and p2r), DUNE (Wire-Cell Toolkit), and ATLAS (FastCaloSim) to identify key application characteristics that have different behaviors for the various portability technologies. Using these results, developers can make more informed decisions on which GPU portability technology is best suited to their application.

Atif, Mohammad [Brookhaven]↗

A Beginner's Guide to Power and Energy Measurement and Estimation for Computing and Machine Learning

Concerns about the environmental footprint of machine learning are increasing. While studies of energy use and emissions of ML models are a growing subfield, most ML researchers and developers still do not incorporate energy measurement as part of their work practices. While measuring energy is a crucial step towards reducing carbon footprint, it is also not straightforward. This paper introduces the main considerations necessary for making sound use of energy measurement tools and interpreting energy estimates, including the use of at-the-wall versus on-device measurements, sampling strategies and best practices, common sources of error, and proxy measures. It also contains practical tips and real-world scenarios that illustrate how these considerations come into play. It concludes with a call to action for improving the state of the art of measurement methods and standards for facilitating robust comparisons between diverse hardware and software environments.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

HamLib: A library of Hamiltonians for benchmarking quantum algorithms and hardware

In order to characterize and benchmark computational hardware, software, and algorithms, it is essential to have many problem instances on-hand. This is no less true for quantum computation, where a large collection of real-world problem instances would allow for benchmarking studies that in turn help to improve both algorithms and hardware designs. To this end, here we present a large dataset of qubit-based quantum Hamiltonians. The dataset, called HamLib (for Hamiltonian Library), is freely available online and contains problem sizes ranging from 2 to 1000 qubits. HamLib includes problem instances of the Heisenberg model, Fermi-Hubbard model, Bose-Hubbard model, molecular electronic structure, molecular vibrational structure, MaxCut, Max- k -SAT, Max- k -Cut, QMaxCut, and the traveling salesperson problem. The goals of this effort are (a) to save researchers time by eliminating the need to prepare problem instances and map them to qubit representations, (b) to allow for more thorough tests of new algorithms and hardware, and (c) to allow for reproducibility and standardization across research studies.

97 MATHEMATICS AND COMPUTING↗

Extreme Ultraviolet and Beyond Extreme Ultraviolet Lithography Using Amorphous Zeolitic Imidazolate Resists Deposited by Atomic/Molecular Layer Deposition

Amorphous zinc-imidazolate (aZnMIm) resists show potential to meet the demands for next-generation high-numerical aperture (high-NA) metal-containing extreme ultraviolet (EUV) resist materials, given their ease of deposition by atomic/molecular layer deposition (ALD/MLD) at thicknesses of 20 nm and below. Here, this study demonstrates that aZnMIm thin films, previously identified as high-resolution electron beam resists, can also function as negative-tone EUV photoresists. Water development achieves high sensitivity (5 mJ/cm 2 ) but leaves significant residue, while acetic acid development results in poor contrast. A hybrid approach─water followed by acetic acid─enables residue-free development with a sensitivity of 181 mJ/cm 2 . Dry development using 1,1,1,5,5,5-hexafluoroacetylacetone (hfacH) is also possible but shows lower sensitivity (375 mJ/cm 2 ) compared to wet development methods. EUV photoelectron spectroscopy (PES), reflectometry/EUV absorption, total electron yield (TEY), residual gas analysis (RGA), and time-of-flight secondary ion mass spectrometry (TOF-SIMS) were used to investigate the effects of EUV irradiation on aZnMIm resists. Reflectometry experiments reveal an aZnMIm EUV absorption coefficient of 6.2 μm –1 , while PES and TEY analyses show that, compared to poly(4-hydroxystyrene) (PHS), a polymer-based reference resist, aZnMIm emits more primary and secondary electrons but generates fewer slow electrons relative to its primary electron emission; its total electron yield is similar to that of poly(methyl methacrylate) (PMMA) resists. When exposed to EUV, aZnMIm predominantly outgasses H 2 , as determined by RGA. TOF-SIMS measurements demonstrate that high-dose EUV exposure only partially fragments the 2-methylimidazole (2MIm) organic linkers, unlike high-dose electron beam exposure, which is known to completely degrade them. Additionally, aZnMIm resists show promise for potential beyond EUV lithography (BEUVL) due to the presence of Zn, which provides higher sensitivity at a wavelength of 6.7 nm compared to other metal ions, such as Sn, that are currently used in the best-performing EUV metal–organic resists. TEY measurements demonstrate that aZnMIm emits nearly twice as many electrons as PHS at 6.7 nm. The BEUV TEY of aZnMIm also surpasses that of PMMA, poly(pentafluorostyrene), and poly(4-iodostyrene), with the latter two being known for their high EUV TEYs. This work provides insight into zeolitic imidazolate framework (ZIF)-based EUV and BEUV resists and highlights their potential for both wet and dry development.

lithography↗

Room-temperature multiferroicity in sliding van der Waals semiconductors with sub-0.3 V switching

The search for van der Waals (vdW) multiferroic materials has been challenging but also holds great potential for the next-generation multifunctional nanoelectronics. The group-IV monochalcogenide, with an anisotropic puckered structure and an intrinsic in-plane polarization at room temperature, manifests itself as a promising candidate with coupled ferroelectric and ferroelastic order as the basis for multiferroic behavior. Unlike the intrinsic centrosymmetric AB stacking, we demonstrate a multiferroic phase of tin selenide (SnSe), where the inversion symmetry breaking is maintained in AA-stacked multilayers over a wide range of thicknesses. We observe that an interlayer-sliding-induced out-of-plane (OOP) ferroelectric polarization couples with the in-plane (IP) one, making it possible to control out-of-plane polarization via in-plane electric field and vice versa. Notably, thickness scaling yields a sub-0.3 V ferroelectric switching, which promises future low-power-consumption applications. Furthermore, coexisting armchair- and zigzag-like structural domains are imaged under electron microscopy, providing experimental evidence for the degenerate ferroelastic ground states theoretically predicted. Non-centrosymmetric SnSe, as the first layered multiferroic at room temperature, provides a novel platform not only to explore the interactions between elementary excitations with controlled symmetries, but also to efficiently tune the device performance via external electric and mechanical stress.

Chen, Rui [University of California, Berkeley, CA ↗

Framework of compressive sensing and data compression for 4D-STEM

Four-dimensional Scanning Transmission Electron Microscopy (4D-STEM) is a powerful technique for high-resolution and high-precision materials characterization at multiple length scales, including the characterization of beam-sensitive materials. However, the field of view of 4D-STEM is relatively small, which in absence of live processing is limited by the data size required for storage. Furthermore, the rectilinear scan approach currently employed in 4D-STEM places a resolution- and signal-dependent dose limit for the study of beam sensitive materials. Improving 4D-STEM data and dose efficiency, by keeping the data size manageable while limiting the amount of electron dose, is thus critical for broader applications. Here we introduce a general method for reconstructing 4D-STEM data with subsampling in both real and reciprocal spaces at high fidelity. The approach is first tested on the subsampled datasets created from a full 4D-STEM dataset, and then demonstrated experimentally using random scan in real-space. The same reconstruction algorithm can also be used for compression of 4D-STEM datasets, leading to a large reduction (100 times or more) in data size, while retaining the fine features of 4D-STEM imaging, for crystalline samples.

4D-STEM↗

Stable Machine‐Learning Parameterization of Subgrid Processes in a Comprehensive Atmospheric Model Learned From Embedded Convection‐Permitting Simulations

Modern climate projections often suffer from inadequate spatial and temporal resolution due to computational limitations, resulting in inaccurate representations of sub-grid processes. A promising technique to address this is the multiscale modeling framework (MMF), which embeds a kilometer-resolution cloud-resolving model (CRM) within each atmospheric column of a host climate model to replace traditional convection and cloud parameterizations. Machine learning offers a unique opportunity to make MMF more accessible by emulating the embedded CRM and reducing its substantial computational cost. Although many studies have demonstrated proof-of-concept success of achieving stable hybrid simulations, it remains a challenge to achieve near operational-level success with real geography and comprehensive variable emulation that includes, for example, explicit cloud condensate coupling. In this study, we present a stable hybrid model capable of integrating for at least 5 years with near operational-level complexity, including coarse-grid geography, seasonality, explicit cloud condensate and wind predictions, and land coupling. Our model demonstrates skillful online performance, achieving a 5-year zonal mean tropospheric temperature bias within 2 K, water vapor bias within 1 g/kg, and a precipitation root mean square error of 0.96 mm/day. Key factors contributing to our online performance include an expressive U-Net architecture and physical thermodynamic constraints for microphysics. With microphysical constraints mitigating unrealistic cloud formation, our work is the first to demonstrate realistic multi-year cloud condensate climatology under the MMF framework. Despite these advances, online diagnostics reveal persistent biases in certain regions, highlighting the need for innovative strategies to further optimize online performance.

Hu, Zeyuan [NVIDIA Corporation, Santa Clara, CA (U↗

The Heterogeneous Integration of Electronic Components

Heterogeneous integration (HI) of electronics components is broadly recognized as a powerful and crucial enabler for the continued growth of computing and communication. From 2010 onwards, the value of HI is increasingly visible in the advanced packaging used in artificial intelligence, high-performance computing, smartphones and communications product implementations. In this Perspective, we argue that HI is crucial to semiconductors and more broadly to the continued evolution of computing and communications. We use leading-edge advanced packaging examples to represent the value, advancements and opportunities for HI. To succeed, it is critical to develop comprehensive HI roadmaps that inform collaborations across the design, manufacturing and reliability spectrum between systems architects, packaging and semiconductor technologists to common goals. Although this article does not provide a full roadmap, we instead detail additional parameters for artificial intelligence, smartphone and other cellular communication devices, and their constituent building blocks including interconnects, power electronics, photonics, thermal management, reliability, modelling and co-design, to foster greater collaboration opportunities among academia, research laboratories and industry.

42 ENGINEERING↗

Broad-range tuning of ferroelectric switching of La x Bi 1−x FeO 3 epitaxial films via digital doping using off-axis co-sputtering

To investigate the scope of ferroelectric behavior in La-substituted BiFeO 3 films, La x Bi 1−x FeO 3 epitaxial films were synthesized using off-axis co-sputtering on SrTiO 3 (001) and DyScO 3 (110) substrates with a SrRuO 3 bottom electrode layer. A digital-doping deposition method was used to enable precise control and continuous tuning of La concentration in high-quality LaxBi 1−x FeO 3 films across a wide range of x = 0.05–0.60, which was systematically investigated using piezoresponse force microscopy. Robust and reversible out-of-plane ferroelectric switching has been observed up to x = 0.35, while films with x ≥ 0.37 exhibit no measurable ferroelectric behavior, indicating a sharp ferroelectric-to-paraelectric phase transition between x = 0.35 and 0.37. This represents the highest reported La concentration in LaxBi 1−x FeO 3 films that retains ferroelectric ordering, highlighting opportunities to engineer ferroelectric and multiferroic properties in complex oxide heterostructures.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Energy Efficiency Scaling for 2 Decades (EES2) Roadmap for Computing

In response to the looming crisis in global energy consumption required for advanced computing applications, the United States Department of Energy (DOE) Advanced Materials and Manufacturing Technology Office (AMMTO) is leading a multi-organizational effort to define a roadmap for energy efficiency scaling for two decades (EES2) with the aim to reduce energy use in all aspects of computation by more than a factor of 1000 in two decades. By July of 2024, over 60 organizations representing industry, academia, and the national laboratories have pledged to work in various aspects of research and development to enable energy efficiency in computing including in the development of the EES2 roadmap, with an initial public release in 2024 as the first phase of an ongoing commitment to energy-efficient and sustainable computation.

Kaarsberg, Tina [U.S. Department of Energy (DOE)]↗

Influence of sizing concentration on strength, stiffness, and porosity in textile grade carbon fiber (TCF)-Epoxy composites: Revealing inverse trends

The effect of fiber sizing (i.e., surface treatment) concentration (0 %, 1.36 %, 1.52 %, 1.94 %, and 2.13 %) on the mechanical properties (tensile, flexural, interlaminar shear strength (ILSS), and low velocity impact) of textile grade carbon fiber (TCF)-epoxy composite is examined. An inverse relationship between the strength and stiffness of the composite is observed with increased sizing concentration. The root mean square (RMS) roughness of the fiber surface increased from 17.8 nm (unsized) to 22.7 nm with 2.13 % sizing concentration. It was found that the tensile strength increased by 131 % from 221.4 ± 18.5 MPa (unsized) to 510.8 ± 28.05 MPa (for 1.36 % sizing) and further by 155 %–563.7 ± 14.95 MPa at 2.13 % sizing. On the contrary, the stiffness is initially increased by 126 % from 33.52 ± 7.80 GPa (unsized) to 75.9 ± 3.21 GPa (for 1.36 % sizing) but reduced with further increase in the sizing concentration. A single fiber pull-out test is simulated using the finite element method to validate the reverse trend in strength and stiffness. The varying sizing concentration is simulated by introducing an interface of varying thickness between fibers and matrix. Simulation results confirm that a thicker interface, corresponding to a higher sizing concentration, decreases interfacial shear stress, enhancing material strength while decreasing stiffness. The reverse trend in strength and stiffness with the sizing concentration aligns with experimental observations. In conclusion, the present study emphasizes the importance of sizing concentration for mechanical properties and provide a design criterion for customized high-strength and high-stiffness applications.

Porosity↗

Modeling CO 2 flow through faulted/fractured reservoirs using tEDFM in corner-point grids

The interest in underground CO 2 storage has increased significantly over the last decade because of the rising concern about global warming due to the growing levels of greenhouse gases in the atmosphere. Considering that CO 2 accounts for 80% of these greenhouse gases, carbon capture, utilization, and storage (CCUS) is regarded as one of the most direct approaches to achieving the net zero carbon target. Although CO 2 storage in deep saline aquifers and depleted gas reservoirs has been studied extensively, most studies use commercial simulators that model faults/fractures by simply modifying the transmissibility in the direction perpendicular to the fault surfaces. Here, this work shows that this simplistic approach ignores the accelerated flow in the directions parallel to the fault plane, leading to significantly higher leakage along the fault surface. To accurately model the flow of CO 2 in faulted reservoirs, we present the first transient embedded discrete fracture model for corner-point grids (tEDFM-CPG). By comparing the results of the tEDFM-CPG to high-resolution reference solutions, we show that this approach is accurate and efficient at predicting CO 2 flow in faulted/fractured reservoirs. Finally, this work presents the use of mixed reality (MR) to efficiently observe CO 2 gas migration in the interior of these corner-point grid systems.

25 ENERGY STORAGE↗

A High-Efficiency Delayed Update Algorithm for Evaluating Slater Determinants in Quantum Monte Carlo

For quantum Monte Carlo simulations of molecular systems or supercells with thousands of electrons, matrix operations related to Slater determinants lead the computational cost. McDaniel et al. [J. Chem. Phys. 2017, 147, 174107] proposed a delayed update algorithm to increase computational efficiency by using matrix–matrix multiplication when updating the inverse matrices of Slater determinants. However, preparing intermediate matrices for applying the Sherman–Morrison–Woodbury formula remained a bottleneck. Here, in this work, we introduce an improved algorithm for CPUs and GPUs that (1) reduces this bottleneck by iteratively updating the intermediate matrices and (2) is efficient at any acceptance ratio, with no cost for rejected moves on CPUs and minimal cost on GPUs. We show the full scheme of integrating the delayed update algorithm into a single-electron move. The high efficiency of our algorithm is demonstrated on CPUs and GPUs for a 512 atom/6144 valence electron calculation, with 12× and 2× overall speed-up compared to traditional rank-1 update schemes in diffusion quantum Monte Carlo, respectively.

Luo, Ye [Argonne National Laboratory (ANL), Argonn↗

Simulating Atmospheric Processes in Earth System Models and Quantifying Uncertainties With Deep Learning Multi‐Member and Stochastic Parameterizations

Abstract Deep learning is a powerful tool to represent subgrid processes in climate models, but many application cases have so far used idealized settings and deterministic approaches. Here, we develop stochastic parameterizations with calibrated uncertainty quantification to learn subgrid convective and turbulent processes and surface radiative fluxes of a superparameterization embedded in an Earth System Model (ESM). We explore three methods to construct stochastic parameterizations: (a) a single Deep Neural Network (DNN) with Monte Carlo Dropout; (b) a multi‐member parameterization; and (c) a Variational Encoder Decoder with latent space perturbation. We show that the multi‐member parameterization improves the representation of convective processes, especially in the planetary boundary layer, compared to individual DNNs. The respective uncertainty quantification illustrates that methods (b) and (c) are advantageous compared to a dropout‐based DNN parameterization regarding the spread of convective processes. Hybrid simulations with our best‐performing multi‐member parameterizations remained challenging and crash within the first days. Therefore, we develop a pragmatic partial coupling strategy relying on the superparameterization for condensate emulation. Partial coupling reduces the computational efficiency of hybrid Earth‐like simulations but enables model stability over 5 months with our multi‐member parameterizations. However, our hybrid simulations exhibit biases in thermodynamic fields and differences in precipitation patterns. Despite this, the multi‐member parameterizations enable improvements in reproducing tropical extreme precipitation compared to a traditional convection parameterization. Despite these challenges, our results indicate the potential of a new generation of multi‐member machine learning parameterizations leveraging uncertainty quantification to improve the representation of stochasticity of subgrid effects.

Behrens, Gunnar [Deutsches Zentrum für Luft‐ und R↗

Navigating the Noise: Bringing Clarity to ML Parameterization Design With O $\boldsymbol{\mathcal{O}}$(100) Ensembles

Abstract Machine‐learning (ML) parameterizations of subgrid processes (here of turbulence, convection, and radiation) may one day replace conventional parameterizations by emulating high‐resolution physics without the cost of explicit simulation. However, uncertainty about the relationship between offline and online performance (i.e., when integrated with a large‐scale general circulation model) hinders their development. Much of this uncertainty stems from limited sampling of the noisy, emergent effects of upstream ML design decisions on downstream online hybrid simulation. Our work rectifies the sampling issue via the construction of a semi‐automated, end‐to‐end pipeline for size ensembles of hybrid simulations, revealing important nuances in how systematic reductions in offline error manifest in changes to online error and online stability. For example, removing dropout and switching from a Mean Squared Error to a Mean Absolute Error loss both reduce offline error, but they have opposite effects on online error and online stability. Other design decisions, like incorporating memory, converting moisture input from specific humidity to relative humidity, using batch normalization, and training on multiple climates do not come with any such compromises. Finally, we show that ensemble sizes of may be necessary to reliably detect causally relevant differences online. By enabling rapid online experimentation at scale, we can empirically settle debates regarding subgrid ML parameterization design that would have otherwise remained unresolved in the noise.

Lin, Jerry [Department of Earth System Sciences Un↗

Unsupervised discovery of extreme weather events using universal representations of emergent organization

Spontaneous self-organization is ubiquitous in systems far from thermodynamic equilibrium. While organized structures that emerge dominate transport properties, universal representations that identify and describe these key objects remain elusive. Here, we introduce a theoretically grounded framework for describing emergent organization that, via data-driven algorithms, is constructive in practice. Its building blocks are spacetime lightcones that embody how information propagates across a system through local interactions. We show that predictive equivalence classes of lightcones—local causal states—capture organized behaviors in complex spatiotemporal systems. Employing an unsupervised physics-informed machine learning algorithm and a high-performance computing implementation, we demonstrate automatically discovering organized structures in two real-world domain science problems. We show that local causal states identify vortices and track their power-law decay behavior in two-dimensional fluid turbulence. We then show how to detect and track familiar extreme weather events—hurricanes and atmospheric rivers—and discover other novel structures associated with precipitation extremes in high-resolution climate data at the grid-cell level.

Rupe, Adam [Pacific Northwest National Laboratory ↗

QCD Predictions for Physical Multimeson Scattering Amplitudes

We use lattice QCD calculations of the finite-volume spectra of systems of two and three mesons to determine, for the first time, three-particle scattering amplitudes with physical quark masses. Our results are for combinations of 𝜋 + and 𝐾 + , at a lattice spacing 𝑎 = 0.063 fm, and in the isospin-symmetric limit. We also obtain accurate results for maximal-isospin two-meson amplitudes, with those for 𝜋 + ⁢𝐾 + and 2⁢𝐾 + being the first determinations at the physical point. Dense lattice spectra are obtained using the stochastic Laplacian-Heaviside method, and the analysis leading to scattering amplitudes is done using the relativistic finite-volume formalism. Results are compared to chiral perturbation theory and to phenomenological fits to experimental data, finding good agreement.

hadron-hadron interactions↗