Search NASASearch

SEARCH · Search NASA

Results for “cloud computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Containers on Switches: A Cluster School Experience

Network switches, such as those from Arista and Mellanox, often have underutilized computational resources in the form of built-in processors and memory. By leveraging these untapped resources, we can optimize functionality and efficiency of computational cluster networks. Our research focuses on deploying containers directly onto these switches to execute various auxiliary tasks ranging from metric logging to system-wide management via post-boot configuration. By doing so, we can significantly enchance the capabilities of the cluster without the need for additional dedicated hardware. Our research involved five distinct scenarios where switch utilization could have a profound impact on HPC Clusters: run cloud-init services via link-local connection; configuring a Telegraf container to export metrics; deploying a caching proxy; creating a reconfigurable IPv6 DHCP/DNS provider for VLAN; and implementing a client detection with Magellan discovery. These scenarios were containerized with podman and docker, and tested both physically on the switch virtually on a QEMU VM both running SONiC OS. Testing and findings indicate that network switches can indeed be used for these scenarios. They offer a wide range of possibilities beyond these applications. They run as expected as containers on the switches, and although there were some minor issues, work-arounds were implemented. Overall, this is a positive result that can be further explored with more scenarios.

97 MATHEMATICS AND COMPUTING

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY

Accelerated data-driven materials science with the Materials Project

The Materials Project was launched formally in 2011 to drive materials discovery forwards through high-throughput computation and open data. More than a decade later, the Materials Project has become an indispensable tool used by more than 600,000 materials researchers around the world. This Perspective describes how the Materials Project, as a data platform and a software ecosystem, has helped to shape research in data-driven materials science. We cover how sustainable software and computational methods have accelerated materials design while becoming more open source and collaborative in nature. Next, we present cases where the Materials Project was used to understand and discover functional materials. We then describe our efforts to meet the needs of an expanding user base, through technical infrastructure updates ranging from data architecture and cloud resources to interactive web applications. Finally, we discuss opportunities to better aid the research community, with the vision that more accessible and easy-to-understand materials data will result in democratized materials knowledge and an increasingly collaborative community.

Horton, Matthew K

Quantum scattering of HC 5 N and para -H 2 on a new potential energy surface

In the interstellar medium (ISM), non-local thermodynamic equilibrium situations are common due to low density, and one needs to consider the effect of molecular collisions in order to interpret the observations. Among the species detected in the ISM, cyanopolyynes, with the general molecular formula HC 2n+1 N (n = 1, 2, …), are characterized by large dipole moments and small rotational constants and constitute an indispensable class of candidates for the sensitive tracers of local density and temperature. We present a study of the collisional (de-) excitation of HC 5 N by para -H 2 (p-H 2 ) in its ground rotational state, namely HC 5 N ( j 1 ) + H 2 ( j 2 = 0) → HC 5 N (j$_1^′$) + H2 (j$_2^′$ = 0), where j 1 (or j$_1^′$) and j 2 (or j$_2^′$) denote the initial (or final) rotational quantum numbers of HC 5 N and H 2 , respectively. We performed the quantum scattering calculations at low collision energy using a new four-dimensional ab initio potential energy surface. In the regime where p-H 2 remains in its rotational ground state, converged cross sections did not require including excited rotational states of p-H 2 in the rotational basis. State-to-state cross sections were computed by means of the quantum-mechanical close-coupling (CC) method and the coupled states (CS) approximation, and rate coefficients for the first 61 levels of HC 5 N were computed for the first time up to 20 K with the CC approach and up to 50 K with the CS method. CC and CS results were found to agree well at temperatures up to 20 K. Finally, these data should allow a more accurate derivation of the HC 5 N abundance in molecular clouds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Impacts of Bulk Microphysics Scheme Structural Choices on Simulations of Rain Initiation Through Drop Coalescence

This study examines how different structural choices in bulk microphysics schemes impact the simulation of warm rain initiation. A single liquid category (SLC) approach prognosing up to four moments of a single drop size distribution (DSD) is compared to the traditional two-category, two-moment approach with separate DSDs for cloud and rain (four total prognostic variables). Different methods for calculating tendencies of the prognostic variables from drop collision-coalescence are also tested: a discretized numerical-integration approach, machine learning via neural networks, lookup tables, and traditional power law fits. Relative to simulations using a bin microphysics model, SLC gives smaller error overall than the two-category approach when numerical integration is used to calculate the collision-coalescence tendencies for both. Replacing the numerical integration with a pre-computed lookup table reduces computational cost with little loss of accuracy. However, using fitted power laws with SLC to represent the collision-coalescence tendencies substantially reduces accuracy and leads to an order of magnitude increase in error. It is also demonstrated that with SLC, reasonably accurate solutions are obtained using only three prognostic moments, while a two-moment SLC scheme leads to substantial error. Overall, both the choice of prognostic moments (e.g., SLC vs. two-category) and method to calculate the collision-coalescence tendencies are important to consider for minimizing errors in bulk schemes. SLC with a sufficiently detailed calculation of the collision-coalescence tendencies provides accurate solutions for a reasonable computational cost, providing a viable alternative to the traditional two-category, two-moment approach for bulk microphysics.

320 (cloud physics and chemistry)

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility

A Machine Learning Framework for Predicting Microphysical Properties of Ice Crystals From Cloud Particle Imagery

The microphysical properties of ice crystals are important because they significantly alter the radiative properties and spatiotemporal distributions of clouds, which in turn strongly affect Earth's climate. However, it is challenging to measure key properties of ice crystals, such as mass or morphological features. Here, we present a proof-of-concept framework for predicting three-dimensional (3D) microphysical properties of ice crystals from in situ two-dimensional (2D) imagery. First, we computationally generated synthetic ice crystals using 3D modeling software along with geometric parameters estimated from the 2021 Ice Cryo-Encapsulation Balloon (ICEBall) field campaign. Then, we used synthetic crystals to train machine learning (ML) models to predict effective density ($ρ_e$), effective surface area ($A_e$), and number of bullets ($N_b$) from synthetic rosette imagery. On unseen synthetic images, our ML models accurately predicted ice crystal properties. ResNet-18 performed best, achieving $R^2$ values of 0.99 and 0.98 for $ρ_e$ and $A_e$, respectively, and MAE of 0.10 for mathematical equation in single view tasks. Stereo view ResNet-18 further reduced RMSE by 40% for $ρ_e$ and $A_e$ and reduced MAE by 0.08 for $N_b$. This work provides a novel ML-driven framework for estimating ice microphysical properties from in situ imagery, which will allow for downstream constraints on microphysical parameterizations, such as the mass-size relationship.

Ko, J. [Columbia Univ., New York, NY (United State

The impact of aerosol mixing state on immersion freezing: insights from classical nucleation theory and particle-resolved simulations

Immersion freezing, initiated by ice-nucleating particles (INPs) in supercooled aqueous droplets, plays an important role in the formation of ice crystals within clouds. The efficiency of immersion freezing depends strongly on INP composition and, crucially, on the mixing state – how chemical species are distributed across the particle population. Here, we quantify the impact of aerosol mixing state on immersion freezing using a combined theoretical and particle-resolved modeling approach. We derive analytical expressions for the frozen fraction of internally and externally mixed INP populations based on classical nucleation theory, showing that the frozen fraction is sensitive to whether ice-active species are present in all particles or only in a subset of the population. We introduce a multi-species immersion freezing scheme into the particle-resolved model PartMC, using the water activity-based immersion freezing model (ABIFM) to compute freezing probabilities for mixed-composition particles. To improve computational efficiency, we implement a Binned Tau-Leaping algorithm and demonstrate an order-of-magnitude speedup with minimal accuracy loss. Simulations reproduce the analytical trends in limiting cases and extend the analysis to more general aerosol populations, where mixing state continues to exert a substantial control on frozen fraction. Sensitivity analyses across particle size, species type, and cooling condition reveal that the mixing state effect is most pronounced when small amounts of highly efficient INPs are mixed with less efficient materials. These findings underscore the need to represent aerosol mixing state explicitly in models of heterogeneous ice nucleation to reduce uncertainty in cloud-phase partitioning.

54 ENVIRONMENTAL SCIENCES

Applying Corrective Machine Learning in the E3SM Atmosphere Model in C++ (EAMxx)

The Simplified Cloud-Resolving E3SM Atmosphere Model (SCREAM) is the newest addition to the family of Earth System Models capable of explicitly resolving convective systems. SCREAM is a kilometer-scale configuration of the advanced E3SM Atmosphere Model (EAMxx), designed for heterogeneous systems. While the enhanced accuracy of kilometer-scale modeling offers significant benefits, it comes with a substantial computational cost, limiting feasible simulation durations to only a few years, even on the fastest supercomputers. Machine learning presents an opportunity for scientists to achieve the high accuracy of storm-resolving models at a significantly reduced cost. Building on the previous success of applying corrective machine learning (ML) to the FV3 model, this study explores the effects of implementing corrective ML in EAMxx-SCREAM. We also address the computational challenges of integrating the corrective ML, which is written in Python, with the C++/Kokkos EAMxx driver, as well as the potential pitfalls of generalizing an approach that was effective with one atmosphere model to another.

54 ENVIRONMENTAL SCIENCES

A Particle-in-cell Method for Plasmas with A Generalized Momentum Formulation, Part III: A family of Gauge Conserving Methods

In this paper, we introduce a new family of spatially co-located field solvers for particle-in-cell applications which evolve the potential formulation of Maxwell’s equations under the Lorenz gauge. Our recent work [2] introduced the concept of time-consistency, which connects charge conservation to the preservation of the gauge at the semi-discrete level. It will be shown that there exists a large family of time discretizations which satisfy this property. Additionally, it will be further shown that for large classes of time marching methods, the satisfaction of the gauge condition automatically implies the satisfaction of Gauss’s law for electricity, with the potential formulation ensuring that that Gauss’s law for magnetism is satisfied by definition. We focus on popular time marching methods including centered differences, backward differences, and diagonally-implicit Runge-Kutta methods, which are coupled to a spectral discretization in space. We demonstrate the theory by testing the methods on a relativistic Weibel instability and a drifting cloud of electrons.

97 MATHEMATICS AND COMPUTING

Spatially Aware Linear Transformer (SAL-T) for Particle Jet Tagging

Transformers are very effective in capturing both global and local correlations within high-energy particle collisions, but they present deployment challenges in high-data-throughput environments, such as the CERN LHC. The quadratic complexity of transformer models demands substantial resources and increases latency during inference. In order to address these issues, we introduce the Spatially Aware Linear Transformer (SAL-T), a physics-inspired enhancement of the linformer architecture that maintains linear attention. Our method incorporates spatially aware partitioning of particles based on kinematic features, thereby computing attention between regions of physical significance. Additionally, we employ convolutional layers to capture local correlations, informed by insights from jet physics. In addition to outperforming the standard linformer in jet classification tasks, SAL-T also achieves classification results comparable to full-attention transformers, while using considerably fewer resources with lower latency during inference. Experiments on a generic point cloud classification dataset (ModelNet10) further confirm this trend. Our code is available at https://github.com/aaronw5/SAL-T4HEP.

Wang, Aaron [Illinois U., Chicago] (ORCID:00000003

Analysis of soot formation from aviation fuels in laminar counterflow flames

Combustion emissions from aviation contribute to the formation of condensation trail (contrail) that can lead to the formation of anthropogenic cirrus clouds. Ice particles that form contrails are observed to have a linear correlation with soot particle number density. Synthetic aviation fuels (SAFs) offer a promising route to mitigate the production of soot particles while also increasing energy security. Although studies have focused on combustion and spray behavior, the detailed investigation of soot formation processes for different jet fuels and their impact on models for computational fluid dynamics (CFD) applications is not well understood. Moreover, experimental measurements of soot for canonical flames using Synthetic aviation fuels (SAF) for model validation remain scarce. To address this, we use employed the Lawrence Livermore National Laboratory (LLNL) detailed soot model based on the discrete sectional method. Additionally, we develop two reduced chemical mechanisms for Jet-A and Alcohol-to-Jet (C1) that are suitable for turbulent flame simulations and couple them with the Hybrid Method of Moments (HMOM). The detailed and reduced model frameworks are validated against experimental measurements of soot volume fraction (ƒ ν ) from a counterflow burner experiment previously reported in the literature. Given the good agreement between modeling results and experimental measurements for the (1) spatial distribution of ƒ ν and (2) the non-linear variation of peak ƒ ν with strain rate, we further investigate the modeled sub-processes (nucleation, condensation, surface growth, and oxidation) using the LLNL model to analyze the assumptions in the reduced model framework. Furthermore, the results indicate a significant contribution from resonant radicals to the surface growth of soot particles, which are not accounted for in the current implementation of HMOM and could help reconcile soot predictions by the reduced model with observations.

Counterflow

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility

Scalable Generation of High-fidelity Synthetic Population Ensembles

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the U.S. via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. The study involves two scenarios: creating ensembles for (1) 17 U.S. metropolitan areas in 2019 and (2) full U.S. Census Divisions in 2023, with each scenario consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system within a research cloud, comprised of virtual containerizations, GPU-enhanced functionality, and orchestrated deployments of UrbanPop’s maturing Likeness Python ecosystem. Results demonstrate we maintained high-fidelity approximations of residential totals by areas of interest and the demographic characteristics of neighborhoods while reducing manual workflow burdens. Finally, we discuss plans to fine-tune and further develop our automated workflows for truly distributed job orchestration to increase computational efficiency, as well as provide an outlook for broadening applications of the ensembles.

Cluster computing

Producing High-fidelity Synthetic Population Ensembles at Scale

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the US via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. Our initial task involves creating ensembles for 17 US metropolitan areas, each consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system comprised of a research cloud, virtual containerization, GPU-enhanced functionality, and a dual API/CLI to interact with UrbanPop’s maturing Likeness Python ecosystem. We observe a reduction in theoretical execution time while maintaining high-fidelity approximations of residential totals by metropolitan area and the demographic characteristics of neighborhoods. We discuss expansion of our approach to produce synthetic population ensembles for the entire US, particularly plans to establish automated workflows for job orchestration to increase computational efficiency, as well as provide outlook for broadening applications of the ensembles.

Gaboardi, James [ORNL] (ORCID:0000000247766826)

High sensitivity of simulated fog properties to parameterized aerosol activation in case studies from ParisFog

Aerosols influence fog properties such as visibility and lifetime by affecting fog droplet number concentrations (N d ). Numerical weather prediction (NWP) models often represent aerosol–fog interactions using highly simplified approaches. Incorporating prognostic size-resolved aerosol microphysics from climate models could allow them to simulate N d and aerosol–fog interactions without incurring excessive computational expense. However, microphysics code designed for coarse spatial resolution may struggle with sub-kilometer-scale grid spacings. Here, we test the ability of the UK Met Office Unified Model to simulate aerosol and fog properties during case studies from the ParisFog field campaign in 2011. We examine the sensitivity of fog properties to variations in N d caused by modifications to simulated aerosol activation. Our model, with a 500 m horizontal resolution and interactive aerosol and cloud microphysics, significantly underpredicts N d , although it only slightly underestimates the cloud condensation nuclei concentration. With an updated version of the Abdul-Razzak and Ghan (2000) activation scheme, we produce N d that are more consistent with those predicted by a cloud parcel model under fog-like conditions. We activate droplets only by adiabatic cooling. We incorporate more realistic hygroscopicities for sulfate and organic aerosols and explore the sensitivity of simulated N d to unresolved updrafts. We find that both N d and simulated fog liquid water content are very sensitive to the updated activation scheme but remain less affected by the update to hygroscopicities. Our improvements offer insights into the physical processes regulating N d in stable conditions, potentially laying foundations for improved operational fog forecasts that incorporate interactive aerosol simulations or aerosol climatologies.

Ghosh, Pratapaditya [Carnegie Mellon University, P

Baltimore Social-Environmental Collaborative (BSEC) Doppler Lidar & Derived Products

This repository contains all processed Doppler‐lidar outputs from the PSU lidar deployed for the Baltimore Social‐Environmental Collaborative (BSEC) project. Vertical Stare Scans (fixed‐beam, vertical profiling): 1 Hz backscatter intensity (m⁻¹ sr⁻¹), signal‐to‐noise ratio (unitless), and Doppler vertical‐velocity (m s⁻¹) on ~30 m range gates, stored as CF-compliant NetCDF. Wind Profiles (horizontal‐wind retrieval): daily NetCDF outputs of retrieved horizontal wind speed (m s⁻¹) and direction (degrees), computed from the angled‐scan returns. Profile Statistics (summary statistics on the vertical velocity): 15 min windows (default) of mean, variance, skewness, kurtosis, high-frequency variance, etc., as a function of height; saved as CF-compliant NetCDF files. Boundary Layer Height (BLH) (fuzzy-logic output): 15 min BLH estimates (m), with lower/upper fuzzy bounds (m) and a quality flag (0–4) indicating data status (e.g., no data, good, below range, ran out of signal, cloud-topped). Cloud Base Height (Haar-gradient detection): 15 min estimates of cloud-base height (m) with a cloud-detection quality flag (0–3: none, low, moderate, high). All five product streams are organized by year and date under their own top-level folders (01_Vertical_Stare_Scans/ through 05_Cloud_Height/). Each folder contains a data_ /YYYY/ subdirectory with daily CF-compliant NetCDF outputs (96 windows per day at 15 min intervals). Global attributes in each file include creation history, version (2.0.0), institution, and source. Instrument & MeasurementsThe PSU Doppler Lidar samples aerosol backscatter (m⁻¹ sr⁻¹), signal-to-noise ratio, and radial velocity at ~1 Hz. Vertical stare scans point the beam straight up; after collecting angled scans through multiple elevation angles, the "Wind Profiles" product contains the fully retrieved horizontal wind speed and direction. Data were collected continuously at ~30 m range resolution, with a typical height ceiling of ~12 km. How to Use Open any NetCDF with Python's xarray, MATLAB, or similar CF-compliant tools. Stare scans and angled-scan retrievals (Wind Profiles) are CF-compliant daily NetCDF files. Profile-Statistics, BLH, and Cloud Height files are daily 15 min summaries (96 time steps per file). Inspect the included variables (e.g., vertical_velocity_variance, wind_speed, BLH, cloud_base_height) for your analyses. Use the quality flags (BLH_flag, cloud_flag) to filter out poor-quality retrievals. For more information or questions about processing methods, please contact:Nicholas E. Prince ⟨nec5299@psu.edu⟩Penn State Department of Meteorology & Atmospheric Science

Air Quality