Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Position Papers for the 2024 ASCR Workshop on Energy-Efficient Computing for Science

On behalf of the Advanced Scientific Computing Research (ASCR) program in the US Department of Energy (DOE) Office of Science, we are organizing a Workshop on Energy-Efficient Computing for Science (EECS). Energy efficiency involves coordination across all the interoperating components of a computing system—in particular, applications, algorithms, system software, programming models, data management, and the hardware on which they run. Looking 10-15 years into the future, the goal is to dramatically lower the energy costs of the computational platforms (from the data center to the edge) serving DOE science while expanding the capabilities of these systems, broadening their applicability to science challenges of interest to DOE and the nation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Adaptive continuity-preserving simplification of street networks

While street network data are nearly universally available, their representation is usually transportation-based. However, for many types of analyses, e.g., urban morphology or network science, unprocessed transportation-based street network data is unsuitable, making a cumbersome manual simplification process necessary. To address this challenge, in this paper we propose an algorithm for simplification of street networks, based on the detection of network portions that need to be simplified, and continuity-preserving heuristics that generate new geometries. The algorithm, released in the open-source Python package neatnet, facilitates the generation of morphological networks and generalises to various geographical contexts without a need to alter the parameters, while offering better performance than other available solutions.

Fleischmann, Martin [Charles University, Prague, C↗

Advanced Interactive 3D Visualization Tool for Customizable Analyses of Tomography Datasets in Material Science

Current methods for visualizing and analyzing 3D tomography datasets in materials science often lack the interactivity and depth required for detailed structural insights. This limitation restricts a researchers' ability to accurately interpret complex data, which is critical for advancing material innovations and understanding structural properties. To address this issue, we have developed a novel, web-based interactive 3D visualization and analysis tool from the Trame framework that offers customizable features to enhance data interpretability. The tool allows users to adjust parameters such as visible range, slice planes, data rotation, and layering, providing a more detailed and dynamic view of complex structures. Its user-friendly web interface increases the accessibility and ease of use for both novice and experienced researchers, to visualize large volumetric datasets. The tool supports a diverse range of data formats, making it versatile for various research applications. Unique capabilities include real-time data manipulation, automated feature detection, context-sensitive feedback, and real-time volume calculations and distributions per sliced region or layer, alongside the ability to quickly generate high-quality screenshots and videos for presentations and reports. These advancements offer a comprehensive solution for enhanced 3D data exploration, significantly improving the analysis process and communication of results in materials science.

36 - MATERIALS SCIENCE↗

Data Imbalance, Uncertainty Quantification, and Transfer Learning in Data‐Driven Parameterizations: Lessons From the Emulation of Gravity Wave Momentum Transport in WACCM

Abstract Neural networks (NNs) are increasingly used for data‐driven subgrid‐scale parameterizations in weather and climate models. While NNs are powerful tools for learning complex non‐linear relationships from data, there are several challenges in using them for parameterizations. Three of these challenges are (a) data imbalance related to learning rare, often large‐amplitude, samples; (b) uncertainty quantification (UQ) of the predictions to provide an accuracy indicator; and (c) generalization to other climates, for example, those with different radiative forcings. Here, we examine the performance of methods for addressing these challenges using NN‐based emulators of the Whole Atmosphere Community Climate Model (WACCM) physics‐based gravity wave (GW) parameterizations as a test case. WACCM has complex, state‐of‐the‐art parameterizations for orography‐, convection‐, and front‐driven GWs. Convection‐ and orography‐driven GWs have significant data imbalance due to the absence of convection or orography in most grid points. We address data imbalance using resampling and/or weighted loss functions, enabling the successful emulation of parameterizations for all three sources. We demonstrate that three UQ methods (Bayesian NNs, variational auto‐encoders, and dropouts) provide ensemble spreads that correspond to accuracy during testing, offering criteria for identifying when an NN gives inaccurate predictions. Finally, we show that the accuracy of these NNs decreases for a warmer climate (4 × CO 2 ). However, their performance is significantly improved by applying transfer learning, for example, re‐training only one layer using ∼1% new data from the warmer climate. The findings of this study offer insights for developing reliable and generalizable data‐driven parameterizations for various processes, including (but not limited to) GWs.

54 ENVIRONMENTAL SCIENCES↗

Benchmarking universal machine learning interatomic potentials for rapid analysis of inelastic neutron scattering data

The accurate calculation of phonons and vibrational spectra remains a significant challenge, requiring highly precise evaluations of interatomic forces. Traditional methods based on the quantum description of the electronic structure, while widely used, are computationally expensive and demand substantial expertise. Emerging universal machine learning interatomic potentials (uMLIPs) offer a transformative alternative by employing pre-trained neural network surrogates to predict interatomic forces directly from atomic coordinates. This approach dramatically reduces computation time and minimizes the need for technical knowledge. In this paper, we produce a phonon database comprising nearly 5000 inorganic crystals to benchmark the performance of several leading uMLIPs. We further assess these models in real-world applications by using them to analyze experimental inelastic neutron scattering data collected on a variety of materials. Through detailed comparisons, we identify the strengths and limitations of these uMLIPs, providing insights into their accuracy and suitability for fast calculations of phonons and related properties, as well as the potential for real-time interpretation of neutron scattering spectra. Our findings highlight how the rapid advancement of AI in science is revolutionizing experimental research and data analysis.

inelastic neutron scattering↗

A prospective on machine learning challenges, progress, and potential in polymer science

Abstract Artificial intelligence and machine learning (ML) continue to see increasing interest in science and engineering every year. Polymer science is no different, though implementation of data-driven algorithms in this subfield has unique challenges barring widespread application of these techniques to the study of polymer systems. In this Prospective, we discuss several critical challenges to implementation of ML in polymer science, including polymer structure and representation, high-throughput techniques and limitations, and limited data availability. Promising studies targeting resolution of these issues are explored, and contemporary research demonstrating the potential of ML in polymer science despite existing obstacles are discussed. Finally, we present an outlook for ML in polymer science moving forward. Graphical Abstract

Struble, Daniel C. (ORCID:0009000093410612)↗

Adamantine 1.0: A Thermomechanical Simulator for Additive Manufacturing

Adamantine is a thermomechanical simulation code that is written in C++ and built on top of deal.II (Arndt et al., 2023), p4est (Burstedde et al., 2011), ArborX (Lebrun-Grandié et al., 2020), Trilinos (The Trilinos Project Team, 2020), and Kokkos (Trott et al., 2022). Adamantine was developed with additive manufacturing in mind and it is particularly well adapted to simulate fused filament fabrication, directed energy deposition, and powder bed fusion. Adamantine employs the finite element method with adaptive mesh refinement to solve a nonlinear anisotropic heat equation, enabling support for various additive manufacturing processes. It can also perform elastoplastic and thermoelastoplastic simulations. It can handle materials in three distinct phases (solid, liquid, and powder) to accurately reflect the physical state during different stages of the manufacturing process. To enhance simulation accuracy, adamantine incorporates data assimilation techniques (Asch et al., 2016). This allows it to integrate experimental data from sensors like thermocouples and infrared (IR) cameras. This combined approach helps account for errors arising from input parameters, material properties, models, and numerical calculations, leading to more realistic simulations that reflect what occurs in a particular print.

36 MATERIALS SCIENCE↗

Optimising the processing and storage of visibilities using lossy compression

The next-generation radio astronomy instruments are providing a massive increase in sensitivity and coverage, largely through increasing the number of stations in the array and the frequency span sampled. The two primary problems encountered when processing the resultant avalanche of data are the need for abundant storage and the constraints imposed by I/O, as I/O bandwidths drop significantly on cold storage. An example of this is the data deluge expected from the SKA Telescopes of more than 60 PB per day, all to be stored on the buffer filesystem. While compressing the data is an obvious solution, the impacts on the final data products are hard to predict. In this paper, we chose an error-controlled compressor – MGARD – and applied it to simulated SKA-Mid and real pathfinder visibility data, in noise-free and noise-dominated regimes. As the data have an implicit error level in the system temperature, using an error bound in compression provides a natural metric for compression. MGARD ensures the compression incurred errors adhere to the user-prescribed tolerance. To measure the degradation of images reconstructed using the lossy compressed data, we proposed a list of diagnostic measures, exploring the trade-off between these error bounds and the corresponding compression ratios, as well as the impact on science quality derived from the lossy compressed data products through a series of experiments. We studied the global and local impacts on the output images for continuum and spectral line examples. We found relative error bounds of as much as 10%, which provide compression ratios of about 20, have a limited impact on the continuum imaging as the increased noise is less than the image RMS, whereas a 1% error bound (compression ratio of 8) introduces an increase in noise of about an order of magnitude less than the image RMS. For extremely sensitive observations and for very precious data, we would recommend a 0.1% error bound with compression ratios of about 4. These have noise impacts two orders of magnitude less than the image RMS levels. At these levels, the limits are due to instabilities in the deconvolution methods. We compared the results to the alternative compression tool DYSCO, in both the impacts on the images and in the relative flexibility. MGARD provides better compression for similar error bounds and has a host of potentially powerful additional features.

Techniques: interferometric↗

Fusion and Fission Energy and Science Directorate and Information Technology Services Directorate HPC Cluster Reduction, Consolidation, and Savings in Data Center Space, Power, and Cooling

This report evaluates the benefits of decommissioning six legacy FFESD purchased HPC clusters and consolidating services and workloads into a new HPC cluster named HELIOS. The findings demonstrate significant reductions in the data center power and cooling requirements, data center footprint, and operational overhead, while simultaneously increasing computational capacity.

97 MATHEMATICS AND COMPUTING↗

Simulating Atmospheric Processes in Earth System Models and Quantifying Uncertainties With Deep Learning Multi‐Member and Stochastic Parameterizations

Abstract Deep learning is a powerful tool to represent subgrid processes in climate models, but many application cases have so far used idealized settings and deterministic approaches. Here, we develop stochastic parameterizations with calibrated uncertainty quantification to learn subgrid convective and turbulent processes and surface radiative fluxes of a superparameterization embedded in an Earth System Model (ESM). We explore three methods to construct stochastic parameterizations: (a) a single Deep Neural Network (DNN) with Monte Carlo Dropout; (b) a multi‐member parameterization; and (c) a Variational Encoder Decoder with latent space perturbation. We show that the multi‐member parameterization improves the representation of convective processes, especially in the planetary boundary layer, compared to individual DNNs. The respective uncertainty quantification illustrates that methods (b) and (c) are advantageous compared to a dropout‐based DNN parameterization regarding the spread of convective processes. Hybrid simulations with our best‐performing multi‐member parameterizations remained challenging and crash within the first days. Therefore, we develop a pragmatic partial coupling strategy relying on the superparameterization for condensate emulation. Partial coupling reduces the computational efficiency of hybrid Earth‐like simulations but enables model stability over 5 months with our multi‐member parameterizations. However, our hybrid simulations exhibit biases in thermodynamic fields and differences in precipitation patterns. Despite this, the multi‐member parameterizations enable improvements in reproducing tropical extreme precipitation compared to a traditional convection parameterization. Despite these challenges, our results indicate the potential of a new generation of multi‐member machine learning parameterizations leveraging uncertainty quantification to improve the representation of stochasticity of subgrid effects.

Behrens, Gunnar [Deutsches Zentrum für Luft‐ und R↗

ARM Cloud and Precipitation Measurements and Science Group (CPMSG) 2024 Workshop Report

The mission of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility is to improve the understanding and representation of cloud and aerosol processes and their interaction with the Earth's surface in Earth system models (ESMs) by providing comprehensive field observations and supporting advanced data analytics. The ARM Cloud and Precipitation Measurements and Science Group (CPMSG) was chartered in March 2019 to help improve the performance and scientific impact of ARM measurements of clouds and precipitation. The group aims to identify and address gaps in measurement capabilities, maximize the scientific impact of ARM data, and effectively serve the scientific community. To achieve these goals, the group includes experts in cloud and precipitation science, as well as representatives from ARM infrastructure, including instrument mentors, engineers, data quality officers, and data product translators. Prior to CPMSG, early discussions on cloud and precipitation measurements primarily focused on improving radar systems, but have since evolved to include a broader scope involving radiometers and other instruments. Since its formation, the CPMSG has gathered feedback using science traceability matrices. CPMSG aims to keep these as living documents to show the measurement needs, scientific drivers, roadblocks, maturity of measurements and retrievals, and pathways to model improvements. The group meets quarterly to discuss and prioritize measurement and operational improvements.

54 ENVIRONMENTAL SCIENCES↗

Harnessing Satellite Data Alone for Mapping Global Thermal Anisotropy

Mapping thermal anisotropy across global lands is critical for advancing a wide range of Earth science studies. However, a comprehensive understanding of global thermal anisotropy intensity (TAI) and its governing factors remains missing. We introduce a novel data-driven methodology to quantify global TAI exclusively using multi-angle MODIS land surface temperature time series observations. Our analysis reveals distinct seasonal and diurnal TAI patterns, with global mean summertime TAI exceeding 2.9°C. Furthermore, we identify strong associations between TAI and key surface and atmospheric parameters, such as leaf area index and downward shortwave radiation. Our findings advocate for a paradigm shift from model-based to data-driven approaches in correcting thermal anisotropy, thereby addressing a critical bottleneck in Earth observation.

54 ENVIRONMENTAL SCIENCES↗

Space radiation measurements during the Artemis I lunar mission

Space radiation is a notable hazard for long-duration human spaceflight. Associated risks include cancer, cataracts, degenerative diseases and tissue reactions from large, acute exposures. Space radiation originates from diverse sources, including galactic cosmic rays, trapped-particle (Van Allen) belts5 and solar-particle events. Previous radiation data are from the International Space Station and the Space Shuttle in low-Earth orbit protected by heavy shielding and Earth’s magnetic field and lightly shielded interplanetary robotic probes such as Mars Science Laboratory and Lunar Reconnaissance Orbiter. Limited data from the Apollo missions and ground measurements with substantial caveats are also available. Here we report radiation measurements from the heavily shielded Orion spacecraft on the uncrewed Artemis I lunar mission. At differing shielding locations inside the vehicle, a fourfold difference in dose rates was observed during proton-belt passes that are similar to large, reference solar-particle events. Interplanetary cosmic-ray dose equivalent rates in Orion were as much as 60% lower than previous observations. Furthermore, a change in orientation of the spacecraft during the proton-belt transit resulted in a reduction of radiation dose rates of around 50%. These measurements validate the Orion for future crewed exploration and inform future human spaceflight mission design.

79 ASTRONOMY AND ASTROPHYSICS↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Intercomparison of Deep Learning Model Architectures for Atmospheric River Prediction

With a rapid surge in the application of machine learning (ML) for a diverse range of tasks in climate science, the present study addresses a challenge for climate scientists when selecting the optimal ML or deep learning (DL) architecture for a given application. In particular, a DL intercomparison study was performed with a focus on forecasting the position of atmospheric rivers (ARs) on short-range time scales (up to 5-day lead times). AR predictions from multiple DL architectures, including various types of convolutional autoencoders and a vision transformer (ViT), were compared against ECMWF ERA5 reanalysis and hindcasts from a global climate model. DL models with similar trainable parameters were trained on ERA5 reanalysis data and AR positions derived from a thresholding algorithm to ensure a fair comparison among the DL models. Each model’s performance and accuracy in forecasting AR location and key input fields within a 5-day window were assessed using metrics of root-mean-square error, anomaly correlation, and mean intersection over union. The ViT architecture outperformed other autoencoder models in most of the metrics. Incorporating additional meteorological fields only yielded slight improvements in forecasting certain fields at longer lead times. The results also suggest that a smaller number of input time steps or smaller number of autoregressive steps can achieve better prediction skills, while also improving the overall computational efficiency. This research offers valuable insights into the strengths and weaknesses of different DL techniques for AR forecasting, hopefully guiding the development of improved models for forecasting this phenomenon.

54 ENVIRONMENTAL SCIENCES↗

Privacy Preservation from High-Performance Computing to Autonomous Science [Industrial and Governmental Activities]

High-Performance Computing (HPC) and Leadership-Class Supercomputing are driving forces behind scientific advancements, enabling researchers to tackle complex challenges in physics, chemistry, biology, and engineering. These systems power vast simulations and data analyses, fueling discoveries in fields ranging from materials science to climate modeling. However, their use often involves processing sensitive data—such as proprietary industry simulations, biomedical records, and national security computations—posing significant privacy concerns. In conclusion, this issue is amplified in collaborative environments like Department of Energy (DOE) user facilities, where HPC resources are shared across institutions to foster innovation.

Kotevska, Olivera [Oak Ridge National Laboratory (↗

“One Table to Rule Them All”: How a Single Table can Enable Extensive Insights, Analytics and Assessment on Human Mobility Data

While much research has been conducted in Human Mobility Science, most studies on the analytics/insights part generally focus on one of the following: processing and analytics on human stop-trip behavior, design of individual mobility metrics (often in silos), calculation and characterization of only a handful (typically 5-6) of human mobility metrics on geospatial-temporal human mobility data of interest. Although human mobility research offers a vast and diverse array of available metrics, most individual studies typically compute only a small subset of five or six metrics at a time when analyzing trajectory datasets of human mobility across different areas of interest. This paper is motivated by the critical need to repeatedly compute an extensive array of human mobility metrics across several trajectory datasets and perform individual metric-level benchmarking to establish a new, standardized Test and Evaluation (T&E) suite for the field of Human Mobility Science. We first present our findings on the minimal yet sufficient pre-processing required to reliably and efficiently compute a wide range of human mobility metrics. The key findings are specifically related to the proposed Composite Stop Locations table, which serves as a core pre-processing data layer. Subsequently, we present a case study demonstrating how the Composite Stop Locations table facilitates computation of at least 14 distinct human mobility metrics (unlike 5-6 different set of metrics used for studies in the literature) using the popular and open-source OpenPFLOW dataset. Finally, we have also presented an example of our benchmarking methodology to evaluate the quality and performance of the trajectory dataset of interest, assessed across multiple human mobility metrics.

De, Debraj [ORNL] (ORCID:0000000233630020)↗

Quantifying UAS Observation Error Variance Used in Data Assimilation Systems and Its Impact on Predictive Skill

Observation error determines the weights of the observations and background state used in data assimilation to generate analyses. Quantifying observation error is critical for the optimal assimilation of observational data sets. Uncrewed Aircraft System (UAS) observations have shown potential benefits in filling observational gaps in the lower atmosphere; however, characterization of their error characteristics has been limited. To optimize the use of UAS observations in numerical weather prediction, UAS observation error is estimated based on the 3‐cornered hat diagnostic approach which uses three independent estimates of the atmospheric state. This approach is applied to data from the 2018 Lower Atmospheric Profiling Studies at Elevation‐a Remotely‐piloted Aircraft Team Experiment field campaign using collocated UAS and rawinsonde observations along with output from a set of convection‐permitting model simulations. The estimated observation error values for UAS temperature, wind, and relative humidity measurements were found to be only weakly dependent on height AGL with mean values equal to 0.5°C, 0.8 m s −1 , and 3%, respectively. Only the newly estimated observation error for temperature differed from that previously used to assimilate commercial aircraft observations into global models (1.0°C). However, using this reduced temperature observation error produced more accurate mesoscale analyses and forecasts of both terrain‐driven flows and convection initiation generated by colliding outflow boundaries within the San Luis Valley of Colorado.

54 ENVIRONMENTAL SCIENCES↗