Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Validation of Modern Nuclear Data Processing in SCALE

The nuclear data (ND) community is continuously developing more accurate, diversified, and comprehensive data for radiation transport modeling to support the nuclear science community. As these community efforts progress, it is crucial that ND processing tools like AMPX (used for SCALE [1] ND) also be developed in parallel to incorporate these new data into transport codes and actually deliver those data to end users. AMPX is a mature, well-tested code that was developed by many people at Oak Ridge National Laboratory (ORNL) over the course of the past few decades. A large portion of the AMPX codebase, however, was outdated, difficult to maintain, and incompatible with modern code development tools. Some of the most important parts of the AMPX code have now been replaced with modern C++ code that can be maintained more cost-effectively and can be tested more rigorously.

AMPX↗

Addressing Limitations of the Endpoint Slippage Analysis

Some rate of oxidation and reduction side-reactions will inevitably coexist in most rechargeable batteries. While parasitic reduction traps electrons, parasitic oxidation donates electrons to the cell’s inventory and may cause temporary capacity gain. Consequently, capacity measurements can provide unreliable information about the total extent of side-reactions occurring in the cell. The most widely used method to determine the rate of both these parasitic processes involves analyzing the slippage of endpoints, which consists in tracking the termination of cell charge and discharge when data is represented along a cumulative capacity axis. Here, we argue that this approach could lead to inaccuracies when applied to certain systems, which includes Si electrodes in Li-ion batteries and hard carbon in Na-ion batteries. This inaccuracy originates from the smooth nature of the voltage profiles of these materials at low and high alkali-ion content, causing the termination of charge and discharge to be dictated by voltage changes at both the positive and negative electrodes. We analyze this issue in quantitative terms and propose equations that can provide true rates of parasitic processes from experimental endpoint slippage data. This work shows that, in battery science, well-established analytical approaches may not be directly transferrable to new electrode systems.

25 ENERGY STORAGE↗

ASCR Workshop Position Paper: Challenges and Opportunities in High Energy Physics

High energy particle physics and cosmology concern themselves with estimating fundamental parameters of nature, such as the masses and interactions of fundamental particles like the Higgs boson and the rate of expansion of the universe. In doing so, they analyze exabyte-scale datasets, some of the largest in all of science, and face many challenges in subsequent data analysis. These challenges are shared between the two disciplines, but we focus on particle physics to highlight one specific domain. In particle physics, the standard method for estimating parameters involves performing Monte Carlo (MC) integration as a function of both parameters of interest and nuisance parameters using an expensive simulator, counting the number of observed collision events (i.i.d. samples) from an experiment in the corresponding integration domains, and forming a Poisson likelihood function. This likelihood function is then used in a Frequentist manner to construct a maximum likelihood point estimate (MLE) and confidence set for the parameters. To sufficiently populate the high-dimensional integration domains, simulators consume billions of CPU-hours annually and produce hundreds of petabytes of intermediate output data. Several techniques have been developed to: optimize definitions of the integration domains so as to be maximally sensitive to a particular subset of parameters, efficiently estimate the integrals, and build robust surrogate models by interpolating between integral evaluations at different parameter points. One can view this whole endeavor as classical Simulation-Based Inference (SBI).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The right conditions for high-precision dynamic temperature and heat capacity measurement via pyrometry and conductivity

The pursuit of accurate bulk temperature T under extreme conditions has been a long-standing goal of the high pressure science community, complicated by a lack of data to inform models. To reach these extremely high-pressure, high-temperature (high P − T) conditions, a combination of dynamic and heated static experiments (e.g., diamond or gem anvil cel experiments) are used. For example, in a diamond anvil cell (DAC) experiment, a sample placed in the DAC is first pressurized. Following pressurization, the sample T is increased either by heating the entire DAC (usually using resistive heating, and limited to ∼1000K) or by applying intense laser power to the sample surfaces. In a dynamic experiment, the process of pressurizing the sample also heats it. In the case of shock physics experiments, such heating is substantial, easily reaching thousands of Kelvin; in our work we have seen T ∼17000K. Most methods of measuring temperature at ambient are not compatible with experiments under these high-pressure, high-temperature conditions: thermocouples break, melt, or have conductivity properties that differ from ambient where they are calibrated; thermometers would melt; both are too slow. As a result most methods are based on non-contact techniques such as x-ray diffraction broadening, neutron scattering, or optical methods. Of these, optical methods using the visible and near-infrared region of the spectrum are the most commonly used as the sources and detectors are readily available. In the case of optical methods the optical depth, and therefore the measurement location, is limited to the surface. When a window or anvil material is used, heat flows from the sample into the window/anvil. Likewise, if the sample undergoes a change in thermodynamic state, such as expansion upon release, different T may be expected. As a result, the surface or apparent temperature T app measurement will differ from the bulk or interior temperature that is desired. This surface measurement must be related to the bulk measurement using thermal transport models and material models. While it is tempting to conclude that one should just use x-ray methods that directly probe the interior, even these methods have been shown to depend on thermal transport and material models. Regardless of the method used to create the high P − T condition, therefore, we must understand the role of thermal transport and material models upon our interpretation of the T measurement, as well as the errors and uncertainties associated with the choice of models used in the analysis. This is a substantial area of research and this paper is by no means a complete survey of the relevant sources of uncertainty. For example, we have yet to begin to address alternate transport models in a detailed manner (e.g., Tan-Ahrens), or the many models that use additional layers to approximate melting, turbulence, or epitaxial phenomena). Likewise, we have not explored the impact upon uncertainty of thermal models that use temperature-dependent thermal transport coefficients, or the wide range of material models that can be applied. Instead, this paper focuses on using one simple model, the Urtiew-Grover model, to understand the sources of error in T measurement so that we may identify how best to focus future research efforts to return the best improvements and avoid working on over-optimizing a single type of measurement. To this end, we work through some of the best and worst case scenarios for T measurement.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

MITgcm-AD v2: Open source tangent linear and adjoint modeling framework for the oceans and atmosphere enabled by the Automatic Differentiation tool Tapenade

The Massachusetts Institute of Technology General Circulation Model (MITgcm) is widely used by the climate science community to simulate planetary atmosphere and ocean circulations. A defining feature of the MITgcm is that it has been developed to be compatible with an algorithmic differentiation (AD) tool, TAF, enabling the generation of tangent-linear and adjoint models. These provide gradient information which enables dynamics-based sensitivity and attribution studies, state and parameter estimation, and rigorous uncertainty quantification. Importantly, gradient information is essential for computing comprehensive sensitivities and performing efficient large-scale data assimilation, ensuring that observations collected from satellites and in-situ measuring instruments can be effectively used to optimize a large uncertain control space. As a result, the MITgcm forms the dynamical core of a key data assimilation product employed by the physical oceanography research community: Estimating the Circulation and Climate of the Ocean (ECCO) state estimate. Although MITgcm and ECCO are used extensively within the research community, the AD tool TAF is proprietary and hence inaccessible to a large proportion of these users. The new version 2 (MITgcm-AD v2) framework introduced here is based on the source-to-source AD tool Tapenade, which has recently been open-sourced. Another feature of Tapenade is that it stores required variables by default (instead of recomputing them) which simplifies the implementation of efficient, AD-compatible code. The framework has been integrated with the MITgcm model’s main branch and is now freely available.

Adjoints↗

The ePIC Simulation Campaign Workflow on the Open Science Grid

The ePIC collaboration is realizing the first experiment of the future Electron-Ion Collider (EIC) at the Brookhaven National Laboratory that will allow for a precision study of the nucleons and the nucleus at the scale of sea quarks and gluons through the study of electron-proton/ion collisions. This paper will discuss the current workflow for running centralized simulation campaigns for ePIC on the Open Science Grid (OSG) infrastructure. This involves monthly releases of ePIC software and container deployments to CVMFS, generation of input datasets in HepMC format according to collaboration-defined policy, using Snakemake in CI/CD for validation and benchmarking, and submitting jobs to the OSG condor scheduler for opportunistic running on available resources. File transfers utilize XrootD, and Rucio is used for data management. The workflow is continuously refined to improve daily throughput (currently 50-100k core hours per day) and minimize job failures. Since May 2023, monthly simulation campaigns employing the workflow have cumulatively used over 20 million core hours on the OSG and produced over 350 TB of simulation data. The campaigns incorporate simulations for the broad science program of the EIC and are actively used for the detector and physics studies in preparation of the Technical Design Report (TDR).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Raw Lidar and Camera Data Synchronized with Precipitation and Present Weather Data

As part of the sensor characterization task of the SMART 2.0 project, this dataset includes raw data from three spinning lidars ([Ouster OS2-128](https://ouster.com/products/scanning-lidar/os2-sensor/), [Velodyne Puck (VLP-16)](https://velodynelidar.com/products/puck/), and [Velodyne Ultra Puck (VLP-32)](https://velodynelidar.com/products/ultra-puck/)), one camera ([Mako G-319](https://www.alliedvision.com/en/camera-selector/detail/mako/g-319/)), and one present weather sensor ([Vaisala FD-70](https://www.vaisala.com/en/products/weather-environmental-sensors/forward-scatter-fd70)). All data were synchronized, with the log start time indicated in the file name (HHMMSS). The data can be filtered by date, log time (HHMMSS), sensor, frame ID, and weather classification. These data were gathered statically at the Argonne Testbed for Multiscale Observational Science (ATMOS). Two target stop signs were placed in view of the sensors to contribute a target for comparing sensor data under different conditions. The weather data for each day are stored in netCDF “.nc” files. The lidar data contain the X, Y, Z, intensity, reflectivity, and ring from Ouster OS2-128 rev6, Velodyne VLP-16, and Velodyne VLP-32 lidars. ![raw lidar image](LiDAR_pointcloud_ATMOS.png)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Inferring building height from footprint morphology data

As cities continue to grow globally, characterizing the built environment is essential to understanding human populations, projecting energy usage, monitoring urban heat island impacts, preventing environmental degradation, and planning for urban development. Buildings are a key component of the built environment and there is currently a lack of data on building height at the global level. Current methodologies for developing building height models that utilize remote sensing are limited in scale due to the high cost of data acquisition. Other approaches that leverage 2D features are restricted based on the volume of ancillary data necessary to infer height. Here, we find, through a series of experiments covering 74.55 million buildings from the United States, France, and Germany, it is possible, with 95% accuracy, to infer building height within 3 m of the true height using footprint morphology data. Our results show that leveraging individual building footprints can lead to accurate building height predictions while not requiring ancillary data, thus making this method applicable wherever building footprints are available. The finding that it is possible to infer building height from footprint data alone provides researchers a new method to leverage in relation to various applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

GRIDCERF - Geospatial Raster Input Data for Capacity Expansion Regional Feasibility

The Geospatial Raster Input Data for Capacity Expansion Regional Feasibility (GRIDCERF) data package is a high-resolution product to evaluate siting suitability for renewable and non-renewable power plants in the conterminous United States. GRIDCERF offers hundreds of individual suitability layers for use with both renewable and non-renewable power plant technology configurations in a harmonized format that can be easily ingested by geospatially-enabled modeling software. It also provides pre-compiled technology-specific suitability layers and allows for user customization to robustly address science objectives when evaluating varying future conditions. GRIDCERF data can be directly used with the CERF (Capacity Expansion Regional Feasibility) model to site power plants at a 1km resolution. GRIDCERF includes composite technology siting suitability raster layers for the following utility scale technology configurations. Note that, in addition to technology sub-types shown below, various cooling types are also included (recirculating, pond, once-through, recirculating-seawater, dry-hybrid, or dry) for various technologies. Biomass Conventional (with or without CCS) IGCC (with or without CCS) Coal Conventional (with or without CCS) IGCC (with or without CCS) Natural Gas Combined-cycle (CC) (with or without CCS) Turbine Geothermal Enhanced Geothermal Systems (EGS) - Class 1 through Class 5 resource potential Nuclear Gen 2 Light Water Reactor (LWR) Gen 3 Small Modular Reactor (SMR) Gen 3 AP1000 Refined Liquids Combined-cycle (CC) (with or without CCS) Turbine Solar Photovoltaic (PV) - for capacity factors in the range of 6-18% Utility-scale Concentrating Solar Power (CSP) - for capacity factors in the range of 24-46% Tower Wind (Onshore) - for capacity factors in the range of 5-50% 80m hub height 100m hub height 120m hub height 140m hub height Wind (Offshore) - for capacity factors in the range of 25-60% 100m hub height 140m hub height 160m hub height

capacity expansion↗

NANT Site - Microwave Radiometer Thermodynamic Retrievals TROPoe / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous brightness temperature measurements at 35 channels observed with a microwave radiometer MP3000A operated by NOAA Physical Sciences Laboratory on Nantucket Island for WFIP3. Additional input data in TROPoe are cloud base height from a collocated ceilometer operated by NOAA GML and temperature, water vapor mixing ratio, and pressure from a sensor attached to the MWR housing. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at Upton, NY.

17 WIND ENERGY↗

2024 roadmap on magnetic microscopy techniques and their applications in materials science

Considering the growing interest in magnetic materials for unconventional computing, data storage, and sensor applications, there is active research not only on material synthesis but also characterisation of their properties. In addition to structural and integral magnetic characterisations, imaging of magnetisation patterns, current distributions and magnetic fields at nano- and microscale is of major importance to understand the material responses and qualify them for specific applications. In this roadmap, we aim to cover a broad portfolio of techniques to perform nano- and microscale magnetic imaging using superconducting quantum interference devices, spin centre and Hall effect magnetometries, scanning probe microscopies, x-ray- and electron-based methods as well as magnetooptics and nanoscale magnetic resonance imaging. The roadmap is aimed as a single access point of information for experts in the field as well as the young generation of students outlining prospects of the development of magnetic imaging technologies for the upcoming decade with a focus on physics, materials science, and chemistry of planar, three-dimensional and geometrically curved objects of different material classes including two-dimensional materials, complex oxides, semi-metals, multiferroics, skyrmions, antiferromagnets, frustrated magnets, magnetic molecules/nanoparticles, ionic conductors, superconductors, spintronic and spinorbitronic materials.

2D materials↗

Legacy Survey of Space and Time Data Preview 1: calibrations dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the calibrations dataset type. These are a collection of calibration datasets such as biases, darks, and flats used to construct the data release. This release contains 496 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 1: raw dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the raw dataset type. These are unprocessed images from the LSST Commissioning Camera. This release contains 16,125 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 1: visit_image dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the visit_image dataset type. These are individual processed and calibrated sky images obtained from a single observation with a single filter. This release contains 15,972 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 1: difference_image dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the difference_image dataset type. These are images created by subtracting a template image from a visit image. This release contains 15,972 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 1: deep_coadd dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the deep_coadd dataset type. These are the combination of multiple processed, calibrated, and background- subtracted images, for a patch of sky, for each of the six filters. This release contains 2,644 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗