Search NASA⌕ Search

SEARCH · Search NASA

Results for “data harmonization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Harmonic analysis of discrete tracers of large-scale structure

It is commonplace in cosmology to analyze fields projected onto the celestial sphere, and in particular density fields that are defined by a set of points e.g. galaxies. When performing an harmonic-space analysis of such data (e.g. an angular power spectrum) using a pixelized map one has to deal with aliasing of small-scale power and pixel window functions. We compare and contrast the approaches to this problem taken in the cosmic microwave background and large-scale structure communities, and advocate for a direct approach that avoids pixelization. We describe a method for performing a pseudo-spectrum analysis of a galaxy data set and show that it can be implemented efficiently using well-known algorithms for special functions that are suited to acceleration by graphics processing units (GPUs). The method returns the same spectra as the more traditional map-based approach if in the latter the number of pixels is taken to be sufficiently large and the mask is well sampled. The method is readily generalizable to cross-spectra and higher-order functions. It also provides a convenient route for distributing the information in a galaxy catalog directly in harmonic space, as a complement to releasing the configuration-space positions and weights, and a route to spectral apodization. Finally, we make public a code enabling the application of our method to existing and upcoming datasets.

79 ASTRONOMY AND ASTROPHYSICS↗

Constraining gravity with a new precision 𝐸 𝐺 estimator using Planck + SDSS BOSS data

The 𝐸 𝐺 statistic is a discriminating probe of gravity developed to test the prediction of general relativity (GR) for the relation between gravitational potential and clustering on the largest scales in the observable Universe. We present a novel high-precision estimator for the 𝐸 𝐺 statistic using CMB lensing and galaxy clustering correlations that carefully matches the effective redshifts across the different measurement components to minimize corrections. A suite of detailed tests is performed to characterize the estimator’s accuracy, its sensitivity to assumptions and analysis choices, and the non-Gaussianity of the estimator’s uncertainty is characterized. After finalization of the estimator, it is applied to Planck CMB lensing and SDSS CMASS and LOWZ galaxy data. We report the first harmonic space measurement of 𝐸 𝐺 using the LOWZ sample and CMB lensing and also updated constraints using the final CMASS sample and the latest Planck CMB lensing map. We find $\hat{𝐸}$$^{Planck+CMASS}_{𝐺}$ = 0.3⁢6$^{+0.06}_{−0.05}$⁢(68.27%) and $\hat{𝐸}$$^{Planck+LOWZ}_{𝐺}$ = 0.4⁢0$^{+0.11}_{−0.09}$⁢(68.27%), with additional subdominant systematic error budget estimates of 2% and 3%, respectively. Using Ω m,0 constraints from Planck and SDSS BAO observations, Λ⁢CDM-GR predicts 𝐸$^{GR}_ {𝐺}$⁡(𝑧 =0.555) = 0.401 ± 0.005 and 𝐸$^{GR}_{𝐺}$⁡(𝑧 =0.316) = 0.452 ± 0.005 at the effective redshifts of the CMASS and LOWZ based measurements. We report the measurement to be in good statistical agreement with the Λ⁢CDM-GR prediction and report that the measurement is also consistent with the more general GR prediction of scale independence for 𝐸 𝐺 . Furthermore, this work provides a carefully constructed and calibrated statistic with which 𝐸 𝐺 measurements can be confidently and accurately obtained with upcoming survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

Announcing the Biomedical Data Translator: Initial Public Release

ABSTRACT The growing availability of biomedical data offers vast potential to improve human health, but the complexity and lack of integration of these datasets often limit their utility. To address this, the Biomedical Data Translator Consortium has developed an open‐source knowledge graph–based system—Translator—designed to integrate, harmonize, and make inferences over diverse biomedical data sources. We announce here Translator's initial public release and provide an overview of its architecture, standards, user interface, and core features. Translator employs a scalable, federated, knowledge graph framework for the integration of clinical, genomic, pharmacological, and other biomedical knowledge sources, enabling query retrieval, inference, and hypothesis generation. Translator's user interface is designed to support the exploration of knowledge relationships and the generation of insights, without requiring deep technical expertise and gradually revealing more detailed evidence, provenance, and confidence information, as needed by a given user. To demonstrate Translator's application and impact, we highlight features of the user interface in the context of three real‐world use cases: suggesting potential therapeutics for patients with rare disease; explaining the mechanism of action of a pipeline drug; and screening and validating drug candidates in a model organism. We discuss strengths and limitations of reasoning within a largely federated system and the need for rich concept modeling and deep provenance tracking. Finally, we outline future directions for enhancing Translator's functionality and expanding its data sources. Translator represents a significant step forward in making complex biomedical knowledge more accessible and actionable, aiming to accelerate translational research and improve patient care.

Research & Experimental Medicine↗

Transformation rate maps of dissolved organic carbon in the contiguous US

Riverine dissolved organic carbon (DOC) plays a vital role in regional and global carbon cycles. However, the processes of DOC conversion from soil organic carbon (SOC) and leaching into rivers are insufficiently understood, inconsistently represented, and poorly parameterized, particularly in land surface and Earth system models. As a first attempt to fill this gap, we propose a generic formula that directly connects SOC concentration with DOC concentration in headwater streams, where a single parameter, the transformation rate from SOC in the soil to DOC leaching flux (P r ), accounts for the overall processes governing SOC conversion to DOC and leaching from soils (along with runoff) into headwater streams. We then derive high-resolution P r maps over the contiguous US (CONUS) using SOC data from two different sources: the Harmonized World Soil Database v1.2 (HWSD) and SoilGrids 2.0. Both maps are developed following the same five major steps: (1) selecting independent catchments where observed riverine DOC data are available with reasonable quality; (2) estimating catchment-average SOC for the independent catchments; (3) estimating the P r values for these catchments based on the generic formula and catchment-average SOC; (4) developing a predictive model of P r with machine learning (ML) techniques and catchment-scale climate, hydrology, geology, and other attributes; and (5) deriving a national map of P r based on the ML model. For evaluation, we compare the DOC concentration derived using the P r map and the observed DOC concentration values at evaluation catchments. The resulting mean absolute scaled error and coefficient of determination are 0.73 and 0.47 for the HWSD-based model and 0.58 and 0.72 for the SoilGrids-based model, respectively, suggesting the effectiveness of the overall methodology. Efforts to constrain uncertainty and evaluate sensitivity of P r to different factors are discussed. To illustrate the use of such maps, we derive a riverine DOC concentration reanalysis dataset over CONUS. The two P r maps, robustly derived and empirically validated, lay a critical cornerstone for better simulating the terrestrial carbon cycle in land surface and Earth system models. Our findings not only set a foundation for improving our predictive understanding of the terrestrial carbon cycle at the regional and global scales, but also hold promises for informing policy decisions related to decarbonization and climate change mitigation. The data presented in this study are publicly available at https://doi.org/10.5281/zenodo.14563816 (Li et al., 2024).

54 ENVIRONMENTAL SCIENCES↗

Short and medium range structure in elastic deformation of metallic and covalent glasses

Here, we present a concise methodology to analyze structural response to the applied stress in amorphous solids, including metallic glasses (MG), glassy selenium, silica and polycarbonate, using high energy x-ray diffraction and atomic pair distribution function (PDF) analysis. To assess the structural anisotropy induced by applied axial stress, diffraction data were expanded into spherical harmonics. Using Bessel transformation, components of the structure function were converted into isotropic and anisotropic PDFs. The PDFs were compared to the expected model behavior for ideal elastic deformation to separate homogeneous affine strain from local non-affine strains. In metallic glass the range of non-affine deformation is limited to the nearest neighbor shell, suggesting local strain relaxation under stress that occurs even in the elastic regime. Beyond the second atomic shell strain is uniform. However, in glassy silica, polycarbonate and selenium strong local bonding inhibits local displacements and strain in short range order is accommodated by rotation of local units. Interestingly, beyond a molecular unit, deformation in covalent systems is similar to MG, and response of the medium range order scales with the macroscopic stress.

glassy structure↗

High Resolution Siting Suitability of Various Power Plant Technologies

Energy sector planning models determine the aggregate need for new generation, but these models are typically at the state or regional scale and are not equipped to address the wide range of location- and technology-specific issues that are increasingly a factor in power plant siting. These animations demonstrate the aggregate siting suitability of various power plant technology configurations, considering technology-specific factors that can prohibit development. The data presented is from the GRIDCERF (Geospatial Raster Input Data for Capacity Expansion Regional Feasibility) data package. GRIDCERF is a harmonized, open-source geospatial product that can be used to evaluate siting suitability for renewable and non-renewable power plants in the conterminous United States. The animations presented here demonstrate a curated selection of the full suite of technology configurations available. GRIDCERF provides the necessary inputs for models that simulate power plant siting for regional capacity expansion planning such as the Capacity Expansion Regional Feasibility (CERF) model.

Mongird, Kendall [Pacific Northwest National Labor↗

Quantifying Microstructure Variability in Laser Powder Bed Fusion 316 L Stainless Steel Microstructures with Spatial Statistics

Here, we have explored data-driven methods for material microstructure quantification that improve sensitivity to microstructural changes compared to traditional approaches. The methods integrate multiple microstructural properties, including grain morphology, crystallographic orientation, and material phase information. The simpler method employs maps of the Euclidean distance transformation metric to evaluate the morphology of grain boundary networks. The more intensive approach employs generalized spherical harmonic mapping for crystallographic orientations, per-pixel phase information, and a variational auto-encoder for dimensionality reduction and results in a multidimensional clustering of by microstructure similarity. Applied to an experimental dataset of additively manufactured steel, both methods detected slight variations in samples produced under nominally identical processing conditions. Both methods were able to distinguish between samples from multiple (nominally identical) builds, while the generalized spherical harmonics-based method could additionally cluster data samples rotated at two orientations on the build plate. The improved sensitivity of the methods, demonstrated through comparison with traditional microstructure characterization techniques, offers advantages for microstructure quantification and comparisons in advanced manufacturing applications.

SS316L↗

Bayesian reconstruction of anisotropic flow fluctuations at fixed impact parameter

The cumulants of the distribution of anisotropic flow are measured accurately in Pb+Pb collisions at the LHC as a function of centrality classifiers (charged multiplicity and/or transverse energy). Using Bayesian inference, we reconstruct from these measurements the probability distribution of anisotropic flow in the ``theorists' frame'' where the impact parameter has a fixed magnitude and orientation, up to ∼70% centrality. The variation of flow fluctuations with impact parameter displays direct evidence of viscous damping, which is larger for higher Fourier harmonics, in line with expectations from hydrodynamics. We use intensive measures of non-Gaussian flow fluctuations, which have reduced dependence on centrality. Here, we infer from ATLAS data the magnitude of these intensive non-Gaussianities in each Fourier harmonic. They provide data-driven estimates of response coefficients to initial anisotropies, without resorting to any specific microscopic model of initial conditions. These estimates agree with viscous hydrodynamic calculations.

Bayesian methods↗

Data-Driven Analysis of Multipactor Dynamics via Dynamic Mode Decomposition

Multipactor effect is a performance-limiting kinetic plasma effect that can occur in high-power microwave and radio frequency (RF) devices. Multipactor effect is of special concern in vacuum or near-vacuum conditions such as those in particle accelerators and spaceborne devices. In this work, we present a data-driven reduced-order model (ROM) based on dynamic mode decomposition (DMD) for modeling of multipactor effects. We study multipactor effects and the resulting nonlinear harmonic generation by processing high-fidelity data generated from electromagnetic particle-in-cell (EMPIC) simulations using the DMD algorithm. We also investigate time-delay embedding extensions of DMD with improved generalizability and accuracy for modeling the electron plasma current density behavior. Here, the results show that DMD provides valuable insights into multipactor phenomena by extracting relevant modal spatiotemporal patterns and frequencies. In addition, DMD offers the potential to time extrapolate EMPIC simulations at a minimal cost, thereby reducing overall simulation time.

43 PARTICLE ACCELERATORS↗

Inverse problem in the large momentum effective theory framework

One proposal to compute parton distributions from first principles is the large momentum effective theory (LaMET), which requires the Fourier transform of matrix elements computed nonperturbatively. Lattice quantum chromodynamics (QCD) provides calculations of these matrix elements over a finite range of Fourier harmonics that are often noisy or unreliable in the largest computed harmonics. It has been suggested that enforcing an exponential decay of the missing harmonics helps alleviate this issue. Using nonperturbative data, we show that the uncertainty introduced by this inverse problem in a realistic setup remains significant without very restrictive assumptions, and that the importance of the exact asymptotic behavior is minimal for values of 𝑥 where the framework is currently applicable. We show that the crux of the inverse problem lies in harmonics of the order of 𝜆 = 𝑧⁢𝑃 𝑧 ∼ 5–15, where the signal in the lattice data is often barely existent in current studies, and the asymptotic behavior is not firmly established. We stress the need for more sophisticated techniques to account for this inverse problem, whether in the LaMET or related frameworks like the short-distance factorization. We also address a misconception that, with available lattice methods, the LaMET framework allows a “direct” computation of the 𝑥-dependence, whereas the alternative short-distance factorization only gives access to moments or fits of the 𝑥-dependence.

Dutrieux, Hervé [Aix-Marseille Université, Marseil↗

Topsoil bulk geochemical compositions - An updated harmonized global dataset

Mineral weathering is a key biogeochemical process because of the capacity of minerals to stabilize organic matter. However, predicting soil weathering status across large spatial areas still isn’t possible due to a lack of global data and theoretical frameworks. To address this knowledge gap, multiple global datasets of bulk topsoil geochemical compositions have been harmonized using R. These datasets document topsoil bulk geochemical compositions across five continents (n = ~16,000 observations). Source data for these observations include the EuroGEOSurveys Geochemical Baseline Database (FOREGS), the US Geological Survey National Geochemical Database (NASGLP), the Geochemical Atlas of Australia (GAA), the US Geological Survey Alaska Geochemical Database (AGD84), the National Cooperative Soil Survey (NCSS), the European Geochemical Mapping of Agricultural Soil (GEMAS), Ecorespira-Amazon (ERA), the New Zealand Geochemical Baseline Survey (NZ_GBS), and the African Soil Information Service (AFSIS). Major elements observed include Aluminum (Al), Calcium (Ca), Iron (Fe), Potassium (K), Magnesium (Mg), Sodium (Na), Titanium (Ti), Manganese (Mn), Phosphorus (P), Carbon (C), and Sulfur (S). This data package includes the harmonized dataset itself, and the R scripts necessary to harmonize these datasets, in addition to metadata that describes all columns, files, and databases used in this project. Methods & Sampling Step 1 – Databases of geochemical data identified This study aimed to leverage existing measurements of topsoil geochemical data. Databases were first identified and deemed appropriate for inclusion if they were measuring soils and performed these measurements on the <2mm soil fraction. Databases such as NCSS and AGD84 needed more post processing to include in the database and this was done using the NCSS_datamerge_031626 R file and Alaska_USGSmerge_031626 R file, respectively. Step 2 – Database harmonization Once appropriate databases were identified, they were harmonized for ease of analysis using the R script Database_Harmonization_031826. This included removing columns from original datasets that would not be used in analysis (removed columns are noted in the code). Then, data cleaning procedures specific to each dataset were undertaken. This includes standardizing columns to include units and adding metadata columns regarding procedures for analyzing specific elements. Functions for standardizing measurements and units are outline in R files: calculate element_mg_kg_031626, calculate_oxide_wt_perc_031626, change_oxide_caps_031626, and conv_2_numeric_031626. This also included adding a unique identifier for each sample to identify it with its respective database (see CD_ID in data dictionary). Geographic information: Data reflect a compilation of datasets collected globally. Geographic areas covered by each of the datasets include: - EuroGEOSurveys Geochemical Baseline Database (FOREGS) - European continent - North American Soil Geochemical Landscapes (NASGLP) - continental United States and limited parts of Canada (see database key for more details) - National Geochemical Survey of Australia (GAA) - Australia - Alaska geochemical database (AGDB4) - Alaska - National Cooperative Soil Survey (NCSS) - Global measurements, but concentrated in the continental United States - Geochemical data for arable land and land under permanent grass cover in continental Europe (GEMAS) - continental Europe - Ecorespira-Amazon (ERA) - Geochemical data from the Amazon basin - Geochemical baseline data for New Zealand (NZGBS) - New Zealand - Geochemical data collected across continental Africa (AfSIS) - Measurements across Africa

EARTH SCIENCE > LAND SURFACE > SOILS↗

Harmonized Database of Western U.S. Water Rights (HarDWR)

From Lisk et al. (2024): "In the arid and semi-arid western U.S., access to water is regulated through a legal system of water rights. Individuals, companies, organizations, municipalities, and tribal entities have documents that declare their water rights. State water regulatory agencies collate and maintain these records, which can be used in legal disputes over access to water. While these records are publicly available data in all western U.S. states, the data have not yet been readily available in digital form from all states. Furthermore, there are many differences in data format, terminology, and definitions between state water regulatory agencies. Here, we have collected water rights data from 11 western U.S. state agencies, harmonized terminology and use definitions, formatted them consistently, and tied them to a western U.S.-wide shapefile of water administrative boundaries. We demonstrate how these data enable consistent regional-scale western U.S. hydrologic and economic modeling."

Economics↗

Investigating the Global Biogeophysical Impact of Area and Mass Based Wood Harvest in a Vegetation Demography Model

Wood harvesting alters land surface properties and energy redistribution, but there is a lack of studies estimating these changes on a global scale. We coupled a vegetation demographic model, the Functionally Assembled Terrestrial Ecosystem Simulator, with the E3SM land model to perform offline model simulation to investigate the land biogeophysical responses, including canopy coverage, leaf area index, albedo, surface roughness length, and energy fluxes, to historical wood harvest on the global scale. In this study, we found 50% less harvested carbon (C) when choosing the area-based harvest rate as driving data that has not been spatially harmonized, compared to reharmonized mass-based harvesting. By considering the uncertainty from reconstruction of historical wood harvest time series and the choice of wood harvest approach in the model, continuous wood harvest (1850–2015) results in 5%–10% of canopy coverage loss, contributing 0.5%–1% increase of albedo over disturbed land, which is much stronger than a non-demographic land surface model. Changes in energy flux from the wood harvest are negligible (<1%), but the responses of land surface properties vary (up to 30%) due to differences in model structure between the single canopy, sun-shade leaf model and vegetation demographic model.

Shu, Shijie [Lawrence Berkeley National Laboratory↗

Harmonized Database of Western U.S. Water Rights (HarDWR) v.1

Abstract In the arid and semi-arid Western U.S., access to water is regulated through a legal system of water rights. Individuals, companies, organizations, municipalities, and tribal entities have documents that declare their water rights. State water regulatory agencies collate and maintain these records, which can be used in legal disputes over access to water. While these records are publicly available data in all Western U.S. states, the data have not yet been readily available in digital form from all states. Furthermore, there are many differences in data format, terminology, and definitions between state water regulatory agencies. Here, we have collected water rights data from 11 Western U.S. state agencies, harmonized terminology and use definitions, formatted them for consistency, and tied them to a Western U.S.-wide shapefile of water administrative boundaries.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Enabling probabilistic learning on manifolds through double diffusion maps

Here, we present a generative learning framework for probabilistic sampling that extends Probabilistic Learning on Manifolds (PLoM), which is designed to generate statistically consistent realizations of a random vector in a finite-dimensional Euclidean space, informed by a (representative) set of observations. In its original form, PLoM constructs a reduced-order probabilistic model by combining three main components: (a) kernel density estimation to approximate the underlying probability measure, (b) Diffusion Maps to characterize the manifold of the data, and (c) a reduced-order Itô Stochastic Differential Equation (ISDE) to sample from the learned distribution. However, its sampling dynamics are posed in the ambient space and the retained number of reduced coordinates is chosen by projection-reconstruction error. In practice, this often (i) requires more coordinates than the data’s intrinsic dimension to achieve stable sampling and (ii) lacks a smooth, basis-independent lifting back to the data domain; moreover, standard Diffusion Maps emphasize harmonic eigenfunctions and can miss non-harmonic latent structure. We address these limitations by decoupling geometry learning from sampling: a first Diffusion Maps pass identifies non-harmonic coordinates on which we formulate a full-order ISDE directly in the latent space, while Double Diffusion Maps captures multiscale geometric features and Geometric Harmonics (GH) learns a smooth lifting map to the ambient variables that is independent of the particular diffusion basis. This hybrid design preserves the system’s dynamical richness with a compact geometric representation and enables principled out-of-sample inference. The effectiveness and robustness of the proposed method are illustrated through two numerical studies: one based on data generated from two-dimensional Hermite polynomial functions and another based on high-fidelity simulations of a detonation wave in a reactive flow.

Double diffusion maps↗

Assessing methods in fusion and fitting for time series construction in remote sensing-based earth observations

This study evaluates the comparative performance of spatiotemporal fusion and time-series fitting methods for constructing high-spatiotemporal-resolution remote sensing time-series data. Due to in-class similarity of fusion methods and fitting methods, we employ the Fit-FC (Fitting, spatial Filtering, and residual Compensation) model as a representative fusion method and the linear harmonic fitting model as a representative fitting method. Both Fit-FC and the linear harmonic fitting are widely used for high-spatiotemporal-resolution time-series data construction, and we modify the original Fit-FC model to enable automatic time-series fusion. To ensure data representativeness, we use 3 years (2019–2021) of Harmonized Landsat and Sentinel-2 surface reflectance datasets and Terra MCD43A4 products. Eight experimental regions are selected worldwide to guarantee generalization of the comparative performance between fusion and fitting methods, covering diverse land-use types (cropland, developed land, forest, and grassland) and varying climatological conditions. Time-series of NDVI and surface reflectance are analyzed under both actual observations and simulated data-missing scenarios. The constructed time-series data reveals that (1) the modified Fit-FC and linear harmonic fitting model achieve excellent performance in constructing high-resolution time-series images; (2) the fusion method outperforms the fitting method in constructing time-series of NDVI and surface reflectance images in cropland-, forest-, and grassland-dominated regions; (3) both methods achieve comparable performance in developed-dominated regions; (4) the fusion method is more robust to missing data, and better captures abrupt phenological transitions under conditions of continuous missing data; (5) the fitting method is computationally more efficient, making it suitable for large-scale time-series image reconstruction. This study provides valuable insights for selecting optimal strategies to generate high-resolution time-series images across diverse application scenarios and lays a foundation for extensions to other vegetation indices or land surface variables.

54 ENVIRONMENTAL SCIENCES↗

From models to reality: a systematic review on simulated and measured residential heat pump energy savings

High-performance HVAC solutions are central to residential energy management. A substantial share of these are electric, reversible-cycle systems, with heat pumps representing the largest portion of current and near-term adoption. This review synthesizes peer-reviewed and grey literature on residential space heating and cooling heat pumps. The academic literature is dominated by modeling (73.8%), with limited field measurement (13.1%). Grey literature from United States serve as a supplemental resource providing measured savings. Conversions from electric-resistance heating consistently show the largest site energy reductions, while oil/propane baselines yield moderate savings, and gas baseline scenario often deliver small and region-dependent savings. This study cross-checks the grey literature measured data with simulation data filtered from the ResStock dataset. The comparison indicates a discrepancy between simulations and measured data: simulated site EUIs are typically lower than measured EUIs, but percentage energy savings fall in similar ranges, implying simulations capture directional effects while underestimating energy use. Factors associated with variability and model–measurement differences include system characterization and control representation (e.g., backup heat engagement, thermostat/setpoint strategies, commissioning/installation quality), occupant behavior, weather normalization, metering scope, and envelope characterization. This paper also outlines the proposed methodology for comparing simulation and measured data for heat pumps. It emphasizes the metrics used for comparison and units harmonization, building characteristics matching, and compact metadata are needed for simulations to match measured data. The proposed methodology is expected to improve the credibility of simulated savings as measured evidence grows.

Yu, Lili↗

Enabling pan-repository reanalysis for big data science of public metabolomics data

Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, here we develop pan-repository universal identifiers and harmonized cross-repository metadata. This ecosystem facilitates discovery by integrating diverse data sources from public repositories including MetaboLights, Metabolomics Workbench, and GNPS/MassIVE. Our approach simplified data handling and unlocks previously inaccessible reanalysis workflows, fostering unmatched research opportunities.

El Abiead, Yasin↗