Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Data for A Hybrid Biophysical-Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81–0.94) and H (R2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

AI/ML↗

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE↗

Search for the Higgs boson decays to a ρ 0 , ϕ, or K ⁎0 meson and a photon in proton-proton collisions at $\sqrt{s} = 13$ TeV

Three rare decay processes of the Higgs boson to a ρ(770) 0 , Φ(1020), or K ⁎ (892) 0 meson and a photon are searched for using $\sqrt{s} = 13$ TeV proton-proton collision data collected by the CMS experiment at the LHC. Events are selected assuming the mesons decay into a pair of charged pions, a pair of charged kaons, or a charged kaon and pion, respectively. Depending on the Higgs boson production mode, different triggering and reconstruction techniques are adopted. The analyzed data sets correspond to integrated luminosities up to 138 fb -1 , depending on the reconstructed final state. After combining various data sets and categories, no significant excess above the background expectations is observed. Upper limits at 95% confidence level on the Higgs boson branching fractions into ρ(770) 0 $γ$, Φ(1020)$γ$, and K ⁎ (892) 0 are determined to be 3.7 x 10 -4 , 3.0 x 10 -4 , and 3.0 x 10 -4 , respectively. In case of the ρ(770) 0 $γ$ and Φ(1020)$γ$ channels, these are the most stringent experimental limits to date.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data Placement Optimization for ATLAS in a Multi-Tiered Storage System within a Data Center

Scientific experiments and computations, especially in High Energy Physics, are generating and accumulating data at an unprecedented rate. Effectively managing this vast volume of data while ensuring efficient data analysis poses a significant challenge for data centers, which must integrate various storage technologies. This paper proposes addressing this challenge by designing and developing a precise data popularity prediction model utilizing state-of-theart AI/ML techniques. This model is crafted from the analysis of ATLAS data and access patterns. It enables us to migrate infrequently accessed data to more economical storage media, such as tape drives, while storing frequently accessed data on faster yet costlier storage media like HDD or SSD. This strategic approach ensures data is placed optimally into the appropriate storage classes, thereby maximizing storage capacity while minimizing data access latency for end-users. Furthermore, the paper includes a performance evaluation of the prediction model using various key metrics such as F1 score, accuracy, precision and recall. Finally, we present a prototype use case, leveraging real-world file access data to assess the model’s impact on performance.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Digital twin framework for PIP-II linac: AI-driven multi-scale modeling from ion source to 800 MeV

The PIP-II superconducting linac at Fermilab is designed to deliver multi-megawatt proton beams for neutrino physics and other high-intensity applications. To expedite commissioning and enhance operational reliability, we have developed an EPICS-based data flow framework that seamlessly integrates digital twins (DT) with physical twins (PT). These digital twins comprise high-fidelity beam dynamics models or data-driven surrogate models connected to their physical counterparts through real-time diagnostics and advanced machine-learning algorithms.Central to this framework is Linac_Gen, an accelerated simulation tool that incorporates convolutional neural networks, random forests, and genetic algorithms to provide up to a tenfold speedup in optimizing the accelerator geometry model. An EPICS translator layer ensures interoperability by efficiently mapping lattice parameters across diverse simulation platforms.Our EPICS-based framework supports multiple operational modes—monitoring, passive learning, closed-loop control, and online learning—covering the entire machine lifecycle. By leveraging HPC resources and multi-objective optimization techniques, the digital twin enables adaptive trajectory correction, real-time fault detection, and predictive modeling of beam stability. This comprehensive approach paves the way for robust, high-intensity operation and data-driven accelerator R&D at Fermilab.

Pathak, Abhishek [Fermilab]↗

RatXcan: A framework for cross-species integration of genome-wide association and gene expression data

Genome-wide association studies (GWAS) have implicated specific alleles and genes as risk factors for numerous complex traits. However, translating GWAS results into biologically and therapeutically meaningful discoveries remains extremely challenging. Most GWAS results identify noncoding regions of the genome, suggesting that differences in gene regulation are the major driver of trait variability. To better integrate GWAS results with gene regulatory polymorphisms, we previously developed PrediXcan (also known as “transcriptome-wide association studies” orTWAS), which maps SNPs to predicted gene expression using GWAS data. In this study, we developed RatXcan, a framework that extends this methodology to outbred heterogeneous stock (HS) rats. RatXcan accounts for the close familial relationships among HS rats by modeling the relatedness with a random effect that encodes the genetic relatedness. RatXcan also corrects for polygenic-driven inflation because of the equivalence between a relatedness random effect and the infinitesimal polygenic model. To develop RatXcan, we trained transcript predictors for 8,934 genes using reference genotype and expression data from five rat brain regions. We found that the cis genetic architecture of gene expression in both rats and humans was sparse and similar across brain tissues. We tested the association between predicted expression in rats and two example traits (body length and BMI) using phenotype and genotype data from 5,401 densely genotyped HS rats and identified a significant enrichment between the genes associated with rat and human body length and BMI. Thus, RatXcan represents a valuable tool for identifying the relationship between gene expression and phenotypes across species and paves the way to explore shared biological mechanisms of complex traits.

Genetics & Heredity↗

Bridging the Gap: User-Centric Energy Monitoring for Policy-Driven Application Optimization in HPC Data Centers

Application energy optimization in HPC data centers face two critical gaps. Systematic methodologies that connect data center policies to application decisions and accessible monitoring tools that enable data-driven optimization. We address both gaps through two complementary pillars. First, we present a methodology based on extended weighted Energy Delay Product (EDP) to translate data center operational priorities and integrate energy considerations into the energy optimization workflow which starts from continuous monitoring through targeted optimization. Second, we present a user-space monitoring tool, Omnistat, that enables this methodology by providing developers with direct access to actionable energy telemetry. Through deployment on the Frontier supercomputer and case studies exploring performance-energy trade-offs, we show how these pillars help energy as an integral optimization target for developers as active participants in data center efficiency.

Shin, Woong [ORNL] (ORCID:0000000172077814)↗

R&D GREET Battery Carbon Footprint Calculator

The Battery Carbon Footprint (CF) Calculator was developed to help U.S. battery manufacturers meet the carbon footprint reporting requirements of the EU Battery Regulation (EU) 2023/1542. The calculator incorporates several major battery carbon footprint frameworks, including the Joint Research Centre's Rules for the Calculation of the Carbon Footprint of Electric Vehicle Batteries (CFB-EV), RECHARGE's Product Environmental Footprint Category Rules for High Specific Energy Rechargeable Batteries for Mobile Applications (PEFCR), the Catena-X Product Carbon Footprint Rulebook (CX-PCF Rules), Battery Pass's Battery Carbon Footprint: Rules for Calculating the Carbon Footprint of the "Distribution" and "End-of-Life and Recycling" Life Cycle Stages, the Global Battery Alliance's Greenhouse Gas Rulebook: Generic Rules, Version 2.1, and the Ministry of Economy, Trade and Industry's draft Carbon Footprint Calculation Method for Automotive Batteries. The tool pairs these frameworks with foreground data from Argonne's R&D GREET models and integrates user-supplied background data covering battery manufacturing and supply chain activities. By bringing multiple international methodologies together in a single platform, the calculator enables manufacturers to evaluate product carbon footprints, improve data consistency, and prepare for evolving regulatory compliance and global market reporting requirements.

Zhang, Jingyi↗

Merged aerosol size distribution from SMPS and OPC for SAIL

This dataset contains merged aerosol number size distribution data for the Surface Atmosphere Integrated Field Laboratory (SAIL) campaign. The merged size distribution data were constructed by combining measurements from a scanning-mobility particle sizer (SMPS) and an optical particle counter (OPC), covering a size range of 0.01–35 µm. The merging methodology follows the approach described by Hand and Kreidenweis (2002) and Marinescu et al. (2019). All aerosol data from the ARM archive were corrected to standard temperature (273.15 K) and pressure (101.3 kPa).

merged size distribution↗

J/ ψ -hadron correlations at midrapidity in pp collisions at $\sqrt{s}$ = 13 TeV

We report on the measurement of inclusive, non-prompt, and prompt J/ψ-hadron correlations by the ALICE Collaboration at the CERN Large Hadron Collider in pp collisions at a center-of-mass energy of 13 TeV. The correlations are studied at midrapidity (|y| < 0.9) in the transverse momentum ranges p T < 40 GeV/c for the J/ψ and 0.15 < p T < 10 GeV/c and |η| < 0.9 for the associated hadrons. The measurement is based on minimum bias and high multiplicity data samples corresponding to integrated luminosities of L int = 34 nb −1 and L int = 6.9 pb −1 , respectively. In addition, two more data samples are employed, requiring, on top of the minimum bias condition, a threshold on the tower energy of E = 4 and 9 GeV in the ALICE electromagnetic calorimeters, which correspond to integrated luminosities of L int = 0.9 pb −1 and L int = 8.4 pb −1 , respectively. The azimuthally integrated near and away side yields of associated charged hadrons per J/ψ trigger are presented as a function of the J/ψ and associated hadron transverse momentum. The measurements are discussed in comparison to PYTHIA calculations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Integrating very-high-resolution imagery, Sentinel-2 time-series data, and machine learning to map shrub fractional abundance across arid and semi-arid ecosystems in China

Shrub fractional abundance (SFA), the proportion of shrub cover per unit area, serves as a critical indicator of environmental aridity and ecosystem health in arid and semi-arid regions, particularly across the Mongolian steppe. However, large-scale SFA mapping in Mongolian steppe ecosystems remains challenging due to the small crown size of shrubs, their sparse distribution, and spectral overlap with coexisting low vegetation (e.g., grasses and herbs), which hinders accurate detection using coarser-resolution satellite data or traditional field surveys. To address these challenges, we developed a two-step approach that integrates very-high-resolution (VHR) imagery, time-series Sentinel-2 data, and deep learning techniques. First, we generated high-accuracy benchmark maps of individual shrub crowns from 0.5 m VHR imagery by combining manual segmentation with a hybrid deep learning framework (Dino V2 and convolutional neural networks). Second, we used these shrub crown maps as training data to build an XGBoost model for predicting SFA from 20 m Sentinel-2 time-series data, leveraging phenological information to improve estimation. We validated our approach across 70 sites (1km 2 each) in the Inner Mongolia Autonomous Region, which is representative of Mongolian steppe ecosystems. From VHR imagery, we mapped 1.31 million shrub crowns with an accuracy of R 2 = 0.92. Scaling up with Sentinel-2 data yielded regional SFA maps with an R 2 = 0.60. Further SHAP (SHapley Additive exPlanations) analysis on the developed XGBoost model revealed that phenological metrics (particularly observations in early-May, mid-July, and late-September), which distinguish shrub phenology from that of other land cover types (e.g., grasses and bare soil), were the most influential predictors of SFA. Finally, our regional SFA maps uncovered unimodal relationships between shrub distribution and climate variables, peaking at mean annual minimum temperatures near 0 °C and annual precipitation around 200 mm. Collectively, these findings demonstrate how the integration of multi-source remote sensing and machine learning can overcome historical limitations in SFA mapping, enabling accurate, spatially continuous assessments across vast Inner-Mongolian steppe ecosystems. Our framework has the potential to be applied to other steppe ecosystems and dryland ecosystems across the Mongolian steppe and beyond, offering a foundation for improved monitoring and ecological impact assessments in the face of global climate changes.

Arid and semi-arid landscapes↗

Study of WH production through vector boson scattering and extraction of the relative sign of the W and Z couplings to the Higgs boson in proton-proton collisions at $\sqrt{s} = 13$ TeV

A search for the production of a W boson and a Higgs boson through vector boson scattering (VBS) is presented, using CMS data from proton-proton collisions at $\sqrt{s} = 13$ TeV collected from 2016 to 2018. The integrated luminosity of the data sample is 138 fb −1 . Selected events must be consistent with the presence of two jets originating from VBS, the leptonic decay of the W boson to an electron or muon, possibly also through an intermediate τ lepton, and a Higgs boson decaying into a pair of b quarks, reconstructed as either a single merged jet or two resolved jets. A measurement of the process as predicted by the standard model (SM) is performed alongside a study of beyond-the-SM (BSM) scenarios. The SM analysis sets an observed (expected) 95% co­fidence level upper limit of 14.3 (9.9) on the ratio of the measured VBS WH cross section to that expected by the SM. The BSM analysis, conducted within the so-called 𝜅 framework, excludes all scenarios with 𝜆 WZ < 0 that are consistent with current measurements, where 𝜆 WZ = 𝜅 W ∕𝜅 Z and 𝜅 W and 𝜅 Z are the HWW and HZZ coupling mod­fiers, respectively. The significance of the exclusion is beyond 5 standard deviations, and it is consistent with the SM expectation of 𝜆 WZ = 1.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Light Meson Spectroscopy with GlueX and Beyond

The GlueX experiment at Jefferson Lab was specifically designed for precision studies of the light-meson spectrum. For this purpose, a photon beam with energies up to 12 GeV is directed onto a liquid hydrogen target contained within a hermetic detector with near-complete neutral and charged particle coverage. Linear polarization of the photon beam with a maximum around 9 GeV provides additional information about the production process. In 2018, the experiment completed its first phase, recording data with a total integrated luminosity above 400 pb?1. We highlight a selection of results from this world-leading data set with emphasis on the search for light hybrid mesons. In the mean time, the detector underwent significant upgrades and is currently recording data with an even higher luminosity. The future plans of the GlueX experiment to explore the meson spectrum with unprecedented precision are summarized.

Austregesilo, Alexander↗

Artificial Intelligence for Data Center Operations (AIOps): Cooperative Research and Development (Final Report)

High performance computing data centers will increasingly need to rely on automation to keep pace with exascale growth in compute capability and to manage and optimize the data center environment and facility resources. Artificial intelligence and machine learning approaches provide the means to improve HPC data center operational efficiency, by learning historical trends and training models to operate on real-time data collected from both IT and facilities sources. NREL has developed methods of real-time collection, aggregation and streaming of these data in the ESIF HPC Data Center and has collected a significant dataset of relevant metrics across computer systems, racks, environmental, building and utility sources for research into various predictive analytics problems. HPE's Advanced Technology Group (ATG) is doing comprehensive research into exascale monitoring and management for High Performance Computing (HPC) systems (hereinafter HPE's Data Monitoring/ Management Technology). NREL and HPE will collaborate to add Artificial Intelligence (AI) to NREL's real-time data collection/ aggregation/ streaming system and HPE's Data Monitoring/ Management System, with the goal of improving the operational efficiency of NREL's Energy Systems Integration Facility (ESIF) HPC Data Center through data analytics on both historical and real-time data from IT systems and facilities operations. This collaboration will consist of efforts in Data Management, Data Analytics, and AI/ML Optimization for both manual and autonomous intervention in data center operations. This will be a multi-year, multi-staged effort with a goal towards building capabilities for an Advanced Smart Facility, and demonstration of these techniques in the NREL ESIF HPC Data Center.

97 MATHEMATICS AND COMPUTING↗

Preliminary Study on Fine-Grained Power and Energy Measurements on Grace Hopper GH200 with Open-Source Performance Tools

The increasing adoption of tightly integrated, heterogeneous architectures, combined with the slowdown of Moore’s law, has made application power and energy-driven optimizations critical to efficiently use high-performance computing systems. This paper introduces a newly developed open-source toolkit that seamlessly integrates the Linux real-time hardware monitoring program hwmon with the Performance Application Programming Interface and the Score-P performance measurement system, thereby enabling fine-grained power and energy measurements for high-performance computing applications. Our primary target platform is the Wombat test bed, which is a system based on the NVIDIA GH200 superchip. The toolkit can capture transient power peaks with high temporal resolution (50 ms) and, thanks to Score-P integration, can map power metrics to specific code regions, thereby providing actionable information on power-intensive operations and inefficiencies. The toolkit also provides a holistic view of both the power and the energy consumption of the entire GH200 superchip by covering all major components: the Grace CPU, the Hopper GPU, and the I/O subsystem. Experiments that use Locally Self-consistent Multiple Scattering, which is an application for first-principles calculations of materials developed at Oak Ridge National Laboratory, have demonstrated the tool’s ability to identify transient power spikes and uncover opportunities for energy-aware optimizations. Additionally, we introduce a Python-based utility for converting Open Trace Format 2 traces to Parquet format, thus enabling advanced data analysis for numerical integration methods applied to power data for accurate energy profiling.

Hernandez Mendoza, Oscar [ORNL] (ORCID:00000002538↗

Searches for Pair-Produced Multijet Resonances Using Data Scouting in Proton-Proton Collisions at s = 13 TeV

Searches for pair-produced multijet signatures using data corresponding to an integrated luminosity of 128 fb - 1 of proton-proton collisions at s = 13 TeV are presented. A data scouting technique is employed to record events with low jet scalar transverse momentum sum values. The electroweak production of particles predicted in R -parity violating supersymmetric models is probed for the first time with fully hadronic final states. This is the first search for prompt hadronically decaying mass-degenerate higgsinos, and extends current exclusions on R -parity violating top squarks and gluinos.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗