Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

PV Degradation Modeling: Applying Geospatial Workflows with "PVDeg"

Accurate degradation modeling is essential for predicting photovoltaic (PV) module performance, estimating longevity and informing design decisions. With degradation rates varying significantly by location, geospatial analysis is critical for PV and broader applications, such as agrivoltaics, weathering and environmental data analysis. This work presents PVDeg, an open-source tool designed for geospatial degradation analysis. PVDeg integrates meteorological data from global sources, including the National Solar Radiation Database (NSRDB) and Photovoltaic Geographical Information System (PVGIS), with degradation models. The toolkit enables users to customize geospatial workflows by integrating weather data, material parameters, and user-defined Python functions. It facilitates accelerated downloads of NSRDB and PVGIS datasets and optimizes geospatial point selection to preserve data density in regions of interest. Additionally, PVDeg provides a local database for storage and spatial queries, supporting large-scale analyses without the need for high-performance computing (HPC) resources. PVDeg provides a foundational workflow that extends its utility beyond PV applications, enabling researchers to analyze geospatial processes across discipline.

14 SOLAR ENERGY↗

Data and Code for Understanding Generative AI Content with Embedding Models

This repository contains code for the experiments in the paper "Understanding Generative AI Content with Embedding Models". Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully hand-crafting data representations based on domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models -- which are trained to be useful across many contexts -- we demonstrate that simple and well-studied dimensionality-reduction techniques such as Principal Component Analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence (AI).

Vargas, Max [Pacific Northwest National Laboratory↗

Advancing set-conditional set generation: Diffusion models for fast simulation of reconstructed particles

The computational intensity of detector simulation and event reconstruction poses a significant difficulty for data analysis in collider experiments. This challenge inspires the continued development of machine learning techniques to serve as efficient surrogate models. We propose a fast emulation approach that combines simulation and reconstruction. In other words, a neural network generates a set of reconstructed objects conditioned on input particle sets. To make this possible, we advance set-conditional set generation with diffusion models. Using a realistic, generic, and public detector simulation and reconstruction package (COCOA), we show how diffusion models can accurately model the complex spectrum of reconstructed particles inside jets.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Distributed Augmentation, Hypersweeps, and Branch Decomposition of Contour Trees for Scientific Exploration

Contour trees describe the topology of level sets in scalar fields and are widely used in topological data analysis and visualization. A main challenge of utilizing contour trees for large-scale scientific data is their computation at scale using highperformance computing. To address this challenge, recent work has introduced distributed hierarchical contour trees for distributed computation and storage of contour trees. However, effective use of these distributed structures in analysis and visualization requires subsequent computation of geometric properties and branch decomposition to support contour extraction and exploration. In this work, we introduce distributed algorithms for augmentation, hypersweeps, and branch decomposition that enable parallel computation of geometric properties, and support the use of distributed contour trees as query structures for scientific exploration. Finally, we evaluate the parallel performance of these algorithms and apply them to identify and extract important contours for scientific visualization.

97 MATHEMATICS AND COMPUTING↗

TDCOSMO - XVII. New time delays in 22 lensed quasars from optical monitoring with the ESO-VST 2.6m and MPG 2.2m telescopes

We present new time delays, the main ingredient of time delay cosmography, for 22 lensed quasars resulting from high-cadence r-band monitoring on the 2.6 m ESO VLT Survey Telescope and Max-Planck-Gesellschaft 2.2 m telescope. Each lensed quasar was typically monitored for one to four seasons, often shared between the two telescopes to mitigate the interruptions forced by the COVID-19 pandemic. The sample of targets consists of 19 quadruply and 3 doubly imaged quasars, which received a total of 1918 hours of on-sky time split into 21 581 wide-field frames, each 320 seconds long. In a given field, the 5-σ depth of the combined exposures typically reaches the 27th magnitude, while that of single visits is 24.5 mag – similar to the expected depth of the upcoming Vera-Rubin LSST. The fluxes of the different lensed images of the targets were reliably de-blended, providing not only light curves with photometric precision down to the photon noise limit, but also high-resolution models of the targets whose features and astrometry were systematically confirmed in Hubble Space Telescope imaging. This was made possible thanks to a new photometric pipeline, lightcurver, and the forward modelling method STARRED. Finally, the time delays between pairs of curves and their uncertainties were estimated, taking into account the degeneracy due to microlensing, and for the first time the full covariance matrices of the delay pairs are provided. Of note, this survey, with 13 square degrees, has applications beyond that of time delays, such as the study of the structure function of the multiple high-redshift quasars present in the footprint at a new high in terms of both depth and frequency. The reduced images will be available through the European Southern Observatory Science Portal.Key words: methods: data analysis / surveys / distance scale

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A model to assess Zircaloy’s mechanical property changes following a transient beyond critical heat flux

Maintaining the integrity of nuclear fuel rods is essential for ensuring public health and safety in nuclear power generation. During reactor operation, this integrity is confirmed by demonstrating compliance with established regulatory acceptance criteria. For moderate-frequency events, such as limiting transients and anticipated operational occurrences (AOOs), the current fuel integrity criterion is based on preventing boiling transition. This criterion assumes that prevention of boiling transition will prevent excessive cladding heating and, thus, fuel failure during normal operations. While conservative, this approach places significant constraints on core design, fuel cycle economics, and a plant’s ability to perform major power uprates, leading to suboptimal fuel utilization and inefficient carbon-free energy production. A more efficient approach could be achieved by revising the failure criterion to a material-specific limit rather than strictly preventing the boiling transition, since boiling transition per se is not a cause of fuel cladding failure. Here, as a result, a new licensing framework based on material properties, termed time-at-temperature (t@T), is needed. This approach would allow for brief periods of post–critical heat flux operation during an AOO without compromising safety. Implementing the t@T licensing strategy requires a robust technical foundation in material properties, which must be established through comprehensive data collection on both unirradiated and irradiated fuel and cladding materials. This foundation would enable the development of a safety basis that ensures safe operation while providing greater flexibility and efficiency for reactor operation. This paper documents a thorough review of the available data to establish a baseline knowledge that can inform the development of cladding mechanical models, as well as identify experimental data gaps that need to be addressed in future research. Machine learning and data informatics were utilized to extract the importance of parameters on the t@T parameter. Industry tools were used to perform baseline analyses to define the relevant transient conditions for data analysis. The subsequent review successfully identified applicable experimental data, as well as sufficient data to evaluate changes in cladding mechanical properties following an AOO transient. Rather than developing new models, this work coupled existing irradiation annealing and recrystallization models to calculate changes in hardness, yield stress, and ultimate tensile stress following an AOO event. The findings from this review were summarized to highlight the experimental data needs required to fill remaining gaps and support the development of future t@T licensing methodologies.

Cladding performance↗

High-throughput micro-scale bandgap mapping for perovskite-inspired materials with complex composition space

Abstract To realize the full promise of high-throughput experimental workflows, the rate of sample synthesis must be matched by that of characterization. Of growing interest are contactless optical techniques that can rapidly measure material homogeneity and properties. Here, we present a hyperspectral imaging method to measure local optical bandgap distributions within samples, utilizing spatially-resolved reflectance spectra coupled with automated data analysis. We collect approximately one million optical bandgap data across the compositional space of Cs 3 (Bi x Sb 1-x ) 2 (Br y I 1-y ) 9 perovskite-inspired materials. Our results show non-monotonic bandgap variations (i.e., bandgap bowing) along six composition gradient sequences, in addition to identifying samples with multiple bandgaps in statistics. High-throughput transient absorption spectroscopy reveals that within these compositions, the depletion of the ground state carriers to excited states occurred at discrete energy levels with independent carrier dynamics, consistent with the bandgap observation and indicative of phase separation. This work demonstrates the potential for rapid optical measurements to assess material quality and homogeneity in a high-throughput experimental setting, supporting screening and recipe optimization of optoelectronic material candidates with desired carrier dynamics and optical properties.

Science & Technology - Other Topics↗

Preliminary Study on Fine-Grained Power and Energy Measurements on Grace Hopper GH200 with Open-Source Performance Tools

The increasing adoption of tightly integrated, heterogeneous architectures, combined with the slowdown of Moore’s law, has made application power and energy-driven optimizations critical to efficiently use high-performance computing systems. This paper introduces a newly developed open-source toolkit that seamlessly integrates the Linux real-time hardware monitoring program hwmon with the Performance Application Programming Interface and the Score-P performance measurement system, thereby enabling fine-grained power and energy measurements for high-performance computing applications. Our primary target platform is the Wombat test bed, which is a system based on the NVIDIA GH200 superchip. The toolkit can capture transient power peaks with high temporal resolution (50 ms) and, thanks to Score-P integration, can map power metrics to specific code regions, thereby providing actionable information on power-intensive operations and inefficiencies. The toolkit also provides a holistic view of both the power and the energy consumption of the entire GH200 superchip by covering all major components: the Grace CPU, the Hopper GPU, and the I/O subsystem. Experiments that use Locally Self-consistent Multiple Scattering, which is an application for first-principles calculations of materials developed at Oak Ridge National Laboratory, have demonstrated the tool’s ability to identify transient power spikes and uncover opportunities for energy-aware optimizations. Additionally, we introduce a Python-based utility for converting Open Trace Format 2 traces to Parquet format, thus enabling advanced data analysis for numerical integration methods applied to power data for accurate energy profiling.

Hernandez Mendoza, Oscar [ORNL] (ORCID:00000002538↗

Automated qualification data tool for high temperature metallic materials

This report describes a framework for storing, processing, and displaying qualification data for high temperature mechanical properties. The framework automates the process of generating design data from mechanical test results, for example for a data qualification report for the ASME Boiler \& Pressure Vessel Code. The framework has three parts: a data storage model with common formats for several types of typical mechanical property tests, a backend based on the \pycreep Python library for correlating and extrapolating the data to generate design material properties and allowable stresses, and a demonstration user interface for displaying, sorting, and filtering the data and exploring different options for modeling the design mechanical properties. The report discusses the options available for data processing, with illustrations from real test data on Alloy 617, Alloy 709, Alloy 740H, and Laser-Powder Bed Fusion 316H. The framework is complete for ASME type data analysis and will be used to store test data generated by the Department of Energy, Office of Nuclear Energy, Advanced Materials and Manufacturing Technologies sponsored qualification programs. Future work could extend the tool to other types of material properties and/or expand the demo user interface to make it accessible across the AMMT program.

36 MATERIALS SCIENCE↗

Computational epidemiological tools for pandemic analysis, understanding, and response

This suite of software tools is being developed to enhance and analyze computational epidemiological models that incorporate realistic disease dynamics and human behavior, with the goal of supporting epidemic and pandemic response. Specifically, the tools enable data analysis, feature extraction, data synthesis, machine learning model development, and prediction of key public health outcomes, such as cases, hospitalizations, deaths, and behavioral responses, for airborne infectious diseases like COVID-19 and influenza.

Butts, David↗

Data & Code from Phoenix CPPP Phase 2 Analysis

This data and code package supports the analysis presented in “Beyond Surface Cooling: Comprehensive Field Assessment of Reflective Pavement Thermal Performance in Phoenix, Arizona” and provides fully reproducible workflows for evaluating the thermal performance of cool pavement treatments in a hot urban environment. The dataset integrates multi-modal field measurements collected across residential and nonresidential settings, including mobile air temperature traverses, stationary air temperature monitoring, residential mean radiant temperature (MRT) measurements, subsurface temperature profiles, and controlled testbed observations. The data package contains raw and processed datasets in comma-separated value (CSV) format, accompanying metadata files describing site characteristics and measurement protocols, and R scripts (.R files) used for data cleaning, time synchronization, spatial and temporal matching, quality control filtering, statistical comparison, and figure generation. All analyses were conducted using R (version ≥ 4.2.0) with commonly available packages (e.g., tidyverse, lubridate, data.table, ggplot2). No proprietary software is required to reproduce results. Field campaigns were designed to quantify the effects of high-reflectance pavement coatings on surface temperature, near-surface air temperature, subsurface heat propagation, and radiative heat exposure. Temporal alignment procedures include standardized timestamp conversion and nearest-neighbor matching of high-frequency sensor measurements to stop-based metadata within defined tolerance windows to ensure comparability across instruments. The workflows generate summary statistics, treatment–control contrasts, depth-dependent thermal gradients, and time-series visualizations used in the associated publication. By integrating mobile, stationary, radiative, and subsurface measurements within a unified and transparent processing framework, this package enables comprehensive evaluation of cool pavement performance across multiple thermal exposure pathways and supports reuse in future urban heat mitigation and climate resilience studies.

AIR TEMPERATURE↗

SPRUCE Whole Ecosystem Warming (WEW) Environmental Data and Water Table Summaries, Marcell Experimental Forest, Minnesota, 2015-2024

This data set contains observations of photosynthetically active radiation (PAR), precipitation, soil temperature, soil volumetric water content, air temperature, relative humidity, and normalized water table depth that are summarized on a daily, weekly, monthly, and annual basis for each of the SPRUCE plots. Observations span 2015-2024. This dataset draws on several datasets (Hanson et al. 2016; Hanson et al. 2020; and Warren, unpublished data) and compiles these environmental observations into useful formats for data analysis. These environmental metrics can be used to understand the environmental conditions inside SPRUCE environmental chambers throughout the durations of the experiment and can be paired with other data for modeling and analysis. R code used to generate these files is provided as part of the data package. This dataset contains four data files in comma separate (.csv) format and a compressed folder (*.zip) containing three R (*.r) scripts. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format. User note: Users must cite the original dataset/s along with this dataset when publishing any analyses using this dataset. Details on the dataset used to compile each variable are available in the header row of the files and in the user guide.

air temperature↗

Machine Learning-Driven Reliability Estimation of PV Inverters Considering Alert-Ambient Variability

Weather-induced spatio-temporal degradation limits outdoor PV inverter lifetime and reliability, necessitating advanced data analysis. This study employs a top-down, data-driven approach utilizing multiple machine learning (ML) algorithms to estimate inverter reliability in a 1.4 MW PV power plant, considering factors such as irradiance, humidity, temperature, time of day, and weather conditions. An extensive alert dataset from 17 identical inverters, including alert types, propagation, and frequency, reveals significant correlations with environmental factors and inverter output power, enabling the construction of a performance reliability model. Dual-stage supervised-ML models are evaluated for accuracy, with the ‘classification-regression’ model by an artificial neural network (ANN) tested on the averaged “Alert-Ambient” dataset, which is outperformed by ‘clustering-regression’ models using random forest (RF) and K-Nearest Neighbors (KNN) on individual inverter datasets. K-means clustering applies principal component analysis to reduce dimensions, achieving improved accuracy beyond the 80% achieved by ANN on the averaged dataset. Second-stage regression estimates inverter reliability with a mean square error of 0.0195 on the averaged dataset and as low as 0.002 on individual inverter datasets using RF. Furthermore, these findings highlight the method's suitability for estimating PV inverter output reliability under ambient conditions, essential for digital twin development and related applications.

14 SOLAR ENERGY↗

Outdoor Deployment Data for a Four-Terminal GaAs//Si Tandem Solar Mini-Module

This dataset contains the complete outdoor measurement and analysis data for a mechanically stacked, four-terminal (4T) gallium arsenide (GaAs)//silicon (Si) tandem solar mini-module deployed from October 2019 to January 2021 at the Solar Radiation Research Laboratory (SRRL) in Golden, Colorado, USA. The data support a performance modeling and degradation analysis framework for tandem photovoltaic devices, as described in the accompanying publication. The dataset includes: (1) current–voltage (J–V) characteristics of each sub-cell measured approximately every five minutes, with extracted performance parameters; (2) spectral irradiance from an EKO MS-710 WISER spectroradiometer, along with derived spectral mismatch ratios (SMR) and average photon energy (APE); (3) one-minute resolution meteorological data from the co-located SRRL weather station and GPS-derived precipitable water vapor (PWV); (4) pre-deployment laboratory characterization (external quantum efficiency, J–V curves, standard test conditions parameters); (5) outdoor-extracted temperature and PWV correction coefficients; and (6) PVcircuit equivalent-circuit simulation outputs used for model validation. Degradation rates of −4.1 ± 0.2 %/year (GaAs) and −2.5 ± 0.9 %/year (Si) were determined using a filtering and normalization methodology adapted for fixed-tilt tandem modules. All data are provided in open, portable formats (Apache Parquet, CSV, JSON) to enable full reproducibility of the published analysis.

14 SOLAR ENERGY↗

Portable Distortion Free Large Solid Angle Coverage Detector

Neutron single crystal diffractometers require large solid angle coverage for optimum performance. This can be achieved by tiling flat detectors in a cylindrical or spherical geometry, but results in large gaps in detector coverage and parallax distortion. The detector edges exhibit degraded resolution, distortion, and gamma rejection. Since the detector edge regions are a significant fraction of the detector active area, they must be removed from the experimental data set, requiring extra beam time to collect enough analysis data. A spherical detector with a continuous surface would effectively address this issue while eliminating most boundary ‘dead’ areas. Here we report on the development of a novel hemispherical shaped neutron detector using seamlessly tiled readout modules to form the desired shape. The heart of the detector is a specially developed curved neutron scintillator coupled to high resolution silicon photomultiplier (SiPM) Anger cameras via custom made fiber optic tapers (FOTs). The detector has been assembled and initial tests have been conducted at the High Flux Isotope Reactor (HFIR) beamlines at Oak Ridge National Laboratory (ORNL). Here, in this work, we describe details of the scintillator design, fabrication and characterization, evaluation of individual detector modules, the details of the detector design implementation, and evaluation of the assembled detector at ORNL beamlines.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Pando

SAND2025-02006O Pando is a distributed data analysis software tool. It is designed to handle large-scale graph analysis problems, often with a specific focus on blockchain/cryptocurrency data. Pando handles scalability by running on a distributed cluster of servers. Users can customize the output using the program’s plugin/extension design methodology. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Gabert, Kasimir↗

Probabilistic data fusion and physics-informed machine learning: A new paradigm for modeling under uncertainty, and its application to accelerating the discovery of new materials

In this report we summarize the work conducted by PI Perdikaris and his group under this Early Career project DE–SC0019116 during the period of 09/01/2018 – 08/31/2023. The central aim of the work was to introduce a new paradigm for scientific data analysis that can seamlessly synthesize rigorous mathematical modeling with data of variable fidelity (e.g., measurements at multiple scales/resolutions or predictions of variable fidelity models) and multiple modalities (e.g., images, time–series, or scattered measurements). The setting we are interested in involves complex systems that are partially observed and whose dynamical behavior could be hard to model or totally unknown. The inherent uncertainty associated with this setting necessitates a departure from the classical deterministic realm of modeling and scientific computation, and, consequently, our main building blocks can no longer be crisp deterministic numbers and governing laws, but instead we must operate with probabilistic models.

97 MATHEMATICS AND COMPUTING↗