Search NASASearch

SEARCH · Search NASA

Results for “Data Summarization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Impacts of Year-to-Year Weather Variability and Inter-Panel Spacing on Crop Yields in a Massachusetts Agrivoltaics System

This presentation summarizes recently published work on agrivoltaic field data and irradiance modeling. The body of agrivoltaic field data is still growing, and crop responses to different solar configurations under different local climates are highly varied. We investigate the impact of adding spacing between adjacent solar panels in a fixed-tilt system to improve light diffusion to crops. For four crops (broccoli, peppers, kale, Swiss chard) grown across 3 years in an agrivoltaic system in Massachusetts, we found that only kale had a linearly increasing trend as the inter-panel spacing increased from 0.6 m to 1.5 m (2 ft to 5 ft). However, there were significant year-to-year differences in the yield of agrivoltaic versus control fields. Agrivoltaic and full sun fields produced equivalent yields in a hot, dry year, whereas the full-sun control beds produced more salable yield for all four crops in a warm, wet year. This demonstrates variability of agricultural outcomes and the need for more multi-year studies to ensure agrivoltaic impacts are not under- or overestimated.

14 SOLAR ENERGY

Report priority gaps in high temperature thermodynamic data (Interim Progress Report)

This interim progress report (Level 4 Milestone Number M4SF-26LL010203023) summarizes research conducted at Lawrence Livermore National Laboratory (LLNL) within the Argillite Host Rock Properties & Processes SF-26LL01020302. Our focus is to assess gaps in data availability and understanding for radionuclide thermodynamics within the context of a “hot repository” concept and expand SUPCRT-NE database development to address higher temperatures needed for a DPC DGR disposal concept. The database is intended to inform the argillite GDSA baseline model. The leading European thermochemical database (Thermochimie) is only applicable to temperatures below 80°C. Thus, a US effort to integrate and expand upon other international thermodynamics database efforts is needed, particular if a “hot repository” concept moves forward.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Size-resolved Eddy-Covariance Particle Flux Measurement during the TRACER Campaign (Final Report)

The main goal of the TRacking Aerosol Convection interactions ExpeRiment (TRACER) campaign was to study aerosol–cloud interactions during deep convection over the Houston area. This project deployed a suite of instrumentation with the aim to (1) quantify turbulent vertical particle fluxes during at DOE-ARM sites, including TRACER, (2) assess hygroscopic growth factors and hygroscopicity parameters of the material driving modal aerosol growth during new particle formation and growth events, (3) derive turbulent aerosol mass fluxes using co-located Doppler LIDAR measurements, and (4) create quality-controlled PI data products to support future research utilizing data collected during the TRACER campaign. This report summarized the main findings from the deployments at two DOE-ARM sites. Briefly, we found that new particle formation may occur aloft, in a residual layer, near the top of the boundary layer. Small grown particles appear later due to downward mixing with daytime turbulence. The species that are responsible for aerosol modal growth had hygroscopicity parameters varying between 0.05 and 0.34. These values systematically depended on the wind sector, suggesting that the chemical composition of the precursors differed. This work demonstrated that lidar retrievals of the elastic backscatter and Doppler velocity can be used to obtain surface number emissions of particles with a diameter greater than 0.53 µm. During TRACER, emission particle number fluxes peaked near ∼ 100 cm−2 s−1. Multiple quality-controlled PI data products that will support future TRACER related science were generated and made publically available.

54 ENVIRONMENTAL SCIENCES

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING

Status Update of Permanent Magnet Radiation Resiliency Studies at CEBAF

The proposed energy upgrade of the Continuous Electron Beam Accelerator Facility (CEBAF) incorporates Fixed-Field Alternating-gradient (FFA) arcs utilizing permanent magnet technology. Given the radiation environment within the CEBAF tunnel enclosure, validating the long-term magnetic stability of these materials is a critical step for the project's technical feasibility. This contribution presents an overview of the ongoing permanent magnet radiation resiliency program at Jefferson Lab. We briefly review the experimental methodology used to monitor demagnetization in situ and summarize the operational experience from the initial data-taking campaign. Furthermore, we discuss the upgrades implemented for the second exposure campaign, currently underway, which aims to refine dose correlation and reduce systematic uncertainties. We report on the general status of the program and the roadmap for certifying permanent magnet optics for the proposed upgrade energies.

Bodenstein, R. [Thomas Jefferson National Accelera

Size-resolved particle and black carbon deposition over the cryosphere (Final Report)

Our project focused on providing observational constraints and investigation of aerosol deposition by performing the first unambiguous direct eddy flux covariance measurements of aerosol dry deposition over the cryosphere. The project aimed to make eddy covariance flux measurements of dry deposition and off-line wet deposition measurements of black carbon containing aerosol plus eddy covariance measurements of size-resolved accumulation mode scattering aerosol over the DOE ARM AMF3 Oliktok Point site in Fall 2020. We aimed to compare our measurements to model parameterizations to constrain uncertainties and systematic bias, and improve deposition parameterizations in models. Due to disruptions in field work from COVID-19, our work was delayed and moved to an alternate site on the North Slope of Alaska. This report summarizes the key findings from the project, including data analysis from previous field projects and from the North Slope of Alaska.

54 ENVIRONMENTAL SCIENCES

CAMEO: A Co-design Architecture for Multi-objective Energy System Optimization (Project Report)

CAMEO (Codesign Architecture for Multi-objective Energy System Optimization) is a modular workflow management framework that abstracts co-design problems as Directed Acyclic Graphs (DAG). The framework employs JSON-based workflow specifications that enable systematic decomposition of complex optimization problems into reusable, interchangeable components including data loaders, scenario generators, optimization solvers, and result summarizers.

97 MATHEMATICS AND COMPUTING

United States Nuclear Data Program Annual Report for Fiscal Year 2024

The US Nuclear Data Program (USNDP) Annual Report for Fiscal Year 2024 (FY24) summarizes the work of USNDP for the period of October 1, 2023 through September 30, 2024, with respect to the Work Plan for FY24 that was prepared in 2022. The Work Plan and Final Report for USNDP are prepared for the DOE Office of Science, Office of Nuclear Physics. The support for the nuclear data activity from sources outside the nuclear data program is described in the staffing table and in Appendix A. This leverage amounts to about 14.8 FTE scientific, to be compared with 13.9 FTEs at USNDP units funded by the DOE Office of Science, Office of Nuclear Physics. Since it is often difficult to separate accomplishments funded by various sources, some of the work reported in the present report was accomplished with nuclear data program support leveraged by other funding.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Towards a RAG-based summarization for the Electron Ion Collider

Abstract The complexity and sheer volume of information — encompassing documents, papers, data, and other resources — from large-scale experiments demand significant time and effort to navigate, making the task of accessing and utilizing these varied forms of information daunting, particularly for new collaborators and early-career scientists.To tackle this issue, a Retrieval Augmented Generation (RAG)-based Summarization AI for EIC (RAGS4EIC) is under development. This AI-Agent not only condenses information but also effectively references relevant responses, offering substantial advantages for collaborators. Our project involves a two-step approach: first, querying a comprehensive vector database containing all pertinent experiment information; second, utilizing a Large Language Model (LLM) to generate concise summaries enriched with citations based on user queries and retrieved data. We describe the evaluation methods that use RAG assessments (RAGAs) scoring mechanisms to assess the effectiveness of responses. Furthermore, we describe the concept of prompt template based instruction-tuning which provides flexibility and accuracy in summarization. Importantly, the implementation relies on LangChain [1], which serves as the foundation of our entire workflow. This integration ensures efficiency and scalability, facilitating smooth deployment and accessibility for various user groups within the Electron Ion Collider (EIC) community. This innovative AI-driven framework not only simplifies the understanding of vast datasets but also encourages collaborative participation, thereby empowering researchers. As a demonstration, a web application has been developed to explain each stage of the RAG Agent development in detail. The application can be accessed athttps://rags4eic-ai4eic.streamlit.app.[A tagged version of the source code can be found inhttps://github.com/ai4eic/EIC-RAG-Project/releases/tag/AI4EIC2023_PROCEEDING.]

Instruments & Instrumentation

Performance Year 1 Technical Report - OPEN COG Grid: Extendable Coherent Models-Datasets for Cognitive Power Grids

The OPEN COG Grid project is a collaborative effort between LLNL, NREL, and Texas A&M University (TAMU) to develop synthetic power system datasets that (i) contain all technical information that would be available in a real system, allowing to conduct studies ranging from dynamic simulation to long term planning studies; ii) are accessible to researchers from the broader data sciences community, as oppossed to power system experts only; and (iii) This report summarizes the work conducted during the first 15 months of execution of the project. These activities encompassed: 1. Conduct a survey of existing open data sets and open source power systems simulators, their supported use cases, and accessibility (Chapter 1). 2. Define a new extensible specification for power system data, covering all parameters necessary for most computational use cases (Chapter 2). 3. Collecting real technical system data to complete missing parameters in existing open source datasets (Chapter 3). 4. Develop models that capture the behavior of emergent actors in power grids, neglected by existing datasets; aggregated residential demand response (Chapter 4) and demand response of cryptocurrency miners (Chapter 5). 5. Collect detailed spatial information on distributed energy resources, particular, solar photovoltaic facilities (Chapter 6). The following chapters provide detailed descriptions of these tasks, the assumptions taken, and their findings. In conducting these tasks, the project team produced: two (accepted) conference papers; one journal paper under submission; one draft journal paper pending submission; released one repository with the developed power system data specification, with documentation and examples; and one extended dataset for the Texas power grid under review for release. The team hopes these contributions will enhance access to power system data and remove barriers to the development of new computational techniques for power systems, particularly, those inspired by cognitive sciences.

24 POWER TRANSMISSION AND DISTRIBUTION

Performance of heavy-flavour jet identification in Lorentz-boosted topologies in proton-proton collisions at √(s) = 13 TeV

Measurements in the highly Lorentz-boosted regime provoke increased interest in probing the Higgs boson properties and in searching for particles beyond the standard model at the LHC. In the CMS Collaboration, various boosted-object tagging algorithms, designed to identify hadronic jets originating from a massive particle decaying to bb̅ or cc̅, have been developed and deployed across a range of physics analyses. This paper highlights their performance on simulated events, and summarizes novel calibration techniques using proton-proton collision data collected at √(s) = 13 TeV during the 2016–2018 LHC data-taking period. Three dedicated methods are used for the calibration in multijet events, leveraging either machine learning techniques, the presence of muons within energetic boosted jets, or the reconstruction of hadronically decaying high-energy Z bosons. The calibration results, obtained through a combination of these approaches, are presented and discussed.

Pattern recognition

FORESTR: Finding, Organizing, Representing, Explaining, Summarizing, and Thinning Random forests

Random forests have become popular models used for data driven predictions. As a result, random forests are currently used or being considered for high-consequence mission applications in national security, such as the prediction of yield from optical signals and malware detection. While random forests may provide accurate predictions, the complexity of the algorithm causes a lack of interpretability. Random forests are an ensemble of regression or decision trees. Individual regression and decision trees are interpretable, but ensembles are inherently difficult to interpret due to the compilation of many models. We aim to increase the interpretability of random forests by finding patterns in the ensemble of trees that can be used to “thin” (or remove) trees. As a starting point, in this report, we develop a new distance metric for quantifying the similarity between trees based on their topologies (i.e., shapes). We base the metric on a novel distance metric for graphs that is a proper mathematical distance, is invariant to transformations, has registration between graphs, and computes topological evolutions between graphs. We use the tree distance metric to compute tree statistics such as a “mean tree” and to identify clusters of trees. We apply the developed methodology to a toy dataset and a mission relevant product inspection dataset to demonstrate how the metric can provide insight into random forests. Furthermore, we discuss the limitations of the approach and ideas for future research into how the metric could be used as a thinning tool to develop less complex models.

97 MATHEMATICS AND COMPUTING

Data-Driven Supervised Dimension Reduction for Scientific Discovery (LDRD QTI Report)

This report summarizes the findings of a four months FY24 Advanced Science & Technology (AS&T) LDRD Quick Targeted Investigation (QTI) project focused on the exploration of supervised dimension reduction approaches based on autoencoders. Autoencoders have been extensively employed in literature for unsupervised learning tasks, however, their use for supervised regression tasks, which are common within scientific applications, has been limited. Motivated by linear dimension reduction strategies like Active Subspaces and Adaptive Basis, we explored the possibility of employing autoencoders to discover a non-linear manifold able to represent the original function in fewer dimensions. In this report, we discuss a neural network architecture and we perform a numerical campaign on several problems ranging from simple two-dimensional functions to a model problem for magnetohydrodynamics in five dimensions. In our preliminary results, we show that the proposed approach is found to be superior to linear dimension reduction strategies in representing the target function even with a single latent variable.

97 MATHEMATICS AND COMPUTING

Review of Technical Photovoltaic Key Performance Indicators and the Importance of Data Quality Routines

Technical key performance indicators (KPIs) are important metrics used to assess and quantitatively summarize various aspects of photovoltaic (PV) systems, including long-term performance, economic viability, and carbon footprint. Herein, a group of experts of the International Energy Agency's Photovoltaic Power Systems Programme Task 13 collect and describ the most important technical KPIs used in the industry. Thereby, a set of best practices for reliably handling PV system data is presented and the impact of data quality and climatic variability on KPI calculation is investigated. Further, the effective use of technical KPIs allows triggering data-driven and informed decisions to optimize PV systems and providing a comprehensive overview of how PV systems operate across different conditions and climates. With the worldwide growth of the PV industry, more companies operate/own PV systems in different regions, where the climatic and seasonal profiles differ. This requires context-aware evaluation of KPIs, or the judicious application of multiple KPIs, to ensure that each asset is evaluated correctly. Beyond that, there is untapped potential in the utilization of KPIs through geospatial mapping and extrapolation of fleet KPIs. This study demonstrates that the uncertainty in KPI estimation is not well understood and depends on data quality, climatic variability, and system configuration.

14 SOLAR ENERGY

Strong Lensing by Galaxies

Strong gravitational lensing at the galaxy scale is a valuable tool for various applications in astrophysics and cosmology. Some of the primary uses of galaxy-scale lensing are to study elliptical galaxies’ mass structure and evolution, constrain the stellar initial mass function, and measure cosmological parameters. Since the discovery of the first galaxy-scale lens in the 1980s, this field has made significant advancements in data quality and modeling techniques. In this review, we describe the most common methods for modeling lensing observables, especially imaging data, as they are the most accessible and informative source of lensing observables. We then summarize the primary findings from the literature on the astrophysical and cosmological applications of galaxy-scale lenses. We also discuss the current limitations of the data and methodologies and provide an outlook on the expected improvements in both areas in the near future.

79 ASTRONOMY AND ASTROPHYSICS

GNSS-based Vegetation Optical Depth, Tree Sway, and Evapotranspiration data from the Niwot Ridge Subalpine Forest (US-NR1) AmeriFlux site

This data package contains data and information about Global Navigation Satellite System (GNSS)-based Vegetation Optical Depth (VOD), tree sway motion, and eddy-covariance evapotranspiration (ET) data collected at the Niwot Ridge Subalpine Forest AmeriFlux site (US-NR1). The raw GNSS data were collected between May 2022 and August 2023. Other processed datasets such as tree sway motion and ET data are also included. The goal was to study the water content within a subalpine forest and, more specifically, examine the canopy evaporation process. This data archive includes all data that were used within the following Biogeosciences discussion paper that further summarizes the research objectives and conclusions:Burns, S.P., V. Humphrey, E.D. Gutmann, M.S. Raleigh, D.R. Bowling, and P.D. Blanken, 2025: Using GNSS-based vegetation optical depth, tree sway motion, and eddy-covariance to examine evaporation of canopy-intercepted rainfall in a subalpine forest. EGUsphere [preprint],https://doi.org/10.5194/egusphere-2025-1755This data archive also supplements the 30-min Lawrence Berkeley National Laboratory (LBNL) AmeriFlux dataset for US-NR1 (i.e., https://doi.org/10.17190/AMF/1246088) and updates what was in the 2020 ESS-DIVE US-NR1 archive (https://doi.org/10.15485/1671825) to include data from the years 2020-2025. More specifically, the following updates are provided: (i) five-minute statistics (means, variances, covariances) of all data measured by the US-NR1 data system between Sep 2020 and Jun 2025 in netCDF format, (ii) the electronic logbook of US-NR1 site visits, (iii) a web calendar (in HTML format) documenting activity at the site (a replica of https://urquell.colorado.edu/calendar/), (iv) photos taken at the site between years 2020 and present day (Aug 2025), and (v) several auxiliary datasets, primary related to trees near the site, soil properties, soil moisture and soil temperature, and subcanopy radiation data. The data package is setup so that the web calendar, photos, and electronic logbook can be easily accessed on a local computer using a web browser. The provided data files are in either BINEX or SBF format (for the raw GNSS data), netCDF, CSV, ASCII, or MATLAB format. To obtain a better understanding about the archive, please start by reading the following PDF which is included within the data archive:README_ESS_DIVE_USNR1_2025_readme_first.pdf.

54 ENVIRONMENTAL SCIENCES