Search NASA⌕ Search

SEARCH · Search NASA

Results for “data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

A Comment on “Deep Proteogenomics of a Photosynthetic Cyanobacterium”

Proteomic researchers strive to achieve complete annotation of protein-coding DNA sequences to provide a foundational context for their relevant biological data. A recent deep proteogenomic study using a photosynthetic cyanobacterium Synechocystis sp. PCC 6803 by Spät et al. proposed 64 refined open reading frames (ORFs). By searching LC-MS/MS data from affinity chromatography-isolated protein complexes, our laboratory identified that six of these high-abundance ORFs possess Nterminal initiation start sites that differ than those proposed in the alternative models. Our findings are supported by highly confident MS2 data, phylogenetic analysis, chemical labeling, and established data from two independent research groups. Based on these highquality experimental identifications, we subsequently propose a standardized strategy and set of criteria for future deep proteogenomic efforts to ensure accurate and stringent proteogenomic annotation.

cyanobacteria↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Preliminary Analysis of Nuclear-Powered Data Center Scenarios

This report provides a comprehensive analysis of the potential for nuclear energy to meet the growing energy demands of data centers (DCs). It evaluates the technical, economic, and socio-environmental implications of coupling Nuclear Power Plants (NPPs) with DCs, providing initial responses to several key research questions: What is the potential increased energy demand from DCs in the U.S., in the short, medium and long term? The U.S. is experiencing a rapid increase in energy demand from DCs, with projections indicating a total increase of 24-74 GWy(e) by 2028. Meeting this demand with nuclear energy would require 27–85 GWe of installed capacity. While this surge is expected to slow in the long term, the DC industry needs reliable, scalable, and clean energy sources. How much nuclear capacity can be deployed to meet DC demand and in which timeframe? Several pathways for increasing nuclear capacity were identified, including uprates, restarts of recently retired reactors, power purchase agreements with existing fleet, and new construction. Approximately 20‒28 GWe of nuclear capacity could be dedicated to DCs by the early 2030s. How much High Assay Low Enriched Uranium (HALEU) would be needed to support some nuclear deployment scenarios for DCs? Meeting the deployment targets announced by Google and Amazon for the Kairos Power Fluoride-Salt-Cooled High-Temperature Reactor or KP-FHR (~500 MWe by 2035) and the Xe-100 (~1 GWe by 2040), respectively, requires ramping up 19.75% enriched HALEU production to ~6 t/yr by 2040. What types of nuclear energy/DC coupling options exist, and what are the different benefits/challenges? Five coupling options were analyzed, ranging from grid-connected configurations to colocated, behind-the-meter setups. Key design considerations include the proximity to high- and/or medium-voltage transmission lines, the desired internal fault tolerance, and the sources of alternative/backup power during outages. Each coupling option offers unique benefits and challenges in terms of reliability, system costs, regulation, timeline, etc. A list of NPP/DC deployment scenarios was developed, considering existing or newly built NPP or DC projects. Colocated DCs with new small modular reactors or large reactors on greenfield and brownfield sites are the focus of this report. What types of reactors, especially what size, may be incentivized by DCs? Reactor sizing optimization revealed that the ideal reactor size and number of units depend on DC demand, coupling configurations defined in this report, and other economic factors. Larger reactors are preferred for high-demand DCs and grid-connected systems, while larger number of smaller reactors are better suited for DC configurations without grid backup. Which sites may be compatible with co-located nuclear-powered DCs? Siting those projects is a complicated evaluation factoring local water resources, grid connection availability and reliability, IT infrastructure, local work force, proximity to population zones, etc. For this effort greenfield and brownfield sites such as retired coal-fired plants were used to evaluate this question. This evaluation is not meant to recommend any particular site but it highlights key siting criteria and demonstrates large-scale site availability. What are the socio-economic impacts of co-located nuclear-powered DCs? Those projects generate substantial economic benefits to the local economy, particularly in urban settings. Hyperscale DCs colocated with nuclear power plants (sized around 1 GW of power) can create nearly 1,700 jobs for annual operations and more than 7,300 jobs among the supply chain and local businesses as a result of increased household spending. Rural projects also provide significant benefits, but at lower magnitudes compared to urban deployments.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Time Distribution Analysis for Task Primitives to Support Dynamic Human Reliability Analysis

To support data collection for dynamic human reliability analysis (HRA), this study investigates time distributions for task primitives defined in the Goals, Operators, Methods, and Selection rules (GOMS)–Human Reliability Analysis (HRA) method and Human Reliability data EXtraction (HuREX). GOMS-HRA was developed to provide cognition-based time and human error probability (HEP) information for dynamic HRA calculations within the Human Unimodel for Nuclear Technology to Enhance Reliability (HUNTER) framework, while HuREX is a comprehensive HRA data collection method developed by the Korea Atomic Energy Research Institute (KAERI). In this paper, we examine time distributions by using experimental data collected from the Simplified Human Error Experimental Program (SHEEP) study, which proposes an HRA data collection framework to complement full-scope simulator research and gather input data for dynamic HRA by using simplified simulators such as the Rancor Microworld simulator. This paper investigates whether the time required for GOMS-HRA and HuREX task primitives fits 13 statistical distributions. Additionally, we compare and discuss the time distributions obtained from both student operators and professional operators. The result was that this study identified several time distributions for five GOMS-HRA and four HuREX task primitives. In the future, the results of this study are expected to provide objective reference data on the elapsed time for task primitives and aid in realistically simulating scenarios within dynamic HRA.

Dynamic Human Reliability Analysis↗

MRCI Task 3: Facilitating Data Collection, Sharing, and Analysis Final Technical Summary Report

The Midwest Regional Carbon Initiative (MRCI) Task 3.0 was defined to facilitate development of carbon capture, utilization, and storage (CCUS) in the region by collection and sharing of existing and new technical data from CCUS projects and research. The task also included support for further analysis and assessment of tools by the project team and by researchers working on programs such as National Risk Assessment Partnership (NRAP), machine learning (ML) techniques, and assessment and improvement of CCUS site assessment, operations, and monitoring aspects. Work under Task 3.0 addressed key issues related to CCUS deployment and provided foundational research and datasets to help establish CCUS projects in the MRCI. Report Authors and Principal Technical Contributors: Joel Sminchak, Laura Keister, Mackenzie Scharenberg, Priya Ravi-Ganesh, Autumn Haagsma, Srikanta Mishra, Jared Hawkins, Jared Schuetter, Amy Lang, Jaelen Lewis, Derrick James, Jorge Barrios, Stuart Skopec, and Sanjay Mawalkar (Battelle). Chris Korose, Carl Carmen, Nate Grigsby, Nathan Webb (Illinois State Geological Survey). Principal Investigators: Dr Neeraj Gupta, Dr. Chris Korose.

MRCI,NRAP,data collection,data compilation,legacy ↗

Extraction of the Collins-Soper Kernel from a Joint Analysis of Experimental and Lattice Data

We present a first joint extraction of the Collins-Soper kernel (CSK) combining experimental and lattice QCD data in the context of an analysis of transverse-momentum-dependent distributions (TMDs). Based on a neural-network parametrization, we perform a Bayesian reweighting of an existing fit of TMDs using lattice data, as well as a joint TMD fit to lattice and experimental data. We consistently find that the inclusion of lattice information shifts the central value of the CSK by approximately 10% and reduces its uncertainty by 40%–50%, highlighting the potential of lattice inputs to improve TMD extractions.

Avkhadiev, Artur [Massachusetts Inst. of Technolog↗

Demonstration of new MeV-scale capabilities in large neutrino LArTPCs using ambient radiogenic and cosmogenic activity in MicroBooNE

Large neutrino liquid argon time projection chamber (LArTPC) experiments can broaden their physics reach by reconstructing and interpreting MeV-scale energy depositions, or blips, present in their data. We demonstrate new calorimetric and particle discrimination capabilities at the MeV energy scale using reconstructed blips in data from the MicroBooNE LArTPC at Fermilab. We observe a concentration of low-energy (<3 MeV) blips around fiberglass mechanical support struts along the time projection chamber edges with energy spectrum features consistent with the Compton edge of 2.614 MeV 208 Tl decay 𝛾 rays. These features are used to verify proper calibration of electron energy scales in MicroBooNE’s data to few percent precision and to measure the specific activity of 208 Tl in the fiberglass composing these struts, (11.7 ± 0.2⁢(stat) ± 3.1⁢(syst)) Bq/kg. Cosmogenically produced blips above 3 MeV in reconstructed energy are used to showcase the ability of large LArTPCs to distinguish between low-energy proton and electron energy depositions. An enriched sample of low-energy protons selected using this new particle discrimination technique is found to be smaller in data than in dedicated corsika cosmic-ray simulations, suggesting either incorrect corsika modeling of incident cosmic fluxes or particle transport modeling issues in geant4.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

2025 TEM Workshop

The TEM Data Management Workshop will take place on August 26 from 9 a.m. to 12 p.m. MT, and will be held virtually on TEAMS. The primary goal of this workshop is to engage NSUF users and stakeholders in discussions about the data needs for the utilization of AI and ML in the analysis of TEM data. Key topics to be covered include data storage, data sharing, data tagging, metadata inclusion, standardized data formats, data augmentation, and annotated training datasets. Additionally, the workshop will provide valuable insights into resources such as the Nuclear Research Data System (NRDS) for data storage and sharing, as well as open-source codes for data analysis.

Bachhav, Mukesh↗

Homomorphic Encryption for Electrical Metering Aggregation: Protecting the Privacy of Building Tenants

Electrical meters are devices that measure consumer electricity usage. The data collected by these meters is necessary for utility billing and electrical grid management but can also be used to assess the environmental impact of buildings. Prior research has found that unprotected metering data could potentially be used to infer some information about the behaviors of building tenants by detecting changes in electricity usage. For example, a period of low electricity usage could suggest that the tenants are not in the building. As smart metering becomes more common, there is a growing need for data privacy protections for metering data that do not negatively impact the quality and availability of data used for energy management and billing applications. To identify potential solutions, we developed a Python-based data aggregation platform to analyze the potential efficacy of privacy-enhancing technologies for energy metering applications. This platform aggregates groups of metering sites into virtual buildings, which could potentially detach changes in electrical activity from individual tenants, making it more difficult to track the activity of a specific tenant. To further protect data during analysis, this project utilizes homomorphic encryption as part of its initial approach. Homomorphic encryption offers a means of protecting energy consumption data while permitting mathematical operations to be performed without the need to know the data contents. This allows for data to be processed into usable statistics without revealing energy consumption information. A series of homomorphic encryption libraries were evaluated to determine their applicability and limitations in the context of metering data. The use of these techniques may help to reassure consumers and encourage further adoption of smart grid infrastructure.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Polarization observables 𝑇 and 𝐹 in the 𝛾⁢𝑝→𝜋 0 ⁢𝑝 reaction at CLAS

Pion photoproduction in the $\vec{γ}$$\vec{p}$ → π 0 p reaction has been measured in the FROzen Spin Target (FROST) experiment at the Thomas Jefferson National Accelerator Facility. In this experiment, circularly polarized photons with energies up to 3.082 GeV impinged on a transversely polarized frozen-spin target. Final-state protons were detected in the Continuous Electron Beam Accelerator Facility (CEBAF) Large Acceptance Spectrometer. The polarization observables T and F have been extracted for W from 1445 to 2525 MeV, of which the energy range is much broader, and the precision is better than the existing measurements in higher W ranges. The data generally agree with predictions of present partial-wave analyses but also show marked differences for higher W ranges. By incorporating the present data into the databases, the Scattering Analysis Interactive Data (SAID) fits have been improved with relatively small χ 2 and significant changes in the parameters of the (1910)1/2 + and N(2190)7/2 − have been found with the JüBo model.

particle production↗

Utah FORGE: Neubrex Well 16B(78)-32 DAS Data - April, 2024

This dataset comprises Distributed Acoustic Sensing (DAS) data collected from the Utah FORGE monitoring well 16B(78)-32 (the producer well) during hydraulic fracture stimulation operations conducted in April 2024. The data were acquired continuously over the stimulation period at a temporal sampling rate of 10,000 Hz (10 kS/s) and a spatial resolution of approximately 3.35 feet (1.02109 meters). The measurements were captured using a Neubrex NBX-S4100 Time Gated Digital DAS interrogator unit connected to a single-mode fiber optic cable, which was permanently installed within the casing string. All recorded channels correspond to downhole segments of the fiber optic cable, from a measured depth (MD) of 5,369.35 feet to 10,352.11 feet. The DAS data reflect raw acoustic energy generated by physical processes within and surrounding the well during stimulation activities at wells 16A(78)-32 and 16B(78)-32. These data have potential applications in analyzing cross-well strain, far-field strain rates (including microseismic activity), induced seismicity, and seismic imaging. Metadata embedded in the attributes of the HDF5 files include detailed information on the measured depths of the channels, interrogation parameters, and other acquisition details. The dataset also includes a recording of a seminar held on September 19, 2024, where Neubrex's Chief Operating Officer presented insights into the data collection, analysis, and preliminary findings. The raw data files, stored in HDF5 format, are organized chronologically according to the recording intervals from April 9 to April 24, 2024, with each file corresponding to a 12-second recording interval.

15 GEOTHERMAL ENERGY↗

Toward the Neutrino Discovery Platform: An Auditable, Uncertainty-Bearing Toolchain for MINERvA Open-Data Cross-Section Analysis

The Neutrino Discovery Platform (NDP) aims to accelerate DUNE-era science by making the neutrino program's existing datasets analyzable through fast, reproducible, and auditable workflows. We report a working version of two of its layers, data curation and agentic orchestration, built and tested end to end on MINERvA open data. The guiding lesson throughout is that a cross section is a measurement, and not just a plotted shape, only if it carries a defensible systematic-uncertainty budget, a trustworthy unfolding, and a reproducible record. Using a single medium-energy playlist pair from the MINERvA open-data release (about $2.05\times10^{17}$ protons on target of data), we first reproduced the shapes of two published charged-current inclusive $\nu_\mu$ measurements through a complete extraction ladder: selection, background subtraction, D'Agostini unfolding, efficiency correction, and flux normalization. These shape-level reproductions ran and tracked the published results, but they lacked the systematic-uncertainty machinery that defines a MINERvA cross section. To supply it, we vendored and built the MINERvA Analysis Toolkit and developed a many-universe systematic-uncertainty tool that produces a portable covariance artifact, a parallel event-loop runner, and a per-run auditability harness. Validated against a published covariance release, the toolchain reproduces the released statistical, flux, and muon-energy-scale terms and shows that they account for roughly 63\% of the total variance, with the remainder unreleased. Using this same infrastructure, we then performed a measurement of our own design, the hadronic recoil-energy distribution of low-energy ($E_\nu<2.5$~GeV) charged-current inclusive events, and found data/simulation shape agreement of $\chi^2/\mathrm{ndf}=1.26$. Together these results show that the platform supports original physics and not only reproductions.

Breaux, Auto [Tulane U. (main)]↗

Commercial, industrial, and institutional discount rate estimation for efficiency standards analysis: Sector-level data 1998–2023

Underlying each of the U.S. Department of Energy’s (DOE’s) federal appliance and equipment energy conservation standards are a set of complex analyses of the projected costs and benefits of regulation. Any new or amended standard must be designed to achieve significant additional energy conservation, provided that it is technologically feasible and economically justified (42 U.S.C. 6295(o)(2)(A)). DOE determines economic justification based on whether the benefits exceed the burdens, considering a variety of factors, including the economic impact of the standard on consumers of the product and the savings in lifetime operating cost compared to any increase in price or maintenance expenses (42 U.S.C. 6295(o)(2)(B)). As part of this determination, DOE conducts a life-cycle cost (LCC) analysis, which models the combined impact of appliance first cost and operating cost changes on a representative commercial building sample to identify the fraction of customers achieving LCC savings or incurring net cost at the considered efficiency levels. Thus, the commercial discount rate value(s) used to calculate the present value of energy cost savings within the LCC model implicitly plays a role in estimating the economic impact of potential standard levels. This report provides an in-depth discussion of the commercial discount rate estimation process relying on the Capital Asset Pricing Model (CAPM) to estimate a business’ cost of equity, and by adding a risk adjustment factor to the risk-free rate associated with long-term U.S. Treasury bonds to estimate their cost of debt. It is an update to previous reports on estimating commercial discount rates from firm-level and sector-level financial data (e.g., Fujita, 2021, 2016). Major topics covered in this report include the following: • Discount rate estimation methods and rationale • Data sources used and data limitations • Discount rate distributions for use in standards analysis • Discount rate estimation methods and distributions specific to the small business subgroup analysis.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Codebase release 2.0 for sauce

Low energy nuclear physics experiments are transitioning towards fully digital data acquisition systems. Realizing the gains in flexibility afforded by these systems relies on equally flexible data reduction techniques. In this paper, methods utilizing data frames and in-memory techniques to work with data, including data from self-triggering, digital data acquisition systems, are discussed within the context of a Python package, sauce. It is shown that data frame operations can encompass common analysis needs and allow interactive data analysis. Two event building techniques, dubbed referenced and referenceless event building, are shown to provide a means to transform raw list mode data into correlated multi-detector events. These techniques are demonstrated in the analysis of two example data sets.

Marshall, Caleb (ORCID:0000000211942920)↗

Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems

The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globallyconsistent meta-data. In this paper, we propose a novel datastructure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.

Castellana, Vito G. [BATTELLE (PACIFIC NW LAB)] (O↗

Software and computing for Run 3 of the ATLAS experiment at the LHC

The ATLAS experiment has developed extensive software and distributed computing systems for Run 3 of the LHC. These systems are described in detail, including software infrastructure and workflows, distributed data and workload management, database infrastructure, and validation. The use of these systems to prepare the data for physics analysis and assess its quality are described, along with the software tools used for data analysis itself. An outlook for the development of these projects towards Run 4 is also provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Lessons Learned from AskGDR: Usage and Impact Analysis of the Geothermal Data Repository's AI Research Assistant: Preprint

In October of 2024, the Department of Energy's (DOE) Geothermal Data Repository (GDR) team officially launched AskGDR, an AI research assistant resulting from the integration of a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets. AskGDR allows GDR users to ask deeper questions about the origin of datasets, the methods used to collect them, and the findings they help support. Using Retrieval Augmented Generation (RAG), AskGDR can be used to summarize findings spread across dozens of papers and technical reports or to extract relevant information describing a single data field. However, generative AI is experimental. The National Renewable Energy Laboratory (NREL) has been collecting metrics on AskGDR and documenting lessons learned during its deployment. This paper will outline the efficacy and impact of AskGDR through analysis of its use, operating costs, number and types of questions asked, and the quality of answers provided.

15 GEOTHERMAL ENERGY↗