Search NASA⌕ Search

SEARCH · Search NASA

Results for “DATA BASES”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Uncertainty quantification of a physics-informed model based on sparse identification of a Thermal Energy Distribution System

Integrated energy systems (IES)s are crucial for enhancing the economy and efficiency of power generation sources (e.g., nuclear energy) necessary to unleash American energy dominance. These systems can be integrated with thermal energy storage (TES) and intermittent renewable energies to optimize overall energy use, peak-load regulation, and demand-side responses. However, the stabilization of energy generation, transport, and utilization introduces operational complexities that exceed the challenges of managing each sub-component individually. Currently, though IESs rely on human operators for efficiency and stability, reducing human error risk and enhancing performance through automation is highly desirable. Recent advances at Idaho National Laboratory have demonstrated successful control of the Thermal Energy Distributed System (TEDS). However, the automatic control system depends on a deterministic Sparse Identification of Nonlinear Dynamics with Control (SINDyC) model, which are trained based on simulation data from physics-based simulations. Because of uncertainties in physics-based simulation, SINDyC model results in large discrepancies against experimental data and cannot be reliably used in automatic control. In this paper, we present an innovative approach to address these discrepancies by quantifying uncertainties and developing a more robust model. We first generated trajectories by using first-principles physics codes to encapsulate the experiment. Next, we trained thousands of models by randomly sampling these trajectories. We then collapsed all those models into one probabilistic SINDyC by fitting a multivariate Gaussian distribution onto the resulting coefficient’s distribution. Despite its simplicity, our approach successfully produced 95% confidence intervals that captured the experimental trajectories. It even did so with a higher probability and better U-pooling score across six of the seven relevant quantities of interest (QoIs), as compared to other classical approaches. In conclusion, ongoing research is focusing on generating new experimental trajectories to validate this approach, and on employing Bayesian calibration to refine parametric uncertainties and guide future model development efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Autoencoder-Based Anomaly Detection System for Online Data Quality Monitoring of the CMS Electromagnetic Calorimeter

The CMS detector is a general-purpose apparatus that detects high-energy collisions produced at the LHC. Online data quality monitoring of the CMS electromagnetic calorimeter is a vital operational tool that allows detector experts to quickly identify, localize, and diagnose a broad range of detector issues that could affect the quality of physics data. A real-time autoencoder-based anomaly detection system using semi-supervised machine learning is presented enabling the detection of anomalies in the CMS electromagnetic calorimeter data. A novel method is introduced which maximizes the anomaly detection performance by exploiting the time-dependent evolution of anomalies as well as spatial variations in the detector response. The autoencoder-based system is able to efficiently detect anomalies, while maintaining a very low false discovery rate. The performance of the system is validated with anomalies found in 2018 and 2022 LHC collision data. In addition, the first results from deploying the autoencoder-based system in the CMS online data quality monitoring workflow during the beginning of Run 3 of the LHC are presented, showing its ability to detect issues missed by the existing system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

New approaches to Bayesian uncertainty quantification for Nuclear Science (Final Technical Report)

Inverse problems play a central role in experimentation and theory/data comparisons for many areas of modern Nuclear Physics (NP) and High-Energy Physics (HEP). Bayes’s Theorem is a powerful tool for solving Inverse Problems, providing conceptually transparent and unbiased constraints on theoretical parameters and their uncertainties (“Bayesian Inference”) and enabling the quantification of agreement or tension between models and data. However, analyses based on Bayesian Inference are often challenging for NP and HEP applications, either because of the large number of parameters in the problem, the high computational cost, or both. We propose a multi-institutional collaboration to develop and deploy novel Bayesian analysis tools that advance the scientific scope of a broad range of current and future NP experiments. This project brings together NP domain scientists working on several high-profile NP projects for which new, high-performance Bayesian Uncertainty Quantification (“Bayesian UQ”) methods are essential to carry out the science, and data scientists who are developing state-of-the-art methods applicable to these problems. The NP projects in this proposal comprise measurements of the mass and fundamental nature of the neutrino; study of the Quark-Gluon Plasma that filled the early universe; and mapping of natural and anthropogenic radiation environments. While these NP projects have very different scientific goals, with datasets and analysis approaches that differ significantly, they share common requirements for improving computationally intensive Bayesian analyses using advanced Machine Learning algorithms and will benefit strongly from a coherent effort to develop general solutions. This proposal brings together these projects and forefront ML-based data science algorithms to develop such general solutions. The methods developed in this project will also be more widely applicable, thereby advancing science in the larger Nuclear Physics portfolio.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics↗

Major impacts of widespread structural variation on sorghum

Genetic diversity is critical to crop breeding and improvement, and dissection of the genomic variation underlying agronomic traits can both assist breeding and give insight into basic biological mechanisms. Although recent genome analyses in plants reveal many structural variants (SVs), most current studies of crop genetic variation are dominated by single-nucleotide polymorphisms (SNPs). The extent of the impact of SVs on global trait variation, as well as their utility in genome-wide selection, is not yet understood. In this study, we built an SV data set based on whole-genome resequencing of diverse sorghum lines (n = 363), validated the correlation of photoperiod sensitivity and variety type, and identified SV hotspots underlying the divergent evolution of cellulosic and sweet sorghum. In addition, we showed the complementary contribution of SVs for heritability of traits related to sorghum adaptation. Importantly, inclusion of SV polymorphisms in association studies revealed genotype–phenotype associations not observed with SNPs alone. Three-way genome-wide association studies (GWAS) based on whole-genome SNP, SV, and integrated SNP + SV data sets showed substantial associations between SVs and sorghum traits. The addition of SVs to GWAS substantially increased heritability estimates for some traits, indicating their important contribution to functional allelic variation at the genome level. Our discovery of the widespread impacts of SVs on heritable gene expression variation could render a plausible mechanism for their disproportionate impact on phenotypic variation. This study expands our knowledge of SVs and emphasizes the extensive impacts of SVs on sorghum.

59 BASIC BIOLOGICAL SCIENCES↗

Projected Urban Morphology of the Los Angeles Area by the Year 2100

This dataset provides projections of urban building morphologies for the Los Angeles urban area at 30-meter spatial resolution. It contains 192 raster files that detail two primary building attributes: building footprint fractions (ranging from 0 to 1) and average building heights (ranging from 0 to 75 meters). The projections account for a wide range of future pathways, covering two Shared Socioeconomic Pathway (SSP) scenarios (SSP3 and SSP5), two population scenarios, two developed land intensification scenarios, and four distinct levels of intensification. The dataset was created using dual Generative Adversarial Networks (GANs) trained on 2015 land cover and building properties from the National Land Cover Database (NLCD) and Model America datasets. Supporting information on the dataset has been described in the LAUrbanAreaMorphologyProjections2100_README.txt file.

Pandey, Bhartendu↗

Rapid Monitoring and Defense Approach for Resilience Improvement of Grid Cyber Security

Cyber-physical systems and electric utilities significantly depend on the reliability and efficiency of information and operational technology. However, false data injection attacks based on synchrophasor measurement data pose a serious threat to the safe and reliable operation of modern power systems. Here, to mitigate this problem, a rapid monitoring and defense approach is proposed to defend against cyber attacks. Initially, the Time and Frequency based Convolutional neural Network (TFCN) is proposed to detect different types of attacks. Within the TFCN, the advances are that both time and frequency domain information can be fused without extra spectrum analysis methods, and can save detection time to speed the calculation efficiency using the developed time-frequency block. Next, a comprehensive defense strategy is developed for multiple cyber attacks to ensure the stability and resilience of the power system according to the feedback detection results. The advances of this strategy are that different control strategies can be automatically selected to recover the stability to the greatest extent according to the detected attacks. To verify the effectiveness of the proposed approach, the high-speed frequency measurements collected from the wide-area monitoring system are used. The results demonstrate that the cyber attack detection performance can reach 95.57% accuracy, outperforming both traditional and some advanced neural networks. Importantly, the defense strategy is conducted and verified in a modified IEEE 39 bus system as well, which illustrates profound performance in faster stability restoration.

Comprehensive defense strategy↗

Experimentation in Exploring Photovoltaic Inverter Dynamics Under Different Irradiance Levels Through a Data-Driven Approach

As conventional direct connections of synchronous generators are being phased out, inverter-based resources (IBRs) with grid support functions are increasingly being integrated into power systems. This transition requires the development of accurate dynamic models for IBRs to predict how power systems will adapt to varying levels of IBRs penetration, establish grid code requirements, and ensure compliance. Here, this study introduces an active probing signal-based data-driven modeling technique to accurately derive the dynamics model of a smart photovoltaic inverter operating in Volt-Watt and Freq-Watt modes, in compliance with the IEEE 1547–2018 standard. The paper focuses on investigating how the dynamics of the PV inverter model respond to fluctuations in solar irradiance, utilizing real-time digital simulator experimentation. The experimental analysis demonstrates that the amplitude of dynamics fluctuates with changes in irradiance across both operational modes and confirms the active power’s dependence on irradiance levels. Furthermore, the nature of inverter dynamics varies distinctly between the different modes of activation. Critically, our findings indicate that dynamic models require DC-gain adjustments to accommodate contrasting irradiance levels, highlighting a negative gradient linear relationship between the DC-gain of each model and the irradiance.

14 SOLAR ENERGY↗

CODARcode/MGARD

MGARD is a software providing error-controlled lossy compression and data refactoring based on multi-grid theories. It transforms floating-point scientific data into a multilevel representation, followed by quantization and lossless encoding processes, resulting in a self-describing compressed buffer. It supports diverse data topologies, error control norms, and computing architectures.

Chen, Jieyang [University of Oregon]↗

Spatial Seal Database for Prospective Storage Resources in the USA

The goal of the Spatial Seal Database for Prospective Storage Resources in the USA is to provide relevant information and spatial extents of caprock and seal rocks. A lack of aggregated information is readily available that focuses on the caprock and seal units within sedimentary basins. The EPA class VI permit requires an assessment of the confining zone as part of submitting a permit. The data catalog of seal unit names and relevant properties with the seal spatial extent database aims to help provided important data for carbon storage based assessments. The data catalog and database are designed to show what seal data is available in a sedimentary basin and guide stakeholders to the original data source for those datasets.

Pantaleone, Scott↗

Californium-252 production at the High Flux Isotope Reactor - I: Validation study using campaign data

This paper presents a series of 252 Cf production validation and code-to-code comparison studies performed based on data from the production campaigns at the High Flux Isotope Reactor (HFIR). These studies support efforts to convert HFIR from using highly enriched uranium (HEU) fuel to low-enriched uranium (LEU) fuel. HFIR must maintain its world-class performance and missions following this conversion, and because 252 Cf is a vital neutron-emitting radioisotope used for a variety of high-impact applications (e.g., reactor startup, cancer treatment), the ability to efficiently produce 252 Cf must be preserved. In this work, the HFIRCON, Shift, ORIGEN, and TCOMP codes were deployed, and several sets of data libraries were investigated to better understand the calculation codes and the data biases. As-loaded target composition data, as-run irradiation history data, and post-irradiation measurements from recent multi-cycle irradiation campaigns of the HEU core were used to validate and determine methodology biases. Further, the findings demonstrated a good agreement, with results falling within 3 standard deviations of measurements. This paper lays the ground work for the second paper, which evaluates and compares 252 Cf production and safety metrics with the HEU core and a proposed LEU core.

07 ISOTOPE AND RADIATION SOURCES↗

AQuaRef: machine learning accelerated quantum refinement of protein structures

Cryo-EM and X-ray crystallography provide crucial experimental data for obtaining atomic-detail models of biomacromolecules. Refining these models relies on library-based stereochemical data, which, in addition to being limited to known chemical entities, do not include meaningful noncovalent interactions. Quantum mechanical (QM) calculations could alleviate these issues but are too expensive for large molecules. Here we present a novel AI-enabled Quantum Refinement (AQuaRef) based on AIMNet2 machine learned interatomic potential (MLIP) mimicking QM at substantially lower computational costs. By refining 41 cryo-EM and 30 X-ray structures, we show that this approach yields atomic models with superior geometric quality compared to standard techniques, while maintaining an equal or better fit to experimental data. Notably, AQuaRef aids in determining proton positions, as illustrated in the challenging case of short hydrogen bonds in the parkinsonism-associated human protein DJ-1 and its bacterial homolog YajL.

Zubatyuk, Roman [Carnegie Mellon University, Pitts↗