Search NASASearch

SEARCH · Search NASA

Results for “data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

EMT data generation

The integration of inverter-based resources (IBRs) in power systems is accelerating, bringing with it significant benefits such as reduced greenhouse gas emissions, improved grid resilience, and increased energy independence. Despite these advantages, the widespread adoption of IBRs introduces several challenges, including issues related to grid stability, increased operational complexity, and the need for updated regulatory frameworks. To address these challenges, IEEE released Standard 2800 in 2022, which sets forth the necessary interconnection capabilities and performance criteria for IBRs connected to transmission and sub-transmission systems. This standard outlines the performance requirements to ensure the reliable integration of IBRs into the bulk power system. Furthermore, in 2023, the North American Electric Reliability Corporation (NERC) published a reliability guideline for electromagnetic transient (EMT) modeling of BPS-connected IBRs. This guideline provides recommendations for developing EMT model requirements, performing model quality checks, and implementing verification practices specifically for EMT models representing BPS-connected inverter-based resources in reliability studies conducted by transmission planners and planning coordinators. These standards and guidelines have a profound impact on EMT studies for transmission networks, influencing system stability analyses, grid recovery and resynchronization processes, fault ride-through evaluations, protection and coordination strategies, advanced control methodologies, and the inclusion of IBRs in transient models of transmission networks. As a result, the generation of EMT data is crucial for conducting various transient-based studies to understand the impact of IBRs. EMT data generation use cases serve as the basis for scenarios in event detection and identification use cases, providing comprehensive details about EMT data generation for transmission grids with inverter-based resources. These use cases supply sufficient training and validation datasets for subsequent EMT analysis algorithms.

Xia, Qianxue

Data Generation for Machine Learning Interatomic Potentials and Beyond

The field of data-driven chemistry is undergoing an evolution, driven by innovations in machine learning models for predicting molecular properties and behavior. Recent strides in ML-based interatomic potentials have paved the way for accurate modeling of diverse chemical and structural properties at the atomic level. The key determinant defining MLIP reliability remains the quality of the training data. A paramount challenge lies in constructing training sets that capture specific domains in the vast chemical and structural space. This Review navigates the intricate landscape of essential components and integrity of training data that ensure the extensibility and transferability of the resulting models. We delve into the details of active learning, discussing its various facets and implementations. We outline different types of uncertainty quantification applied to atomistic data acquisition and the correlations between estimated uncertainty and true error. The role of atomistic data samplers in generating diverse and informative structures is highlighted. Furthermore, we discuss data acquisition via modified and surrogate potential energy surfaces as an innovative approach to diversify training data. The Review also provides a list of publicly available data sets that cover essential domains of chemical space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES

Monthly hydropower generation data for Western Canada to support Western-US interconnect power system studies

Hydroelectric power generation in Western Canada significantly contributes to power grid operations of the North American Western Interconnection through substantial generation, some of which is exported to the United States (U.S.). However, the lack of publicly available hydropower generation datasets poses challenges for future market projections and resource adequacy evaluations. We present a simulation-based monthly power system model-ready hydropower generation dataset for 110 facilities in British Columbia and Alberta from 1981 to 2019. These monthly hydropower generation estimates are developed from integrated hydrologic model simulations of runoff and reservoir-operated streamflow, followed by scaling that considers diversion inflow constraints based on hydropower water license information. To address the lack of comparable hydropower generation records, we conduct step-by-step evaluations for simulated runoff, regulated streamflow, and hydropower generation using available observations or estimates. The presented hydropower dataset aims to enhance the representation of hydropower resources in Western Canada, supporting power grid system studies for the Western Interconnection of the U.S. and Canada.

13 HYDRO ENERGY

Reliable Integration of AI Data Centers at Scale – Analysis, Modeling and Synthetic Data Generation

This report analyzes the power consumption of large dynamic digital loads using the open-source MIT supercloud and SURF datasets. With an emphasis on the MIT data, we calculate important power consumption characteristics to help system operators improve generation planning and resource allocation. We also introduce a rudimentary model for generating synthetic load profiles.

97 MATHEMATICS AND COMPUTING

IK-Frag: Frag data generator for the PHITS simulation with the inverse kinematic reaction producing a focused neutron beam

IK-Frag has been developed for the creation of the nuclear cross-section data format, which is named frag data and can be used in PHITS, a Monte Carlo simulation code. IK-Frag focuses on the inverse kinematic reactions between a lithium or beryllium ion and a proton target. These reactions achieve naturally collimated neutrons and potentially reduce the necessity of radiation shielding. IK-Frag enables PHITS users to conduct simulations for the inverse kinematic reactions. The present software aims to contribute to future development of the neutron source system using the inverse kinematic reactions.

43 PARTICLE ACCELERATORS

BSDF Data generation for daylight applications: A call for international standardization

Standardized methods for generating angle-dependent, bidirectional, solar-optical properties for complex fenestration systems do not exist, which means that energy and daylight evaluations in building performance simulations often suffer from major inaccuracies. This position paper provides an overview of state-of-the-art data-driven methods for characterizing light scattering properties of fenestration materials and blind systems (e.g. fabrics, metal slats, patterned glazing), validation via laboratory, simulation and field tests, and salient issues in support of standardization of such methods via the International Standardization Organization (ISO). The ISO standard is intended to provide the fundamental underpinnings for recently mandated daylight standards that rely on bidirectional scattering distribution function data for climate-based daylight modelling and building performance simulations.

Geisler-Moroder, D.

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE

Data Generation Workflow for Meso/Microscale Coupled Offshore Wind Farm Simulations

A robust, simple-to-use workflow is developed in this study which allows mesoscale information from the NOW-23 database to be easily incorporated into microscale wind farm simulations. This process will enable many different wind farm configurations to be simulated under realistic inflow conditions spanning a variety of atmospheric phenomena.

AMR-Wind