Search NASA⌕ Search

SEARCH · Search NASA

Results for “data access”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

OSW Consortium 2 - Validated National Offshore Wind Resource Dataset with Uncertainty Quantification (CRADA Report)

This research has led to the development of the 2023 National Offshore Wind data set (NOW-23), which offers the latest wind resource information for offshore regions in the United States. NOW-23 supersedes, for its offshore component, the Wind Integration National Dataset (WIND) Toolkit, which was published a decade ago and is currently a primary resource for wind resource assessments and grid integration studies in the contiguous United States. By incorporating advancements in the Weather Research and Forecasting (WRF) model, NOW-23 delivers an updated and cutting-edge product to stakeholders. As part of this project, we also developed a summary of the uncertainty quantification in NOW-23, along with NOW-WAKES, a 1-year post-construction data set that quantifies expected offshore wake effects in the US Mid-Atlantic lease areas. Stakeholders can access the NOW-23 data set at https://doi.org/10.25984/1821404.

17 WIND ENERGY↗

EDX ClaiMM: Digital Resources for the Critical Minerals and Materials Community

Securing critical mineral supply chains is essential for transitioning to a clean energy economy and for maintaining national security. Big-data analytics can serve as a cost-effective means of identifying new domestic critical mineral resources but only if data can be easily located and digested. Using ArcGIS Enterprise Sites, EDX ClaiMM was developed to increase the accessibility of critical minerals data, reducing time spent on data collection and integration. Hosted tools provide rapid visualization and exploration of key datasets, unlocking insights to support resource assessments.

Yesenchak, Rachel↗

The Foundational Industrial Energy Dataset (FIED): Open-Source Data on Industrial Facilities

The state of data on industrial energy use has co-evolved over several decades with the demands of industrial energy analysis. The most recent development - analysis in support of decarbonizing the industrial sector - has changed the characteristics of industrial data that are useful for analysts and model developers. Although data and its collection processes may be cast from a conventional viewpoint as objective and free from the influence of social dynamics, this provides an incomplete picture of not only the processes by which information is generated, but also the limitations and opportunities of data to be useful for analysis. The foundational industry energy data set (FIED) is a result of the confluence of trends in open data and the demand for higher resolution industrial energy analysis. The general approach to compiling the FIED involves accessing, filtering, and formatting data published by federal organizations on the Internet for public use. Unlike most industrial energy datasets, which are published by the U.S. Energy Information Administration (EIA), the FIED relies on core datasets from the U.S. Environmental Protection Agency (EPA). The FIED addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization - including estimates of energy use, greenhouse gas emissions, and design capacities - for facilities that are identified by latitude and longitude. This enables local-level analysis of existing combustion equipment, as well as regional comparisons with traditional industrial energy data estimates. The report summarizes the general logic behind compiling the FIED. The FIED itself and its Python code are available from OpenEI and GitHub, respectively.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗

Performance Year 1 Technical Report - OPEN COG Grid: Extendable Coherent Models-Datasets for Cognitive Power Grids

The OPEN COG Grid project is a collaborative effort between LLNL, NREL, and Texas A&M University (TAMU) to develop synthetic power system datasets that (i) contain all technical information that would be available in a real system, allowing to conduct studies ranging from dynamic simulation to long term planning studies; ii) are accessible to researchers from the broader data sciences community, as oppossed to power system experts only; and (iii) This report summarizes the work conducted during the first 15 months of execution of the project. These activities encompassed: 1. Conduct a survey of existing open data sets and open source power systems simulators, their supported use cases, and accessibility (Chapter 1). 2. Define a new extensible specification for power system data, covering all parameters necessary for most computational use cases (Chapter 2). 3. Collecting real technical system data to complete missing parameters in existing open source datasets (Chapter 3). 4. Develop models that capture the behavior of emergent actors in power grids, neglected by existing datasets; aggregated residential demand response (Chapter 4) and demand response of cryptocurrency miners (Chapter 5). 5. Collect detailed spatial information on distributed energy resources, particular, solar photovoltaic facilities (Chapter 6). The following chapters provide detailed descriptions of these tasks, the assumptions taken, and their findings. In conducting these tasks, the project team produced: two (accepted) conference papers; one journal paper under submission; one draft journal paper pending submission; released one repository with the developed power system data specification, with documentation and examples; and one extended dataset for the Texas power grid under review for release. The team hopes these contributions will enhance access to power system data and remove barriers to the development of new computational techniques for power systems, particularly, those inspired by cognitive sciences.

24 POWER TRANSMISSION AND DISTRIBUTION↗

MemFriend: Understanding Memory Performance with Spatial-Temporal Affinity

In HPC applications, memory access behavior is one of the main factors affecting performance. Improving an application’s memory access behavior involves optimizing data layout and/or restructuring code, and requires studying spatial-temporal data locality. Existing data locality analyses focus on single-location metrics and are restricted to evaluating temporal locality. We introduce spatial-temporal affinity metrics that quantify temporal access proximity, forward access correlation, and nearby access correlation between pairs of memory locations. We describe methods for distinguishing between potential vs. realized affinity and for reasoning about affinity at multiple resolutions (3D, 2D, 1D). Finally, we construct spatial-temporal affinity signatures that classify memory behavior and that be used to reason about changes in software (data relayout, code refactoring) or hardware (caching, prefetching). We describe methods for signature visualization, interpretation, and quantitative comparison of signatures. We evaluate our methodology using applications with variants that contrast data structures, data layouts and algorithms. We show that spatial-temporal affinity analysis provides novel insights and enables predictive reasoning about application performance when contrasted with reuse distance analysis.

Suriyakumar, Yasodhadevi↗

Towards a Public Event Display for DUNE

The Deep Underground Neutrino Experiment (DUNE) is a next generation long baseline neutrino experiment based at Fermilab, with a near detector near the beam target and a Far Detector (FD) in South Dakota. As the experiment prepares for its first data runs, creating pathways for public engagement and data transparency is essential. We present the first-ever DUNE event display designed for public outreach and education. Developed using data from the ProtoDUNE detectors at the CERN Neutrino Platform, this tool provides an intuitive and interactive interface that allows non-experts to visualise and explore particle interactions in a Liquid Argon Time Projection Chamber (LArTPC). By translating raw experimental data into a browser-accessible format, we establish the essential infrastructure for DUNE’s pathway to open data. This talk will detail the technical development of the display, its current implementation with ProtoDUNE data, and the strategic roadmap for integrating it into DUNE’s long-term open-access framework.

Sabater, Eva [U. Sussex (main)] (ORCID:00090001748↗

Data analytics for intermodal freight transportation applications

With the growth of intermodal freight transportation, it is important that transportation planners and decision-makers are knowledgeable about freight flow data to make informed decisions. This is particularly true with Intelligent Transportation Systems (ITS) offering new capabilities for intermodal freight transportation. Specifically, ITS enables access to multiple different data sources, but they have different formats, resolutions, and time scales. Thus, knowledge of data science is essential to be successful in future ITS-enabled intermodal freight transportation systems. This chapter discusses the commonly used descriptive and predictive data analytic techniques in intermodal freight transportation applications. These techniques cover the entire spectrum of univariate, bivariate, and multivariate analyses. In addition to illustrating how to apply these techniques manually, this chapter will also show how to apply them using the statistical software R. Additional exercises are provided for those who wish to apply the described techniques to more complex problems.

Huynh, Nathan↗

Benchmarking CO₂ storage simulations: Results from the 11 th Society of Petroleum Engineers Comparative Solution Project

The 11 th Society of Petroleum Engineers Comparative Solution Project (shortened SPE11 herein) benchmarked simulation tools for geological carbon dioxide (CO 2 ) storage. A total of 45 groups from leading research institutions and industry across the globe signed up to participate, with 18 ultimately contributing valid results that were included in the comparative study reported here. This paper summarizes the SPE11 results. A comprehensive introduction and qualitative discussion of the submitted data are provided, together with an overview of online resources for accessing the full depth of data. A global metric for analyzing the relative distance between submissions is proposed and used to conduct a quantitative analysis of the submissions. This analysis attempts to statistically resolve the key aspects influencing the variability between submissions. The study shows that the major qualitative variation between the submitted results is related to thermal effects, dissolution-driven convective mixing, and resolution of facies discontinuities. Moreover, a strong dependence on grid resolution is observed across all three versions of the SPE11. However, our quantitative analysis suggests that the observed variations are predominantly influenced by factors not documented in the technical responses provided by the participants. We therefore identify that unreported variations due to human choices within the process of setting up, conducting, and reporting on the simulations underlying each SPE11 submission are at least as impactful as the computational choices reported.

Nordbotten, Jan M. [Univ. of Bergen (Norway); Norw↗

DOE BSSD Performance Management Metrics Report Q3

Microbiome data is complex, spanning information from microbial genomes within diverse communities, protein and metabolite readouts, and contextual information (metadata) captured from the environments from which these samples were collected. While the variety and scale of microbiome data generation has dramatically expanded over the past twenty years, infrastructure to support data management, sharing, and access has lagged. New ways to improve interoperability across existing resources and advancing community standards are necessary to support how researchers create, use, and reuse data. The National Microbiome Data Collaborative (NMDC) aims to advance a microbiome data sharing network through infrastructure, data standards, and community building.

54 ENVIRONMENTAL SCIENCES↗

Hybrid Storage Solution

With the rise of artificial intelligence and machine learning, data sets used to train models have become increasingly large. The availability, accessibility and integrity of large data sets has become important to the research conducted at Los Alamos National Laboratory. Ceph is a storage solution suitable for use with critical data because of its distributed nature and ability to keep multiple copies of a file in different locations. The amount of data means that bandwidth, latency, and cost are important factors and the reason most storage solutions are on-premises. However, there are distinct advantages to hosting services in the cloud, namely scalability and ease-of-use. In this paper, we explore the possibility of provisioning a hybrid Ceph cluster that leverages the benefits of both cloud architectures and on-premise performance.

97 MATHEMATICS AND COMPUTING↗

Transmission Data-Driven User-Defined Model for Inverter-based and Conventional Power Plants

Recent events in Odessa [1], [2] have shed light on the complexities of integrating large Inverter-Based Resource (IBR) plants with the transmission system, prompting NERC to stress continuous performance monitoring by transmission operators. Challenges such as plant control updates, IBR model revisions, Phase-locked loop loss of synchronism, and protection events have been identified, underscoring the need for enhanced monitoring protocols by regulatory bodies. The recent FERC 901 order underscores the importance of accurate data exchange regarding IBRs for reliability studies. However, limited access to IBR plant-related data hampers effective decision-making for transmission operators (TOP). This paper proposes a method for constructing data-driven User-Defined dynamic Models (UDM) for power plants for validating multiple-event data using field measurements from interconnection bus locations. The problem is formulated as a power plant model identification problem and a multi-task learning approach under partial input observability assumptions is proposed in this work. This approach aims to predict aggregated responses of conventional and IBR power plants during various dynamic physical events which is useful for planning studies under diverse disturbance conditions. Ultimately, this methodology emphasizes the importance of plant visibility to operators in addressing power system challenges, facilitating improved planning and operational studies.

Mahapatra, Kaveri [BATTELLE (PACIFIC NW LAB)]↗

SCALE Procedure for Verified, Archived, Library of Inputs and Data (VALID)

This procedure provides a framework for preparing, reviewing, and storing model inputs and derived data so that individuals with authorized access to the Verified, Archived, Library of Inputs and Data (VALID) repository can use the inputs and data with confidence in their analyses. This procedure uses documented checks and reviews to ensure that the inputs and data were correctly generated using appropriate references. Configuration management is implemented to prevent inadvertent modification of the inputs and data or inclusion of models that have not been reviewed. This procedure also provides guidance to be followed if errors are identified or if input or data revisions are needed.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Adaptable Standards for Discovery, Access, and Usability of Oak Ridge National Laboratory’s Data Portals and Catalogs

Oak Ridge National Laboratory (ORNL) is leveraging its established capabilities and subject matter expertise in data curation, governance, management, national security, and risk assessment and mitigation to support the US Department of Energy (DOE) Grid Modernization Initiative. Using standards modeled by the National Institute of Standards and Technology (NIST), the Data Curation Network (DCN), the Oak Ridge Leadership Computing Facility (OLCF), and other leading organizations in the fields of energy research, high-performance computing, and national and homeland security, ORNL seeks to provide a federated approach to research data discovery, use, and interoperability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives.

access↗

Special Observing Period (SOP) data for the Year of Polar Prediction site Model Intercomparison Project (YOPPsiteMIP)

The rapid changes occurring in the polar regions require an improved understanding of the processes that are driving these changes. At the same time, increased human activities such as marine navigation, resource exploitation, aviation, commercial fishing, and tourism require reliable and relevant weather information. One of the primary goals of the World Meteorological Organization's Year of Polar Prediction (YOPP) project is to improve the accuracy of numerical weather prediction (NWP) at high latitudes. During YOPP, two Canadian “supersites” were commissioned and equipped with new ground-based instruments for enhanced meteorological and system process observations. Additional pre-existing supersites in Canada, the United States, Norway, Finland, and Russia also provided data from ongoing long-term observing programs. These supersites collected a wealth of observations that are well suited to address YOPP objectives. In order to increase data useability and station interoperability, novel Merged Observatory Data Files (MODFs) were created for the seven supersites over two Special Observing Periods (February to March 2018 and July to September 2018). All observations collected at the supersites were compiled into this standardized NetCDF MODF format, simplifying the process of conducting pan-Arctic NWP verification and process evaluation studies. This paper describes the seven Arctic YOPP supersites, their instrumentation, data collection and processing methods, the novel MODF format, and examples of the observations contained therein. MODFs comprise the observational contribution to the model intercomparison effort, termed YOPP site Model Intercomparison Project (YOPPsiteMIP). All YOPPsiteMIP MODFs are publicly accessible via the YOPP Data Portal (Whitehorse: https://doi.org/10.21343/a33e-j150, Huang et al., 2023a; Iqaluit: https://doi.org/10.21343/yrnf-ck57, Huang et al., 2023b; Sodankylä: https://doi.org/10.21343/m16p-pq17, O'Connor, 2023; Utqiagvik: https://doi.org/10.21343/a2dx-nq55, Akish and Morris, 2023c; Tiksi: https://doi.org/10.21343/5bwn-w881, Akish and Morris, 2023b; Ny-Ålesund: https://doi.org/10.21343/y89m-6393, Holt, 2023; and Eureka: https://doi.org/10.21343/r85j-tc61, Akish and Morris, 2023a), which is hosted by MET Norway, with corresponding output from NWP models.

54 ENVIRONMENTAL SCIENCES↗

Methods of Securing Chemical and Pharmaceutical Knowledge and Recommendations for International Institutions to Enhance Research Integrity

Here, this paper examines strategies for securing chemical and pharmaceutical expertise in a globalized research environment, focusing on safeguarding intellectual property and preventing the misuse of sensitive and potentially dual-use information. The product of collective efforts between Pacific Northwest National Laboratory, Carol Davila University of Medicine and Pharmacy, and New Bulgarian University, highlights the challenges and opportunities posed by cross-border research collaborations, particularly in the context of differing regulatory frameworks and research cultures. It explores current mechanisms to prevent data loss and unauthorized access to sensitive information while assessing the effectiveness of existing security measures, frameworks, and international export control regimes. The approach examines the differing methodologies for promoting transparency, trust-building, and mutual accountability in joint research projects to cultivate secure data-sharing practices and intellectual property. It provides recommendations for international institutions to implement security guidelines in framing research priorities, encourages continual training and education programs, and the integration of processes for monitoring research compliance. This partnership aims to advance scientific innovation while maintaining global stability, ensuring compliance with international norms, and safeguarding valuable intellectual property as measures in chemical and pharmaceutical research security practices continue to expand due to international collaboration and knowledge exchange.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Compilation and utilization of a sorghum transcriptome compendium for gene regulatory network analysis and crop trait engineering

Sorghum bicolor (Sorghum) is a drought and heat tolerant C4 grass crop used to produce grain, forage, biofuels, and other bioproducts. Genetic improvement of sorghum hybrid crops is aided by a large and diverse germplasm, sorghum's diploid inbreeding genetics, and a relatively small genome that has facilitated genomic research. Over the past 20 years, the sorghum research community characterized the cytogenetic and recombinant landscapes of sorghum's 10 chromosomes, sequenced and annotated the sorghum genome, and used that information to identify genes/alleles that modulate flowering time, plant height, seed shattering, and other important traits. More recently, >1000 RNA-seq transcriptome profiles were collected from 15 sorghum genotypes to help understand the genetic basis of variation in growth and development of sorghum stems, tillers, roots, and leaves, and the regulation of biosynthetic pathways that produce epicuticular wax, dhurrin, and RFOs, compounds that contribute to sorghum's resilience. Transcriptome studies were designed to identify differentially expressed genes that are co-expressed during development or in response to a treatment to enable construction of gene regulatory networks. Co-expression and network analysis identified transcription factors and their cognate binding sites in target gene promoters and signaling pathways that modulate gene regulatory networks providing gene editing targets for further trait optimization. RNA-seq data from >20 experiments targeting sorghum organs, tissues, cell types, developmental stages, and responses to environmental conditions (i.e., diel, day-length, shading, water-deficit, temperature) has been compiled in a sorghum transcriptome compendium. The goal of this resource paper is to describe compendium content, accessibility, and a compendium data analysis pipeline and to illustrate the types of information that can be derived from the compendium with a focus on the elucidation of gene regulatory networks useful for guiding the improvement of sorghum traits through gene editing.

RNA-seq↗