Search NASA⌕ Search

SEARCH · Search NASA

Results for “data standard”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

FC Site 4.0 - NLR Scanning Lidar (Halo XR+ 235) / Standardized and Quality-Controlled Data

This dataset contains lidar data that have been standardized and quality-controlled through NatLabRockies/FIEXTA/LiDARGO (https://github.com/NatLabRockies/FIEXTA/tree/main/lidargo). Standardization rearranges the lidar data into convenient range vs beamID vs scanID coordinates that facilitate data analysis. The scan geometry (i.e., azimuth, elevation) is shifted on a regular grid based on the most likely angles within the scan file. Quality control of radial wind speed is performed through a generalized version of the dynamic lidar filter (Beck and Kuhn, 2017).

17 WIND ENERGY↗

FC Site 4.2 - NLR Scanning Lidar (Halo XR+ 199) / Standardized and Quality-Controlled Data

This dataset contains lidar data that have been standardized and quality-controlled through NatLabRockies/FIEXTA/LiDARGO (https://github.com/NatLabRockies/FIEXTA/tree/main/lidargo). Standardization rearranges the lidar data into convenient range vs beamID vs scanID coordinates that facilitate data analysis. The scan geometry (i.e., azimuth, elevation) is shifted on a regular grid based on the most likely angles within the scan file. Quality control of radial wind speed is performed through a generalized version of the dynamic lidar filter (Beck and Kuhn, 2017).

17 WIND ENERGY↗

FC Site 1.9 - NLR Scanning Lidar (Halo XR+ 200) / Standardized and Quality-Controlled Data

This dataset contains lidar data that have been standardized and quality-controlled through NatLabRockies/FIEXTA/LiDARGO (https://github.com/NatLabRockies/FIEXTA/tree/main/lidargo). Standardization rearranges the lidar data into convenient range vs beamID vs scanID coordinates that facilitate data analysis. The scan geometry (i.e., azimuth, elevation) is shifted on a regular grid based on the most likely angles within the scan file. Quality control of radial wind speed is performed through a generalized version of the dynamic lidar filter (Beck and Kuhn, 2017).

17 WIND ENERGY↗

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Accessible Content Optimization for Research Needs (ACORN)

ACORN employs a set of automated processes for informing and/or enforcing defined content schemas to create standardized and highly structured data. Because of its standardized data source, ACORN easily applies computer automation to generate communication assets such as PDFs, Powerpoint presentations, and web pages. Built using the memory-safe Rust programming language, ACORN is portable and accessible for use on any Windows, Mac, or Linux machine.

Wohlgemuth, JasonHoward [Oak Ridge National Labora↗

LTE Electrolyzer Data Collection

The goal for NREL is to collect, develop and publish performance metrics relative to low temperature electrolyzer installations. This will be done through the development of: Secure storage solution to house the collection of data from multiple projects Standardization of data to be collected and analyzed. This will be done using data templates developed with the help of partners involved with electrolyzer installations. Analysis that produces metrics of interest for all stakeholders Aggregation of results from multiple projects to view industry progress as a whole Publication of aggregated results in the form of composite data products (CDPs) Collaboration with Idaho National Lab and their work with high temperature electrolyzer installations will enable efficient use of storage and analysis tools.

data↗

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration↗

Reference Correlations for the Density and Viscosity of Molten Alkali and Alkaline Earth Fluoride Salts

While there is a significant body of literature pertaining to thermophysical property measurements of molten salts, there is often a wide degree of variability among independent measurements of the same compounds. As such, the scientific community benefits greatly from an unbiased, independent assessment of duplicate datasets, so that reference correlations which describe these thermophysical properties as functions of temperature can be determined and then commonly used by researchers, scientists, and engineers. With regard to molten fluoride compounds, a significant time has elapsed since density and viscosity reference correlations have been determined; Janz conducted the most recent effort, in 1988, to provide reference correlations for the densities and viscosities of molten fluoride compounds via the National Standard Reference Data System coordinated by the National Bureau of Standards. Since then, new data have been published for molten fluoride compounds, and a new precedent has surfaced for putting forth reference correlations that involve fitting to multiple primary datasets. In this work, reference correlations are put forth for molten alkali and alkaline earth fluoride compounds in an effort to provide updated, improved correlations for general use. For molten alkali fluoride densities, estimated uncertainties with a 95% confidence interval are summarized as follows: LiF (0.63%), NaF (0.48%), KF (0.76%), RbF (0.93%), and CsF (0.75%). For molten alkaline earth fluoride densities, an estimated uncertainty was not able to be quantified for BeF 2 because of limited data; however, estimated uncertainties with a 95% confidence interval are summarized as follows for the remaining alkaline earth fluorides: MgF 2 (1.5%), CaF 2 (0.92%), SrF 2 (1.6%), and BaF 2 (0.23%). For molten alkali fluoride viscosities, uncertainty was not able to be quantified for RbF and CsF because of limited data; however, estimated uncertainties with a 95% confidence interval are summarized as follows for the remaining alkali fluorides: LiF (4.4%), NaF (3.0%), and KF (4.0%). For molten alkaline earth fluoride viscosities, limited consistent data resulted in the recommendation of single datasets (from literature) that are deemed to be the most trustworthy based on the quality of the underlying experimental studies.

Birri, A. [Oak Ridge National Laboratory (ORNL), O↗

Cloud-based Testbed for Adaptive Under-Frequency Load Shedding with High DER Penetration

Increasing penetration of distributed energy resources and behind-the-meter renewables may soon disrupt the efficacy of critical protection schemes, such as under-frequency load shedding (UFLS). Improved data exchange and coordination across the transmission-distribution boundary will be required to maintain reliability of bulk electric system. Standards-based data integration platforms using agreed-upon semantic vocabularies, such as the Common Information Model, will be key to enabling adaptive protection schemes requiring synthesized data from both the bulk power system and behind-the-meter resources. This paper introduces a cloud-based open-source data integration environment and UFLS clustering algorithm being developed to enable adaptive relay coordination between transmission and distribution utilities in the state of Vermont.

Anderson, Alexander A.↗

A practical approach to using the Genomic Standards Consortium MIxS reporting standard for comparative genomics and metagenomics

Comparative analysis of (meta)genomes necessitates aggregation, integration, and synthesis of well-annotated data using standards. The Genomic Standards Consortium (GSC) collaborates with the research community to develop and maintain the Minimal Information about any (x) Sequence (MIxS) reporting standard for genomic data. To facilitate use of the GSC’s MIxS reporting standard, we provide a description of the structure and terminology, how to navigate ontologies for required terms in MIxS, and demonstrate practical usage through a soil metagenome example.

standards, metadata, genome, metagenome, schema, v↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

DOE EV Data Collection - Vehicle Data

Vehicle data consist of electric vehicle performance data collected directly from the vehicle during standard operations. Data were collected using onboard data loggers that were either installed by the project team or preinstalled by the original equipment manufacturer. Data recorded by the data loggers were made accessible via an online web portal or an application programming interface. Different data loggers were used (HEM, ViriCiti, and Geotab), and the method for each vehicle is defined in the vehicle attributes file. Some systems collected data on a “trip-level” basis, in which each row of a table represents a single trip (the period between a key-on and key-off event), whereas other data were collected on a per-day basis, in which each row represents a single day of operation. Data were collected over a range of data collection periods, depending on the project. Data have been anonymized by removing information or decreasing information resolution as necessary so that fleets are not identifiable. Due to the wide range of vehicle types represented and variation in data collection, data parameters and frequencies differ between vehicles and fleets The **Performance Data Daily/Trip Data Dictionaries** contain definitions for each available parameter associated with a vehicle’s operations, aggregated at either a daily or trip level. The parameters available will vary from vehicle to vehicle, but every possible parameter will be defined. The **Vehicle Attributes Data Dictionary** contains definitions for each available parameter associated with a vehicle’s physical and functional attributes and fleet context. The **Vehicle Attributes** table contains specific vehicle characteristics, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. The **Vehicle Data** tables contain the data from each vehicle’s operations, aggregated at either a daily or trip level, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗

Commercial, industrial, and institutional discount rate estimation for efficiency standards analysis: Sector-level data 1998–2023

Underlying each of the U.S. Department of Energy’s (DOE’s) federal appliance and equipment energy conservation standards are a set of complex analyses of the projected costs and benefits of regulation. Any new or amended standard must be designed to achieve significant additional energy conservation, provided that it is technologically feasible and economically justified (42 U.S.C. 6295(o)(2)(A)). DOE determines economic justification based on whether the benefits exceed the burdens, considering a variety of factors, including the economic impact of the standard on consumers of the product and the savings in lifetime operating cost compared to any increase in price or maintenance expenses (42 U.S.C. 6295(o)(2)(B)). As part of this determination, DOE conducts a life-cycle cost (LCC) analysis, which models the combined impact of appliance first cost and operating cost changes on a representative commercial building sample to identify the fraction of customers achieving LCC savings or incurring net cost at the considered efficiency levels. Thus, the commercial discount rate value(s) used to calculate the present value of energy cost savings within the LCC model implicitly plays a role in estimating the economic impact of potential standard levels. This report provides an in-depth discussion of the commercial discount rate estimation process relying on the Capital Asset Pricing Model (CAPM) to estimate a business’ cost of equity, and by adding a risk adjustment factor to the risk-free rate associated with long-term U.S. Treasury bonds to estimate their cost of debt. It is an update to previous reports on estimating commercial discount rates from firm-level and sector-level financial data (e.g., Fujita, 2021, 2016). Major topics covered in this report include the following: • Discount rate estimation methods and rationale • Data sources used and data limitations • Discount rate distributions for use in standards analysis • Discount rate estimation methods and distributions specific to the small business subgroup analysis.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Robust Data-Driven Approach for Mechanical Serial Sectioning

Mechanical serial sectioning (MSS) provides detailed microstructural information across large length scales. By repeatedly removing thin layers of material and imaging the exposed surface, a 3D representation of a specimen’s internal structure can be constructed, enabling failure analysis and feature identification that are otherwise inaccessible via conventional 2D or nondestructive evaluation techniques. Achieving consistent and accurate material removal can be challenging due to system variability, requiring an experienced operator to manually adjust parameters, prolonging data collection times and necessitating post-processing routines to standardize the data. Here, to address these challenges, this paper presents the employment of a one-step model predictive control (MPC) framework tailored to a run-to-run (R2R) controller. The R2R-MPC controller automates the parameter selection process, improving the consistency of material removal through iterative feedback for disturbance rejection and accurate tracking of the target removal rate. Using a data-driven approach, the controller robustly adapts to changing material characteristics. The effectiveness of the R2R-MPC controller is demonstrated through simulation and experimental results and compared to previous data collection procedures.

3D Materials Science↗

Data Readiness for AI: A 360-Degree Survey

Artificial Intelligence (AI) applications critically depend on data. Poor-quality data produces inaccurate and ineffective AI models that may lead to incorrect or unsafe use. Evaluation of data readiness is a crucial step in improving the quality and appropriateness of data usage for AI. R&D efforts have been spent on improving data quality. However, standardized metrics for evaluating data readiness for use in AI training are still evolving. In this study, we perform a comprehensive survey of metrics used to verify data readiness for AI training. This survey examines more than 140 papers published by ACM Digital Library, IEEE Xplore, journals such as Nature, Springer, and Science Direct, and online articles published by prominent AI experts. This survey aims to propose a taxonomy of data readiness for AI (DRAI) metrics for structured and unstructured datasets. We anticipate that this taxonomy will lead to new standards for DRAI metrics that would be used for enhancing the quality, accuracy, and fairness of AI training and inference.

97 MATHEMATICS AND COMPUTING↗