Search NASA⌕ Search

SEARCH · Search NASA

Results for “Research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

Database of Nonaqueous Proton-Conducting Materials

This work presents the assembly of 48 papers, representing 74 different compounds and blends, into a machine-readable database of nonaqueous proton-conducting materials. SMILES was used to encode the chemical structures of the molecules, and we tabulated the reported proton conductivity, proton diffusion coefficient, and material composition for a total of 3152 data points. The data spans a broad range of temperatures ranging from -70 to 260 °C. To explore this landscape of nonaqueous proton conductors, DFT was used to calculate the proton affinity of 18 unique proton carriers. The results were then compared to the activation energy derived from fitting experimental data to the Arrhenius equation. It was found that while the widely recognized positive correlation between the activation energy and proton affinity may hold among closely related molecules, this correlation does not necessarily apply across a broader range of molecules. This work serves as an example of the potential analyses that can be conducted using literature data combined with emerging research tools in computation and data science to address specific materials design problems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

From Data to Discovery: AI's Transformative Role in Thin Film Research

The advancement of thin film technologies is pivotal for progress in numerous fields, including energy, electronics, and quantum computing. However, the traditional trial-and-error approach to materials discovery is inherently slow and inefficient. This presentation will showcase how artificial intelligence (AI) is transforming thin film research by enabling a data-driven paradigm shift. We will highlight our past successes in applying AI to understand radiation damage in thin film oxides, demonstrating how graph analytics can unravel complex material behavior. Additionally, we will provide insights into our current work at the National Renewable Energy Laboratory, where we are leading the charge in autonomous materials science. Backed by a $14M investment in our characterization facility, we are developing AI-guided workflows that seamlessly integrate experimentation and AI-guided decision-making. By harnessing the power of AI, we aim to accelerate the discovery and design of high-performance thin films, propelling innovation across a multitude of industries.

36 MATERIALS SCIENCE↗

Enhancing Discoverability and Management of Atmospheric Data at Scale: Solutions from the ARM Data Center

The Atmospheric Radiation Measurement (ARM) is a multi-laboratory and multi-institutional U.S. Department of Energy (DOE) Office of Science National User Facility. The ARM Data Center (ADC), located at Oak Ridge National Laboratory, collects, archives, and shares vast atmospheric data crucial for climate research. The ADC manages over 7 PB of data from 460 instruments worldwide, processing it into more than 11,000 diverse data products using the Network Common Data Form (NetCDF) for machine-independent accessibility. The primary challenge addressed in this paper is the efficient management and distribution of vast and diverse datasets essential for the climate research community, enhancing accessibility through advanced tools like Data Discovery. The ADC has developed advanced infrastructure and software architecture to handle the continuous influx of heterogeneous data to enhance data discoverability, resulting in increased scientific collaboration. In 2023, users from over 34 countries downloaded and utilized ARM data, resulting in 1,455 publications. The ADC’s efforts have significantly improved the discoverability and usability of atmospheric data, fostering extensive scientific research and collaboration. This paper details the solutions implemented by the ADC team for efficient data discovery and distribution, and it demonstrates ARM’s capability of staging processed data for scientific analysis.

Shah, Chirag [ORNL] (ORCID:0000000203145737)↗

TSDC: Transportation Secure Data Center: Real-World Data for Planning, Modeling, and Analysis

The Transportation Secure Data Center is a centralized repository for high-resolution transportation data from hundreds of travel and transit surveys and studies. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. It houses surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Meanwhile, the Livewire Data Platform empowers research, industry, and academic partners to easily and securely preserve, maintain, share, discover, and gain access to transportation and mobility data. Livewire accommodates a range of datasets, including behavioral, experimental, model, analytical, and raw data at the vehicle, traveler, and system levels. Datasets support mobility research and planning spanning urban science, connected and automated vehicles, fueling and charging infrastructure, mobility decision science, multimodal transportation, vehicle efficiency, and more.

33 ADVANCED PROPULSION SYSTEMS↗

Impact of Domain Knowledge on the Property Prediction of Specialized Machine Learning Models

Developing transferable machine learning models is trending in data-driven materials research. However, how to apply such models to a specific research domain remains unclear. Here, in this work, we choose high-entropy materials as a platform with a specialized data set containing 145,323 DFT-relaxed materials. This data set is used to explore the role of domain-specific knowledge in training effective models. Our tests with three representative graph neural network architectures indicate the model complexity has much smaller influence on performance than the data itself. Specifically, the consideration of low-energy atomic ordering, structures with diverse elemental coverage, and high-order interactions significantly influences the model performance. We also find that domain knowledge-driven sampling can greatly enhance unsupervised learning techniques. This research highlights that developing specialized data sets is more beneficial than further complicating deep learning architectures. Additionally, physics-inspired sampling algorithms are crucially needed for better machine learning models for a specific materials research domain.

36 MATERIALS SCIENCE↗

Breaking the barrier of human-annotated training data for machine learning-aided plant research using aerial imagery

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

59 BASIC BIOLOGICAL SCIENCES↗

Summary Report of the 2nd RCM of the CRP on Updating Fission Yield Data for Applications

The Second Research Coordination Meeting of the IAEA Coordinated Research Project (CRP) on Updating Fission Yield Data for Applications was held in Vienna at the IAEA headquarters from 19 to 23 December 2022, with 23 international experts attending the meeting. The CRP is devoted to evaluation efforts of cumulative and independent fission yields for incident energies from the thermal point up to 14 MeV on actinide targets. Produced fission yield evaluations should include full uncertainty quantification and are expected to combine available experimental data and state-of-the-art model information. The activities undertaken within this CRP were reviewed including the assessment of newly measured data and ongoing evaluation efforts. Technical discussions and the resulting further work plan of this CRP are summarized in this report. The meeting presentations are available at: https://www-nds.iaea.org/index-meeting-crp/2RCM_FY/index.htm.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES↗

IEA Wind Task 49: Reference Site Conditions for Floating Wind Arrays

The commercial-scale deployment of floating offshore wind (FOW) projects is expected to take place in a diverse range of sites that may differ significantly from existing fixed-bottom projects. FOW farms are particularly sensitive to the water depth and the meteorological and oceanographic (metocean) and geotechnical conditions at the project site due to the wave-induced system motions and loads as well as the anchoring system constraints imposed by the seafloor conditions. Uncertainty around the site conditions will permeate through all aspects of project design, leading to suboptimal and overly conservative designs, increased costs, and adversely affected performance. As FOW expands into a global industry, metocean and geotechnical conditions will increasingly vary for projects located in different geographic regions or in far-from-shore, deep-water sites. This study represents the outputs of work package 1 of International Energy Agency Wind Task 49, which focuses on the integrated design of floating wind arrays. The primary goal of this study is to establish the type of parameters and constraints required to characterize FOW array reference sites; provide a realistic and publicly available set of reference site conditions to the FOW community as a baseline set of data for individual research projects; identify and categorize any critical gaps in the existing data or methodologies required to define reference site characteristics; and inform and support the design of reference FOW arrays. A building block concept was developed for synthesizing reference sites for the design of FOW arrays. The building blocks include three classes of site conditions focusing on the techno-economic design of FOW projects: metocean conditions, seabed conditions, and coastal infrastructure. All reference site data produced and collected in this study are publicly available.

17 WIND ENERGY↗

PreSens Dissolved Oxygen and Temperature Data, Old Woman Creek National Estuarine Research Reserve, Huron, OH, 2022-07-06 to 2023-12-15

This dataset contains collected dissolved oxygen (DO) and temperature measurements of the air, surface water, and underlying soil of a wetland at Old Woman Creek Estuarine Research Reserve in Huron, OH. Dissolved oxygen and temperature measurements were made in the first few centimeters of wetland sediment to assess how oxygen profiles changed over time with hydrological events. Measurements taken outside of the sediment were taken as comparison. Measurements were taken by hand with a PreSens Fibox 4 transmitter, paired with an oxygen dipping probe and temperature sensor, at each field outing. Measurements were taken approximately every 2 weeks around 10 AM EST. The PreSens_DO_Temp_Datafile.csv contains the recorded measurements, and the PreSens_InstallationMethods.csv describes the deployment of the sensor.

54 ENVIRONMENTAL SCIENCES↗

Livewire User Guide

The Livewire Data Platform houses a catalog of transportation- and mobility-related project data, as well as a publications database, making it easy to search and share data. It allows transportation researchers, industry, and academic partners to increase the visibility of their projects within the research community, securely share and preserve data, and leverage datasets from other projects. Public data on Livewire are open to anyone with a Livewire account. This guide will help Livewire users understand how to store project data as a data steward, as well as access data as a data consumer.

33 ADVANCED PROPULSION SYSTEMS↗

A Survey of Open-Source Tools for Transmission and Distribution Systems Research

This work presents a review of open-source electric power transmission and distribution systems analysis tools suitable for use by industry professionals and academic researchers. Due to the high complexity of the electric grid, there exist numerous tools and extensive research pertaining to nearly every aspect of the design, operation, and control of transmission and distribution networks. In addition to the commercial tools, a wide range of free, open-source tools, models, and data usable by the scientific community for related research have been developed by different organizations, including both international and US universities and national laboratories. However, due to the absence of a catalog of available tools, models and data, researchers often lack a knowledge of existing capabilities and may develop duplicative software and tools. Increasing awareness of these available resources seeks to accelerate their broader use, leading to more efficient and standardized grid analysis. This review paper (which is part of a larger survey effort that studied over 400 tools in the transmission, distribution, buildings, and electric vehicles space) outlines selected open-source resources that have been developed in power transmission and distribution systems research. It is anticipated that this work can serve as a guide for industry and academic researchers alike, ensuring that research efforts are well-channeled.

24 POWER TRANSMISSION AND DISTRIBUTION↗

CMIP7 Data Request: atmosphere priorities and opportunities

This paper presents a comprehensive overview of the Coupled Model Intercomparison Project Phase 7 (CMIP7) request for data unlocking key research avenues in atmospheric science and provides justification for the resources needed to produce this data. Topics within the CMIP7 Atmosphere Theme centre around processes and feedbacks in atmospheric science such as clouds, aerosols and atmospheric chemistry, atmospheric circulation, temperature variability and extremes, radiative forcings, and Earth system model evaluation. These topics are summarised in this paper as scientific “opportunities” which will be realised through CMIP7 experiments and Earth system model outputs. These opportunities were submitted by a thematic group of atmospheric science community representatives combined with an extended consultation process. The production of these variables will close key gaps and uncertainties identified during previous rounds of CMIP, and will be broadly used by scientific, policy, governmental, industry, and other communities that rely on climate model projections for research and decision making, including supporting the 7th Intergovernmental Panel on Climate Change Assessment Report (AR7). As an author group, we also reflect on the process used to collate this data request and make recommendations to future CMIP governance on implementing a consultation on this scale in the future.

58 GEOSCIENCES↗

Standard Operating Procedure for Optimal Deployment of Meteorological Instrumentation Within the Solar Radiation Research Laboratory: 2024 Edition

The objective of the National Renewable Energy Laboratory's (NREL's) Solar Radiation Research Laboratory (SRRL) is to collect and use high-quality solar radiation data sets for research leading to the widespread adoption of solar technologies. To appropriately populate and track the diverse array of instruments at the NREL-SRRL, NREL has established a Standard Operating Procedure (SOP) for optimal instrument deployment within the SRRL for both the Baseline Measurement System (BMS) and the Research Measurement System (RMS). Using best practices methodologies, the NREL-SRRL maintains a varied and extensive array of solar monitoring equipment to test, evaluate, and characterize the solar sensors used by federal and international agencies as well as the solar industry to determine the solar resource. The SOP provides the industry with guidance for solar resource assessment and is used for procedures in the long-term continuous monitoring of legacy instruments alongside state-of-the-art instruments. Based on the SOP, instruments are annually evaluated for continued deployment. Instruments that do not meet the SOP criteria are decommissioned, and new instruments that meet the criteria are deployed. Streamlining and optimizing the use of this facility ensures that the lab continues to be a world-leading solar calibration and measurement facility. This 2024 edition includes updates to the appendices to reflect the instrument changes from one year to another.

14 SOLAR ENERGY↗

Standard Operating Procedure for Optimal Deployment of Meteorological Instrumentation Within the Solar Radiation Research Laboratory: 2025 Edition

The objective of the National Renewable Energy Laboratory's (NREL's) Solar Radiation Research Laboratory (SRRL) is to collect and use high-quality solar radiation data sets for research leading to the widespread adoption of solar technologies. To appropriately populate and track the diverse array of instruments at the NREL-SRRL, NREL has established a Standard Operating Procedure (SOP) for optimal instrument deployment within the SRRL for both the Baseline Measurement System (BMS) and the Research Measurement System (RMS). Using best practices methodologies, the NREL-SRRL maintains a varied and extensive array of solar monitoring equipment to test, evaluate, and characterize the solar sensors used by federal and international agencies as well as the solar industry to determine the solar resource. The SOP provides the industry with guidance for solar resource assessment and is used for procedures in the long-term continuous monitoring of legacy instruments alongside state-of-the-art instruments. Based on the SOP, instruments are annually evaluated for continued deployment. Instruments that do not meet the SOP criteria are decommissioned, and new instruments that meet the criteria are deployed. Streamlining and optimizing the use of this facility ensures that the lab continues to be a world-leading solar calibration and measurement facility. This 2025 edition includes updates to the appendices to describe the current instrumentation of the NREL-SRRL.

14 SOLAR ENERGY↗

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:↗

Mondo: integrating disease terminology across communities

Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities⁠—including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology—a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene–disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.

biomedical informatics↗