Search NASASearch

SEARCH · Search NASA

Results for “schema”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

HPC ODA Commons [SWR-26-003]

HPC ODA Commons is a community-driven platform for standardizing HPC operational data analytics. HPC sites generate enormous volumes of operational data - scheduler logs, accounting records, monitoring streams - but turning that data into actionable insight is needlessly hard. Each site builds bespoke parsers, schemas, and evaluation pipelines. Results can't be compared across institutions. Promising analytics ideas stay siloed because there's no shared language for describing the data, the experiments, or the outcomes. HPC ODA Commons fixes this by establishing community-governed contracts - versioned schemas, canonical artifacts, and benchmark recipes - that make ODA workflows discoverable, reproducible, and comparable. It pairs these standards with a practical, CLI-first toolkit that lets operators and researchers go from raw logs to standardized results without sending data off-cluster.

Menear, Kevin [National Laboratory of the Rockies

A Prototype Software to Demonstrate a Data Catalog for Hanford Environmental Datasets

Ensuring that data on long-term environmental remediation at the Hanford Site is high-quality, traceable, and easily accessible is an ongoing challenge, complicated by decades of data collection, multiple contractors maintaining data sources, and the wide range of data types. A centralized data catalog, known as the Hanford Environmental Information and Data Index (HEIDI), has been under development as part of the Hanford Environmental Data Management (HEDM) program to address these challenges. HEIDI fulfills a critical need to bring together a wide range of data types and sizes from multiple authoritative data sources, while documenting the data pedigree and quality information (i.e., traceable to the data source/originator). This document describes additional development and maturation of the HEIDI prototype. Key accomplishments included deploying the catalog software, Esri Geoportal Server, on a server accessible to Hanford Local Area Network users, conducting cybersecurity evaluations, investigating integrated authentication solutions, and conducting functional testing of the catalog prototype. The server-based deployment enabled targeted feedback, leading to enhancements including improved accessibility features and an expanded metadata schema. Specifications for the server-based deployment of the prototype catalog and the HEIDI metadata schema are provided in this document to support subsequent HEIDI deployment by the U.S. Department of Energy Richland Operations Office.

54 ENVIRONMENTAL SCIENCES

Consistent performance of large language models in rare disease diagnosis across ten languages and 4917 cases

Background Large language models (LLMs) are increasingly used medicine for diverse applications including differential diagnostic support. The training data used to create LLMs such as the Generative Pretrained Transformer (GPT) predominantly consist of English-language texts, but LLMs could be used across the globe to support diagnostics if language barriers could be overcome. Initial pilot studies on the utility of LLMs for differential diagnosis in languages other than English have shown promise, but a large-scale assessment on the relative performance of these models in a variety of European and non-European languages on a comprehensive corpus of challenging rare-disease cases is lacking. Methods We created 4917 clinical vignettes using structured data captured with Human Phenotype Ontology (HPO) terms with the Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema. These clinical vignettes span a total of 360 distinct genetic diseases with 2525 associated phenotypic features. We used translations of the Human Phenotype Ontology together with language-specific templates to generate prompts in English, Chinese, Czech, Dutch, French, German, Italian, Japanese, Spanish, and Turkish. We applied GPT-4o, version gpt-4o-2024-08-06, and the medically fine-tuned Meditron3-70B to the task of delivering a ranked differential diagnosis using a zero-shot prompt. An ontology-based approach with the Mondo disease ontology was used to map synonyms and to map disease subtypes to clinical diagnoses in order to automate evaluation of LLM responses. Findings For English, GPT-4o placed the correct diagnosis at the first rank 19.9% and within the top-3 ranks 27.0% of the time. In comparison, for the nine non-English languages tested here the correct diagnosis was placed at rank 1 between 16.9% and 20.6%, within top-3 between 25.4% and 28.6% of cases. The Meditron3 model placed the correct diagnosis within the first 3 ranks for 20.9% of cases in English and between 19.9% and 24.0% for the other nine languages. Interpretation The differential diagnostic performance of LLMs across a comprehensive corpus of rare-disease cases was largely consistent across the ten languages tested. This suggests that the utility of LLMs in clinical settings may extend to non-English clinical settings.

Artificial intelligence

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES

Layers Can Be Deceiving: A Hopping Model for Small Molecule Diffusion in TATB Crystal

Sorption of small molecule gases in materials can play a significant role in their long‐term stability and compatibility within multi‐material assemblies. While many material properties of the insensitive high explosive TATB (1,3,5‐triamino‐2,4,6‐trinitrobenzene) are well understood, very little is known regarding its permeability to gases. TATB crystal exhibits a graphitic‐like layered packing structure that evokes a mental schema in which the layers form nanoscopic channels, but it is unclear whether this structure promotes gas transport. Here, we use molecular dynamics (MD) simulations to predict transport of small molecules through TATB single crystal. An approach to fit classical force fields is developed to model TATB interactions with H 2 O, He, Ne, and Ar, which is then combined with steered MD to probe gas transport along selected directions in the crystal. We find that small molecule transport occurs via a hopping mechanism that exhibits distinct jumps between interstitial sites and is substantially faster normal to the layers as compared to through them. This result stems from the finding that intralayer junctions between adjacent TATB molecules are the most stable interstitial sites and that energetic barriers are lower for hopping between adjacent layers. An empirical model for diffusion rate based on the MD data shows that the rate decays exponentially with increasing molecular radius and is negligibly small for all molecules larger than He, including common atmospheric gases. These findings have implications for the interpretation of experiments that measure surface area, material response to extreme conditions, and are expected to help constrain models for material aging.

36 MATERIALS SCIENCE

Scaling open-weight large language models for hydropower regulatory information extraction: A systematic analysis

Information extraction from regulatory and technical documents using large language models (LLMs) involves practical trade-offs between extraction quality and computational cost. We evaluate eight open-weight LLMs spanning 0.6B–70B parameters on hydropower licensing documents and report deployment-oriented evidence under a unified extraction schema and evaluation protocol. Across the model set, we observe clear scale-dependent trends in both baseline extraction quality and the effectiveness of reflective reasoning (self-checking) under our fixed-prompt, no-augmentation setting. Mid-scale models often provide a favorable balance of accuracy and efficiency, whereas the smallest models show limited or inconsistent gains from the reasoning variants tested. Larger models achieve the highest overall F1 scores but incur substantially greater compute and infrastructure requirements. We further find that reliability failure modes can distort conventional metrics in this domain: in particular, high recall can coincide with systematic extraction errors when models fabricate values for fields that are absent from the source text, underscoring the importance of conservative null handling and evidence-grounded evaluation. Overall, our study provides a reproducible resource–performance comparison for open-weight LLM-based extraction in hydropower regulatory documentation and offers practical guidance for model selection under different deployment constraints.

Evaluation protocol

Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools

Large language models (LLMs) show promise in supporting differential diagnosis, but their performance is challenging to evaluate due to the unstructured nature of their responses, and their accuracy compared to existing diagnostic tools is not well characterized. To assess the current capabilities of LLMs to diagnose genetic diseases, we benchmarked these models on 5213 previously published case reports using the Phenopacket Schema, the Human Phenotype Ontology and Mondo disease ontology. Prompts generated from each phenopacket were sent to seven LLMs, including four generalist models and three LLMs specialized for medical applications. The same phenopackets were used as input to a widely used diagnostic tool, Exomiser, in phenotype-only mode. The best LLM ranked the correct diagnosis first in 23.6% of cases, whereas Exomiser did so in 35.5% of cases. While the performance of LLMs for supporting differential diagnosis has been improving, it has not reached the level of commonly used traditional bioinformatics tools. Future research is needed to determine the best approach to incorporate LLMs into diagnostic pipelines.

Reese, Justin T. [Lawrence Berkeley National Labor

PAVC: The foundation for a Pan-Arctic Vegetation Cover database

Field-measured Arctic vegetation cover data is essential for creating accurate, high-quality vegetation structure and composition maps. Extrapolating field data into high-resolution cover maps provides detailed, function-specific information for use in Earth System Models, vegetation classifications, and monitoring vegetation change over time and space. However, field campaigns that collect plant cover vary substantially in scope, method, and purpose, which makes them difficult to unify across data stores, and they are often not designed to meet remote sensing needs. In this work, we synthesized and harmonized field-based fractional cover data from various data stores to create a high-quality, consistent repository schema for remote sensing-based vegetation cover mapping applications. We developed a reproducible workflow for synthesizing visual estimate and point-intercept fractional cover data. The resultant Pan-Arctic Vegetation Cover (PAVC) database contains synthesized fractional cover at both the species and plant functional type levels. The latter includes absolute foliar cover for deciduous shrubs and trees, evergreen shrubs and trees, forbs, graminoids, lichen, bryophytes, and “other” vegetation, as well as absolute cover for litter and top cover for water and bare ground.

Steckler, Morgan R. [Oak Ridge National Laboratory

Design and performance of AI agents interfacing with an atomic layer deposition tool

In this work, we introduce the design of an atomic layer deposition (ALD) reactor augmented with an AI interface for autonomous materials synthesis. Our modular design encapsulates the particularities of the hardware behind a Python interface that communicates with the ALD control software via transmission control protocol. This interface is compatible with model context protocol interfaces used in agentic frameworks. We have integrated our tool with a simple AI agent that leverages a large language model to transform user-supplied queries into ALD processes that are then run in our reactor. Our approach uses a JavaScript object notation schema to encode ALD processes. Our experimental results show that the AI interface does not impose a significant overhead to our control software, at least within our fastest 10 ms scale. We also carried out a detailed evaluation of the agent performance using leading models in two classes of tasks: basic instruction and process discovery tasks, where the agent is presented with a target material and needs to identify the correct ALD process compatible with the reactor configuration. Despite the simplicity of our agent design, we observed that most of the advanced models excelled at the instruction tasks. However, only recent models, such as o1, o3, GPT-5, and Claude Opus 4, performed well in process discovery tasks. We also observed significant variability in the response for the hardest challenges. While the results obtained are promising, we identify areas where AI research could improve the performance of agents for ALD.

47 OTHER INSTRUMENTATION

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram

Populus VariantDB v3.2 facilitates CRISPR and functional genomics research

The success of CRISPR genome editing studies depends critically on the precision of guide RNA (gRNA) design. Sequence polymorphisms in outcrossing tree species pose design hazards that can render CRISPR genome editing ineffective. Despite recent advances in tree genome sequencing with haplotype resolution, sequence polymorphism information remains largely inaccessible to various functional genomics research efforts. The Populus VariantDB v3.2 addresses these challenges by providing a user-friendly search engine to query sequence polymorphisms of heterozygous genomes. The database accepts short sequences, such as gRNAs and primers, as input for searching against multiple poplar genomes, including hybrids, with customizable parameters. We provide examples to showcase the utilities of VariantDB in improving the precision of gRNA or primer design. The platform-agnostic nature of the probe search design makes Populus VariantDB v3.2 a versatile tool for the rapidly evolving CRISPR field and other sequence-sensitive functional genomics applications. The database schema is expandable and can accommodate additional tree genomes to broaden its user base.

59 BASIC BIOLOGICAL SCIENCES

Integrated System Planning: Emerging Software Requirements in the Power Industry

Power system planning software remains fragmented across organizational boundaries, with specialized tools for capacity expansion, production cost modeling, power flow, and dynamic analysis operating on incompatible data models and assumptions. This article argues that the fragmentation is not merely a technical problem but a predictable consequence of Conway's law: software architectures mirror the departmental structures within which they are developed. Regulatory milestones like Federal Energy Regulatory Commission (FERC) Order 888 formalized these divisions, but the roots trace back to the distinct engineering disciplines-mechanical, chemical, and electrical-that staffed generation and transmission planning departments in vertically integrated utilities. As the industry moves toward integrated system planning (ISP) that coordinates generation, transmission, and distribution investment decisions, the software ecosystem must evolve accordingly. We identify five categories of software requirements to enable this transition: coherent data inputs decoupled from individual applications, unified and extensible data schemas, modular component representations that support multiple abstraction levels, lifecycle management of planning datasets, and well-defined application programming interface (API) contracts that separate data exchange from algorithmic control. We examine how these requirements interact with three common workflow patterns-serial gate clearing, sequential multiapplication, and convergence oriented-and discuss the interface design principles each demands. We then outline a vision for platform-based planning architectures where specialized analytical services compose through standardized interfaces and where artificial intelligence (AI)/machine learning (ML) tools augment decision support within a disciplined software infrastructure. The practices proposed here offer a path from today's siloed tool collections toward collaborative planning ecosystems capable of handling the complexity of modern power system transformation.

24 POWER TRANSMISSION AND DISTRIBUTION

Meta‐Analysis and Regression Modeling of the Impacts of Four Indoor Environmental Quality Metrics on Office Performance

Awareness of how buildings interact with the occupant experience—especially human performance—is becoming more prevalent, as seen by increasing interest and investment in healthy built environments. However, there is a need to synthesize the wide array of existing indoor environmental assessment and performance research in a way that can translate directly to building design and operation. Existing research in this area typically focuses on a single isolated metric and has not focused on making the results utilizable by building practitioners. The aim of this research is to investigate existing office performance literature through meta‐analyses and produce regression models for four indoor environmental quality (IEQ) metrics to support critical decision‐making for building operation and renovation. To reach this aim, a literature review was conducted to identify studies that measure the impact of changing ventilation rate, temperature, horizontal illuminance, and noise level in offices on occupant task performance. This repository of field and laboratory studies was analyzed to visualize the trends between the selected IEQ metrics and task performance. The temperature, ventilation rate, and horizontal illuminance regression models showed clear improvement potential when modifying indoor conditions toward the defined high‐performance range, while the regression model for noise level was inconclusive. The discussion notes the importance of designing holistically for all components of these IEQ categories to utilize the results, for example, good filtration on outdoor air for quantifying ventilation impact and uniform overhead lighting with low contrast for quantifying horizontal illuminance impact. The novelty of this work is in considering multiple facets of the indoor environment under a single, unified analysis schema and producing IEQ‐based performance gains that can directly inform cost‐benefit analyses of building design and renovation.

60 APPLIED LIFE SCIENCES

Faraday: A High-temperature Electrolysis Data Explorer

Faraday is a high-temperature electrolysis data visualization tool, which reveals the performance of various button cells under test conditions. These tests and the resulting analytics on their data constitute a state of the industry as the US Department of Energy pushes for the production of hydrogen. Faraday leverages the Idaho National Laboratory's DeepLynx data warehouse to standardize and query button cell data. Faraday programmatically accesses this data in DeepLynx by traversing the schema, represented by a custom ontology. The user interface queries DeepLynx for timeseries data associated with specific button cells in the warehouse, and renders them using JavaScript charts. Additional charting and data analysis techniques are made possible by an auxiliary Python server.

Woodruff, Nathan

Libra

Libra is a Python package that extends support of its parent package, SQLAlchemy. Libra’s primary functionality supports the dynamic creation of object-oriented analogs of SQL tables from a variety of user-defined, text-based schema definition formats and provides quality control and analysis tools and methods

Spears, Brady

Accessible Content Optimization for Research Needs (ACORN)

ACORN employs a set of automated processes for informing and/or enforcing defined content schemas to create standardized and highly structured data. Because of its standardized data source, ACORN easily applies computer automation to generate communication assets such as PDFs, Powerpoint presentations, and web pages. Built using the memory-safe Rust programming language, ACORN is portable and accessible for use on any Windows, Mac, or Linux machine.

Wohlgemuth, JasonHoward [Oak Ridge National Labora

Cyote Insights

CyOTE Insights leverages React, Vite, Typescript, Tailwind, and Daisy UI for the Graphical User Interface. It was designed in a particular style with a dark mode and a light mode. All code is broken down into components and reusable wrapper components for efficiency. All data is stored in Deep Lynx as a central data repository using an ontology based schema. The application serves as a main endpoint for the data in the COREII and CyOTE programs. The main purpose of the application is to display historical attack data in the Operational Technology space. At the time of this writing, it supports 27 historical attack reports compiled from OSINT sources. All of the data is publicly available, but what this application offers is the ability to see many years worth of publications in a detailed dashboard. It will also support future reports that are written using the other applications in the COREII program.

Pluth, AdamJ [Idaho National Laboratory (INL), Ida

Dataset 3: A National Dataset on Actionable Items in Improving Pooled Rideshare, 2025.

Dataset 3: A National Dataset on Actionable Items in Improving Pooled Rideshare.” 2025. Dataset Description: Pooled Rideshare Acceptance Survey - Phase 3 (2025, N = 8,296). This dataset represents the third and final phase of a national survey aimed at understanding user acceptance and preferences related to pooled rideshare (PR) services in the United States. Building on insights from earlier phases, this phase expands both the sample size and the depth of analysis to support policymaking, transportation planning, and service design for sustainable mobility systems. The Phase 3 survey was administered online to a nationally representative sample of 8,296 U.S. adults. The sample includes a wide range of demographics. The survey retained core questions from previous phases while introducing 77 detailed service features (actionable items) to evaluate potential improvements to PR offerings. Each feature was designed to assess whether a specific improvement, such as enhanced safety measures, real-time ride tracking, or user training would increase participants’ willingness to adopt PR services. In addition, behavioral predictors, current rideshare habits, environmental attitudes, and perceived barriers (e.g., safety, privacy, and comfort) were captured. - Phase_3_Final - The dataset contains rows corresponding to individual respondents and columns representing survey items, demographic characteristics, and response values. The data is available in both .CSV and .SAV formats. - Phase_3_Final_MapFile - Accompanying this dataset is a data dictionary explaining each variable, value range, and coding schema. An .XLSX format of the full survey instrument is included to support interpretation and reuse of the dataset.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI