Search NASA⌕ Search

SEARCH · Search NASA

Results for “data standard”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

APOLLO: a facility-scale differentiable virtual accelerator for Fermilab

As the design complexity of modern accelerators grows, there is more interest in using advanced simulations that have fast execution time or yield additional insights like gradients. The FAST/IOTA facility has been working on implementing and experimentally validating an end-to-end digital twin that is both fast and gradient-aware, allowing for rapid prototyping of new software and experiments with minimal beam time costs. Our framework integrates physics and ML codes for linac and ring simulation through a set of generic interfaces between surrogate and physics-based sections. To reproduce device inputs and outputs, system state is exposed as a deterministic discrete event simulator. Because Fermilab is undergoing control system transition, both EPICS and ACNET frontends are supported. Recently, we have begun transitioning to a new community lattice standard, PALS, as well as developing standardized infrastructure for data ingest and normalization to prepare for model calibration during FAST proton injector commissioning. We discuss implementation details as well as challenges, and future plans to extend modelling to main complex proton accelerators like PIPII and Booster.

Kuklev, Nikita [Fermilab]↗

MetaSci Tandem Mass Spectrometry Library

These data are pure standard mass spectral libraries were collected on Orbitrap (Thermo QE-HF) and Q-TOF (SCIEX 5600) mass spectrometers for 192 compounds using 6 different collision energies, respectively. Libraries were acquired while using different solvent additives, precursor adducts, and ion polarities.

47 OTHER INSTRUMENTATION↗

BETTER Together

The Standard Energy Efficiency Data (SEED) and Building Efficiency Targeting Tool for Energy Retrofits (BETTER) platforms are both developed by the Department of Energy and work better together. SEED is a database to manage building characteristics and performance data from a variety of sources. BETTER provides simple energy efficiency measure analyses based on high level data about the building or portfolio of buildings. A demonstration of each platform and their integration will be provided. The inputs for BETTER are building type, floor area, location, utility data, and whether PV shall be included in the analysis. The BETTER analysis can be manually set up through the web application or data can be uploaded with a BuildingSync XML file either directly or through the API. SEED can be the source of this data and the data can be sent to BETTER through the SEED application after the BETTER API token has been entered. The benefit of utilizing SEED is that it has connections to many other sources of data such as ENERGY STAR Portfolio Manager, Audit Template, and Salesforce. Therefore, it is likely that a user of SEED will already have the required inputs for BETTER in SEED already and can create BETTER analyses across their whole portfolio in a couple mouse clicks. This is a major time savings and enables decision makers an easy path to identify buildings that should undergo more detailed audits or retrofit pathways.

ASHRAE↗

Creating a Training Dataset for Semantic Segmentation of Canal Networks for Irrigation Modernization

Canal infrastructure has provided critical irrigation water to the western United States for over a century. To continue providing vital water resources to the semi-arid West, irrigation systems must undergo maintenance and modernization. Many canal companies are resource-constrained, and because funding opportunities often require detailed knowledge of existing infrastructure, they can struggle to secure financial capital. We address this problem by creating training data for a semantic segmentation deep learning model to map canal networks throughout the western United States. To create a diverse and robust training dataset, we labelled 1-m NAIP imagery with the locations of no canals, wet canals, and dry/vegetated canals. Since creating these datasets is time consuming, we first developed a preprocessing methodology to identify canals within our four study areas. We used NAIP imagery and provided canal centerline data to buffer, standardize, and cluster the imagery, automating the labeling process as much as possible. However, this still required manual cleaning and manual classification of canal type. Challenges arose when canals were interrupted (e.g., road culverts or piped sections) or when nearby features shared similar characteristics (e.g., irrigated fields, trees, and shadows). Combining automated preprocessing with manual refinement produced four detailed canal masks to be used in the semantic segmentation model developed by Richard Tapia.

13 - HYDRO ENERGY↗

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES↗

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL), ↗

VA Determinants of Health Data Curation Documentation FY25-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Determinants of Health Data Curation Documentation FY25-Q3

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Community Determinants of Health Data Curation Documentation FY25-Q4

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Community Determinants of Health Data Curation Documentation FY26-Q1

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗

VA Community Determinants of Health Data Curation Documentation FY26-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1 km grid) to another (e.g., U.S. Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, U.S. Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1 km grids. Some economic data may only be available at the ZIP code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., U.S. Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card↗

AI-Ready Control System for the Fermilab Accelerator Complex

Reliable, high-intensity operation of the Fermilab Accelerator Complex is critical to the success of the Long-Baseline Neutrino Facility and Deep Underground Neutrino Experiment. We describe the requirements and infrastructure necessary to support routine use of artificial intelligence and machine learning (AI/ML) in the accelerator control system. Three capabilities are identified: a machine learning operations (MLOps) framework standardizing the lifecycle of AI/ML automation from data management through deployment and monitoring; a data quality framework defining and enforcing standards required to build trustworthy AI/ML applications; and workflow integration with large language models to assist physicists, engineers, and operators with information retrieval, code development, and routine analysis. Use cases spanning beam diagnostics, beam control, and support system automation illustrate the technical requirements across the complex.

43 PARTICLE ACCELERATORS↗

Genesis Mission Data cards

As data-intensive research and artificial intelligence become central to DOE mission science, the need for machine-actionable dataset documentation has grown accordingly. However, many DOE-aligned communities, including the Office of Science, NNSA, and cross-laboratory collaborations, have developed independent metadata practices. This fragmentation creates friction for discovery, federation, and reuse across programs. To address these challenges, this talk introduces the Genesis Data Card: a shared metadata artifact developed in collaboration with a broad DOE community (Jefferson Lab and the National Lab of the Rockies, Oak Ridge, Sandia, Idaho, Berkeley, and Los Alamos). The Genesis Data Card aims to standardize dataset documentation across DOE-aligned initiatives while remaining extensible to discipline-specific needs. This talk will describe the data card template and the supporting code to validate completed data cards, using a companion LinkML schema. I'll walk through the design decisions behind the template, its alignment with existing standards, its treatment of sensitivity and governance metadata, and the phased roadmap toward lifecycle-integrated "xCards" that support autonomous discovery and reuse. The talk closes with current gaps, ongoing work, and how others can contribute datasets and feedback to the shared repository.

McSpadden, Helen [Thomas Jefferson National Accele↗

Learning broken symmetries with approximate invariance

Recognizing symmetries in data allows for significant boosts in neural network training, which is especially important where training data are limited. In many cases, however, the exact underlying symmetry is present only in an idealized dataset, and is broken in actual data, due to asymmetries in the detector, or varying response resolution as a function of particle momentum. Standard approaches, such as data augmentation or equivariant networks fail to represent the nature of the full, broken symmetry, effectively overconstraining the response of the neural network. We propose a learning model which balances the generality and asymptotic performance of unconstrained networks with the rapid learning of constrained networks. This is achieved through a dual-subnet structure, where one network is constrained by the symmetry and the other is not, along with a learned symmetry factor. In a simplified toy example that demonstrates violation of Lorentz invariance, our model learns as rapidly as symmetry constrained networks but escapes its performance limitations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Radiometric Testing of Germicidal UV Products, Round 2: Upper-Room Luminaires (CALiPER Report)

This report analyzes the independently tested performance of eight germicidal ultraviolet (GUV) upper-room luminaires marketed for use in occupied spaces and purchased between March and June 2023. This type of product is mounted to upper walls or ceilings to treat air in the portion of the room above occupants; this allows for safe use of the room when the device is operating, but requires sufficient air mixing between upper and lower portions of the room. Three of the luminaires used UV-emitting LEDs, and the remaining five luminaires used low-pressure mercury (LPM) lamps. Product testing covered radiometric and electrical performance for each luminaire. Initial performance was measured for all eight products, and four were additionally measured after 100 h and 500 h of operation. Measured performance data allowed for comparison against manufacturer or vendor claims if the tested products included such claims. Some products had no performance data available for a given quantity (e.g., UV-C output power), and only four of the eight luminaires had radiant intensity distribution data files in a standard format (e.g., IES LM-63) available for download from product websites. The lack of publicly available performance data makes it difficult for potential buyers and specifiers to identify suitable products and design GUV systems for their specific applications. When products had performance claims, they were sometimes contradictory (e.g., unexplained differences between multiple power values) or ambiguous (e.g., measurement units conflict with quantity, unclear whether luminaire power or lamp power, unclear whether UV output power or UV-C output power). Three of the eight tested luminaires had claimed output power (i.e., radiant flux) values that exceeded measured values by more than an order of magnitude. There was substantial variation in UV-C radiant efficiency, with a measured range of 0.3–1.9% for LED and 0.4–2.1% for LPM, as shown in Figure 1. For example, the LPM luminaire with 0.4% radiant efficiency would need 5 times the amount of electrical energy used by the LPM luminaire with 2.1% radiant efficiency to produce the same amount of UV-C output power. LPM luminaires that had parabolic reflectors aligned with inclined louvers exhibited substantially higher UV-C radiant efficiency than tested luminaires with other designs, potentially cutting energy use by 75%. These results indicate a substantial opportunity for more energy efficient LPM luminaire designs, while demonstrating that UV LED luminaires can offer comparable UV-C radiant efficiency in this application. This may seem surprising, given that LED emitters have lower UV-C radiant efficiency than LPM lamps, but the efficiency-throttling louvers that are generally required for LPM luminaires typically are not needed for LEDs thanks to their directionality. However, lateral beam angles (which describe beam width as viewed from above) were 41–83° for LED luminaires versus 89–110° for LPM luminaires. More luminaires may be required if their lateral beam angles are relatively small, and coverage may be poor if UV-C radiant intensity distribution (i.e., beam shape) is not considered when designing systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Search for supersymmetry using vector boson fusion signatures and missing transverse momentum in pp collisions at $\sqrt{s}$ = 13 TeV with the ATLAS detector

This paper presents a search for supersymmetric particles in models with highly compressed mass spectra, in events consistent with being produced through vector boson fusion. The search uses 140 fb −1 of proton-proton collision data at $\sqrt{s}$ = 13 TeV collected by the ATLAS experiment at the Large Hadron Collider. Events containing at least two jets with a large gap in pseudorapidity, large missing transverse momentum, and no reconstructed leptons are selected. A boosted decision tree is used to separate events consistent with the production of supersymmetric particles from those due to Standard Model backgrounds. The data are found to be consistent with Standard Model predictions. The results are interpreted using simplified models of R-parity-conserving supersymmetry in which the lightest supersymmetric partner is a bino-like neutralino with a mass similar to that of the lightest chargino and second-to-lightest neutralino, both of which are wino-like. Lower limits at 95% confidence level on the masses of next-to-lightest supersymmetric partners in this simplified model are established between 117 and 120 GeV when the lightest supersymmetric partners are within 1 GeV in mass.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Adaptable Standards for Discovery, Access, and Usability of Oak Ridge National Laboratory’s Data Portals and Catalogs

Oak Ridge National Laboratory (ORNL) is leveraging its established capabilities and subject matter expertise in data curation, governance, management, national security, and risk assessment and mitigation to support the US Department of Energy (DOE) Grid Modernization Initiative. Using standards modeled by the National Institute of Standards and Technology (NIST), the Data Curation Network (DCN), the Oak Ridge Leadership Computing Facility (OLCF), and other leading organizations in the fields of energy research, high-performance computing, and national and homeland security, ORNL seeks to provide a federated approach to research data discovery, use, and interoperability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗