Search NASA⌕ Search

SEARCH · Search NASA

Results for “data provenance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Workflow Provenance in the Computing Continuum for Responsible, Trustworthy, and Energy-Efficient AI

As Artificial Intelligence (AI) becomes more pervasive in our society, it is crucial to develop, deploy, and assess Responsible and Trustworthy AI (RTAI) models, i.e., those that consider not only accuracy but also other aspects, such as explainability, fairness, and energy efficiency. Workflow provenance data have historically enabled critical capabilities towards RTAI. Provenance data derivation paths contribute to responsible workflows through transparency in tracking artifacts and resource consumption. Provenance data are well-known for their trustworthiness helping explainability, reproducibility, and accountability. However, there are complex challenges to achieve RTAI, which are further complicated by the heterogeneous infrastructure in the computing continuum (Edge-Cloud-HPC) used to develop and deploy models. As a result, a significant research and development gap remains between workflow provenance data management and RTAI. In this paper, we present a vision of the pivotal role of workflow provenance in supporting RTAI and discuss related challenges. We present a schematic view between RTAI and provenance, and highlight open research directions.

Santos Souza, Renan↗

VA Determinants of Health Data Curation Documentation FY25-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Determinants of Health Data Curation Documentation FY25-Q3

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Community Determinants of Health Data Curation Documentation FY25-Q4

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Community Determinants of Health Data Curation Documentation FY26-Q1

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗

VA Community Determinants of Health Data Curation Documentation FY26-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1 km grid) to another (e.g., U.S. Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, U.S. Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1 km grids. Some economic data may only be available at the ZIP code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., U.S. Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS↗

Cleaned 5-Minute Resolution Air Quality and Meteorological Data from Nine TCEQ CAMS Sites in Houston, Texas (Nov 2021 – Oct 2022)

These data encompass 5-minute air monitoring and meteorological observations collected in the greater Houston, Texas metropolitan region, at nine (9) Continuous Ambient Monitoring Stations (CAMS) operated by the Texas Commission on Environmental Quality (TCEQ) between November 1, 2021 and October 31, 2022. The CAMS sites (CAMS 1, 8, 35, 45, 148, 403, 405, 410, and 1052) were chosen because their instrumentation includes measurements of PM2.5. These sites also provide continuous multi-parameter air-quality and meteorological measurements. Particulate matter (PM2.5, PM10) was sampled along with several trace gases, including ozone (O3), nitrogen oxides (NO, NO2, NOx), sulfur dioxide (SO2), and carbon monoxide (CO). The data set also contains standard surface meteorological parameters (temperature, humidity, pressure, wind speed, and wind direction). Several sites also include AutoGC-based measurements of volatile organic compounds (VOCs). Air monitoring instruments deployed at the selected sites comprise the following systems: BAM-1020 or TEOM (PM2.5), Thermo Scientific TEI 49i (O3), TEI 42i (NOx), and AutoGCs (VOCs). This data set is similar to the data included within the houairq5mX1.00 datastream, except for a few additional quality control steps. A systematic data cleaning and verification process was performed on the data set to ensure its quality and preparation for analysis. Removal of non-numeric status flags (e.g., [LIM], [QAS], [SPZ], [CAL], [PMA], [AQI], [SPN], [MAL]) was accomplished by employing rule-based string parsing to extract valid numerical values. Missing entries were set to -9999; however, invalid or anomalous values (e.g., 99999) were retained as originally reported by the TCEQ to preserve data provenance. The time sequence was verified for completeness, removal of duplicates, and uniformity at 5-minute intervals. Column labeling was standardized, and corresponding values were assessed for physical plausibility. All timestamps in the data set were reported in Coordinated Universal Time (UTC) as provided by the TCEQ. Further, the latitude and longitude coordinates were added for each CAMS site. A subset of the data (June 1–September 30, 2022) has been used in the following publication: Subba et al. 2025. “Implications of sea breeze circulations on boundary layer aerosols in the southern coastal Texas region.” EGUsphere 2025: 1–49, https://doi.org/10.5194/egusphere-2025-2659.

latitude↗

Improving the estimate of higher-order moments from lidar observations near the top of the convective boundary layer

Abstract. Ground-based lidar data have proven extremely useful for profiling the convective boundary layer (CBL). Many groups have derived higher-order moments (e.g., variance, skewness, fluxes) from high-temporal-resolution lidar data using an autocovariance approach. However, these analyses are highly uncertain near the CBL top when the depth of the CBL (zi) is changing during the analysis period. This is because the autocovariance approach is usually applied to constant height levels and the character of the eddies is changing on either side of the changing CBL top. Here, a new approach is presented wherein the autocovariance analysis is performed on a normalized height grid, with a temporally smoothed zi. Output from a large eddy simulation model demonstrates that deriving higher-order moments from time series on a normalized height grid has better agreement with the slab-averaged quantities than the moments derived from the original height grid.

Rosenberger, Tessa E. (ORCID:0000000333205873)↗

Electromagnetic production of kaons on the nucleon

Studies of the electromagnetic production of strange quarks started in the 1950s as something of a curiosity that puzzled experimentalists and theorists alike. Eventually, a nascent understanding of these processes began to take shape through the first pioneering experiments dedicated to explore photo- and electroproduction that were carried out in the period from the 1950s to the 1980s. As the datasets increased, concomitant advances in theoretical models were realized. However, these initial studies also made clear that more precise data was essential to continue to move forward. A paradigm shift occurred in the 1990s with the development of second-generation facilities at ELSA, MAMI, SPring-8, and JLab. High-intensity, high duty-factor accelerators, coupled with novel detector systems and advances in computing and readout electronics, brought nuclear physics experiments forward by orders of magnitude in counting statistics compared to the first-generation efforts. This was an utter boon to strangeness physics investigations, and to date, more than 50 dedicated experiments in kaon photo- and electroproduction have been completed at facilities around the world, leading to a host of experimental observables that have enabled significant advances in the exploration of strongly interacting systems that decay via $s\bar{s}$ quark pair creation. These data have proven to be an essential complementary pathway to study the spectrum and structure of the excited states of the nucleon, and the search for missing and exotic baryon configurations. As well, investigations in these channels are requisite for exploring hypernuclear production as a probe of the $YN$ interaction and for studies of the electromagnetic form factors of strange mesons. This review was designed to provide the first-ever in-depth overview of both the experimental and theoretical progress in the field of the electromagnetic production of strangeness. This work looks back over 70 years of past developments, discusses ongoing work and near-term plans, and details future possibilities being considered for third-generation facilities. Extensive lists of the available datasets and theoretical models are provided, together with a comprehensive supporting bibliography of the field. Throughout this work, the primary impacts of these explorations are highlighted, along with connections to a wide range of related phenomenological applications. An important goal of this review is to provide a complete, (reasonably) self-contained guide into this field prepared at a level that is relevant for both new and seasoned scientists, whether experimentalists, phenomenologists, or theorists, to better understand what has been accomplished by so many dedicated folks-each building on what has come before-and to appreciate the exciting future potential for continued studies in this area.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Testing the thermal Sunyaev-Zel’dovich power spectrum of a halo model using hydrodynamical simulations

Statistical properties of large-scale cosmological structures serve as powerful tools for constraining the cosmological properties of our Universe. Tracing the gas pressure, the thermal Sunyaev-Zel’dovich (tSZ) effect is a biased probe of mass distribution and, hence, can be used to test the physics of feedback or cosmological models. Therefore, it is crucial to develop robust modelling of hot gas pressure for applications to tSZ surveys. Since gas collapses into bound structures, it is expected that most of the tSZ signal is within halos produced by cosmic accretion shocks. Hence, simple empirical halo models can be used to predict the tSZ power spectra. In this study, we employed the HMx halo model to compare the tSZ power spectra with those of several hydrodynamical simulations: the Horizon suite and the Magneticum simulation. We examine various contributions to the tSZ power spectrum across different redshifts, including the one- and two-halo term decomposition, the amount of bound gas, the importance of different masses, and the electron pressure profiles. Our comparison of the tSZ power spectrum reveals discrepancies between the halo model and cosmological simulations that increase with redshift. We find a 20% to 50% difference between the measured and predicted tSZ angular power spectrum over the multipole range ℓ = 10 3 − 10 4 . Our analysis reveals that these differences are driven by the excess of power in the predicted two-halo term at low k and in the one-halo term at high k . At higher redshifts ( z ∼ 3), simulations indicate that more power comes from outside the virial radius than from inside, suggesting a limitation in the applicability of the halo model. We also observe differences in the pressure profiles, despite the fair level of agreement on the tSZ power spectrum at low redshift with the default calibration of the halo model. In conclusion, our study suggests that the properties of the halo model need to be carefully controlled against real or mock data to be proven useful for cosmological purposes.

Ayçoberry, Emma (ORCID:0000000292351195)↗

Using Large Language Models to help customers monitor global threat data

Large Language Models have proven adept at answering general knowledge questions. To make these generative AI tools useful to our mission customers for monitoring global threats, the data sciences team at Sandia is utilizing retrieval augmented generation (RAG) techniques to customize these models with local data. The local data we use consists of data such as research articles and patent abstracts that we've collected over the last several years using automated pipelines.

Herzer, John Andrew [Sandia National Laboratories ↗

Hyperplane decision trees as piecewise linear surrogate models for chemical process design

Recent trends in chemical engineering research point towards an increasing reliance on data-driven modeling approaches. Neural networks, for instance, have proven to be accurate when data is plentiful and high-dimensional, but in many cases, they require computationally-intensive training procedures. Here, in this work, we describe hyperplane decision trees (HT) as a highly expressive and low-compute machine learning model architecture. These models are locally linear and have linear decision boundaries, resulting in a piecewise linear model of the data. This property allows them to be converted into mixed-integer linear constraints which can be globally optimized. Our open-source PyTorch implementation of this method is a fast, flexible, and accessible way to build accurate piecewise linear models of data.

Decision trees↗

DOE CESER 6 GHz Interference Study

DOE CESER has sponsored Idaho National Laboratory (INL) to conduct an objective and independent study of potential 6 GHz interference from outdoor operation of unlicensed devices in the 6 GHz band on fixed service (FS) microwave communication links operated by electrical sector incumbents in that band. INL is collaborating with University of Notre Dame (UND), Electric Power Research Institute (EPRI), Lockard & White, Southern Company, and AT&T, to gather data with real-world 6 GHz interference experiments and identify (1) the potential for interference from unlicensed devices and (2) the interference necessary to cause harm to the incumbents. In addition to the functional assessment, a security assessment of the FCC mandated Automatic Frequency Coordination (AFC) System to regulate use of unlicensed 6 GHz standard power devices is also being conducted. A major objective is to create a science-based and defensible methodology used to produce the necessary data and to derive objective conclusions. This proven methodology can then be used to produce objective data and conclusions for other spectrum bands with similar incumbent uses including 4.4 – 4.9 GHz and 7.125 – 7.4 GHz, identified in the reconciliation bill that was adopted on July 4, 2025, as well as the National Spectrum Strategy discussions that are ongoing. This report contains 6 GHz field experiments and findings in the following real-world scenarios with commercial unlicensed standard power (SP) 6 GHz devices regulated by Automated Frequency Coordination (AFC): • University of Notre Dame (UND) Stadium with a capacity of 80,000 spectators, where Wi-Fi operating in 6 GHz has been deployed recently • Southern Company 6 GHz FS microwave link between Columbus and Fortson, Georgia Following are the following key findings from this study. 1. The AFC is under-protective of FS when line-of-sight exists along the path centerline. Data collected at Southern’s 6 GHz fixed link site shows significant erosion of as much as 21.4-24.4 dB of under-protection that can lead to potential service degradation under typical operating conditions. This first key finding is most likely the result of erroneous use of the RF propagation model. INL will collaborate with EPRI and the AFC Functional Requirements Working Group to submit a change request to the WinnForum TS-1014 standard towards correct use of the propagation model by the AFC. 2. There is additive interference effect of about 3 dB from nearly equal power interferers measured from simultaneous operation of two SP AP's operating co-channel with the FS receive from different locations along the path. This second key finding should be used to add the impact of additive interference of operation of multiple APs in the same geographical area, to the next generation of AFCs. INL will collaborate with FCC on the need for the AFC to consider additive interference. We also recommend that additional experiments are conducted on 6 GHz spectrum interference to further improve the AFC operation as the number of outdoor Wi-Fi devices continues to increase. These proposed steps and recommendations will make the co-existence of the incumbents and the 6 GHz outdoor Wi-Fi providers possible without any impact on the incumbents with a win-win outcome for all.

6 GHz↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗

Active multi-mode data analysis to improve fault diagnosis in AHUs

Faults in heating, ventilation and air conditioning systems can lead to increased energy consumption, occupant comfort issues, and reduced equipment lifetime. Commercial fault detection and diagnosis (FDD) tools has been increasingly deployed in U.S. commercial buildings. While they are helping to achieve energy efficiency and operational reliability, there remain gaps in their fault diagnostic capabilities. The diagnostic results often contain multiple distinct candidate root causes (CRCs) or offer no insight into CRCs. This study developed a novel active rule-based multi-mode data analysis method to enhance diagnostic resolution by applying proven rule sets and additional new rules to data from multiple known operational modes. The proposed method was demonstrated using enhanced air handling unit performance assessment rule sets and validated with the simulated data of two air handling units. New metrics, namely, reduced number of CRCs and improvement ratio, were developed to quantify the improvement of fault diagnostic resolution. The validation results showed that the proposed method effectively reduced the number of CRCs in contrast to analyzing data solely for a single mode of operation. It achieved a median improvement ratio of 80% in 19 test cases.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele↗