Search NASA⌕ Search

SEARCH · Search NASA

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Large Language Model Integration for Knowledge Retrieval and Interaction for the DUNE Experiment

The Deep Underground Neutrino Experiment (DUNE) is a next-generation neutrino experiment that will generate an unprecedented volume of heterogeneous information-from documentation and technical notes to experimental data and reconstruction pipelines. Efficient knowledge retrieval and contextual understanding are increasingly critical for collaboration-wide productivity and onboarding. In this work, we present DUNE-GPT, a prototype framework that leverages large language models (LLMs) and retrieval-augmented generation (RAG) to enable natural-language querying of DUNE's internal documentation and technical resources. The system provides an intelligent interface for DUNE collaborators to interact with experiment-specific knowledge while maintaining data privacy and infrastructure compliance within Fermilab computing resources.

Rafique, A. [Argonne (main)]↗

CONSTRAINT-INDEPENDENT CONSTANT CTOA DETERMINATION FOR DUCTILE STABLE CRACK GROWTH

Crack tip opening angle (CTOA) has been used as a reliable fracture toughness parameter for decades to characterize stable ductile crack growth for thin-walled aerospace structures in the low-constraint conditions. Recently, the CTOA parameter was also applied to the pipeline industry, and a CTOA test standard ASTM E3039 was thus developed for testing a critical constant CTOA. Research showed that the constant CTOA can reasonably describe fracture toughness required to arrest a dynamic crack propagation for a modern gas pipeline. However, the CTOA fracture criterion requires constraint-independent CTOA toughness against stable ductile crack growth. ASTM E3039 recommends a drop weight tearing test (DWTT) specimen for CTOA testing. Since a shallow crack is used, DWTT measured CTOA may depend on constraint level at the crack tip. To understand if it is the case, this paper evaluates the critical CTOA for a set of fracture toughness tests on single edge notched bend (SENB) specimens with shallow and deep cracks based on four CTOA estimation models. In which, the Ln(P)-LLD linear fit model is similar to that used by ASTM E3039 in the CTOA calculation. Fracture test data for X80 pipeline steel and HY80 structural steel are considered in the CTOA evaluation. The results show that the four CTOA models can determine a crack size-independent constant CTOA over stable ductile crack growth for the SENB specimens. As a result, CTOA determined by ASTM E3039 is constraintindependent and transferable to use for an actual crack propagating in a gas pipeline.

Zhu, Xian-Kui↗

White Paper: Scalable Digital Twin Capabilities for Aging and Surveillance of Engineered Systems

This white paper presents a multi-year initiative to develop practical, secure, and scalable digital twin capabilities for engineered systems in aging and surveillance contexts—an approach pioneered at the National Nuclear Security Administration (NNSA) Lawrence Livermore National Laboratory (LLNL) that maps directly onto the needs and ambitions of the Navy for ship- and fleet-level digital twins. LLNL’s work in building part- and process-level digital twins for advanced manufacturing, with a vision to scale up to entire factory floors and, ultimately, enterprise-wide digital twins, offers an adaptable pathway for the Navy as it seeks to modernize lifecycle management, readiness, and predictive maintenance across ships and fleets. For our application, we integrate physics-based modeling with automated data ingestion, processing, and AI-driven calibration, creating hybrid models that are both interpretable and data responsive. We modernized legacy workflows, established centralized data infrastructure, automated experimental pipelines, and demonstrated end-to-end coupling of accelerated aging data with finite element simulations via optimization and surrogate modeling. The result is a generalizable framework that supports part-level digital twins today and lays the groundwork for future system-level twins suitable for Navy applications.

36 MATERIALS SCIENCE↗

Accelerate microstructure evolution simulation using graph neural networks with adaptive spatiotemporal resolution

Abstract Surrogate models driven by sizeable datasets and scientific machine-learning methods have emerged as an attractive microstructure simulation tool with the potential to deliver predictive microstructure evolution dynamics with huge savings in computational costs. Taking 2D and 3D grain growth simulations as an example, we present a completely overhauled computational framework based on graph neural networks with not only excellent agreement to both the ground truth phase-field methods and theoretical predictions, but enhanced accuracy and efficiency compared to previous works based on convolutional neural networks. These improvements can be attributed to the graph representation, both improved predictive power and a more flexible data structure amenable to adaptive mesh refinement. As the simulated microstructures coarsen, our method can adaptively adopt remeshed grids and larger timesteps to achieve further speedup. The data-to-model pipeline with training procedures together with the source codes are provided.

36 MATERIALS SCIENCE↗

Cryptic cycling by electroactive bacterioplankton in Trout Bog Lake

The potential for extracellular electron transfer (EET) is a prevailing genomic feature of humic lake bacterioplankton. However, there has been little evidence for the substantial ecological contribution predicted by genetics. We hypothesized that anoxygenic phototrophic electrotrophs and accompanying heterotrophic electrogens cycle dissolved organic matter (DOM) between oxidized and reduced states. We predicted that such bacterioplankton would exhibit diel-scale oscillations due to the light dependency of photosynthesis. Using Trout Bog Lake in Wisconsin, USA, as our model ecosystem, we profiled the water column with depth-discrete metagenomic, physiochemical, and electrochemical analyses. We observed variation in oxidation reduction potential (ORP) in response to sunlight, initiating at depths populated by anoxygenic phototrophs with EET genes. We developed an automated buoy to measure electric current flow between many pairs of electrodes simultaneously, observing correlation in electron consumption to sunlight. Our results, combined with published metatranscriptomic analysis, indicate the occurrence of electron cycling between phototrophic oxidation (electrotrophic metabolism) by Chlorobium and anaerobic respiration (electrogenic metabolism) by Geothrix, involving DOM. We also repeatedly observed gradual seasonal increases in hypolimnion ORP throughout summer. These diel and seasonal patterns imply that electroactive DOM mediates the ecology of electroactive bacteria in lakes, controlling humic lake methane emissions.IMPORTANCEWe investigated the physical, chemical, and redox characteristics of a bog lake and electrodes hung therein to test the hypothesis that dissolved organic matter is being cycled between oxidized and reduced states by electroactive bacterioplankton powered by phototrophy. To do so, we performed field-based analyses on multiple timescales using both established and novel instrumentation. We paired these analyses with recently developed bioinformatics pipelines for metagenomics data to investigate genes that enable electroactive metabolism and accompanying metabolisms. Our results are consistent with our hypothesis and yet upend some of our other expectations. Our findings have implications for understanding greenhouse gas emissions from lakes, including electroactivity as an integral part of lake metabolism throughout more of the anoxic parts of lakes and for a longer portion of the summer than expected. Our results also give a sense of what electroactivity occurs at given depths and provide a strong basis for future studies.

carbon emissions↗

ML–Enabled FPGA Framework for Fast Quantum State Discrimination in Mid-Circuit Measurement Regimes

Accurate and low-latency quantum state discrimination is essential for protocols involving mid-circuit measurement (MCM) and conditional feed-forward. In superconducting quantum systems, conventional readout pipelines transfer measurement data to host processors for post-processing, introducing millisecond-scale delays that far exceed qubit coherence times. To overcome this bottleneck, we present an in-situ machine learning (ML) inference engine implemented on an FPGA for real-time quantum state discrimination. Our design performs inference directly on digitized readout signals with 40 ns latency, supports both qubit and qutrit readout, and enables conditional operations without host-side intervention. This capability is critical for MCM and for feedback-driven protocols such as quantum error correction. We validate the system on superconducting transmon hardware, demonstrating robust discrimination fidelity across multiple qubit and qutrit channels. We further demonstrate conditional qutrit logic driven by FPGA-resident classification, highlighting the potential of low-latency ML-on-FPGA control for NISQ applications and scalable fault-tolerant quantum computing.

Vora, Neel [Lawrence Berkeley National Laboratory ↗

Advancing Additive Manufacturing Through Artificial Intelligence–Powered, High-Throughput, Nondestructive Characterization and Process Optimization

This Cooperative Research and Development Agreement (CRADA) between Oak Ridge National Laboratory (ORNL) and ZEISS Industrial Metrology has demonstrated the transformative potential of artificial intelligence (AI)-enabled x-ray computed tomography (XCT) to accelerate the qualification and certification of additively manufactured (AM) parts. At the core of this effort is Simurgh, an AI-powered XCT reconstruction framework jointly advanced by ORNL and ZEISS that integrates computer-aided design (CAD) models, physics-based simulations, and deep learning to overcome the long-standing challenges of metal artifact correction, long scan durations, and limited flaw detectability in dense and geometrically complex components. Simurgh enables high-throughput, high-quality 3D reconstruction from sparse and fast scans, which reduces XCT acquisition times by more than an order of magnitude and simultaneously improves defect detection limits by up to fourfold compared with industry-standard approaches. This capability reduces scan costs by more than 50%, lowers labor overhead, and makes XCT characterization economically viable for routine industrial use. By enabling reliable flaw detection in minutes rather than hours, Simurgh facilitates real-time feedback loops for process parameter optimization, which was highlighted in a recent npj Computational Materials (a Nature journal) issue. In the published study, more than 100 alloy coupons were characterized within a single day. This work represents a tenfold acceleration in the development of novel AM alloys and processes compared with conventional workflows. The ZEISS collaboration has also demonstrated the scalability of Simurgh to diverse application domains, including aerospace, nuclear, automotive, and biomedical components; in these applications, ensuring structural integrity is paramount. By drastically reducing barriers to XCT adoption, this partnership has laid the foundation for digital twins and data-driven certification pipelines and directly addressed bottlenecks in qualifying new materials and designs. Together, ORNL and ZEISS have shown that Simurgh advances the state of the art in nondestructive evaluation and aligns with the broader mission of enabling Industry 4.0 manufacturing ecosystems, in which intelligent, cost-effective, rapid quality assurance is integral to accelerating innovation and ensuring safety in critical applications.

36 MATERIALS SCIENCE↗

Advancing Additive Manufacturing Through Artificial Intelligence–Powered, High-Throughput, Nondestructive Characterization and Process Optimization

This Cooperative Research and Development Agreement (CRADA) between Oak Ridge National Laboratory (ORNL) and ZEISS Industrial Metrology has demonstrated the transformative potential of artificial intelligence (AI)-enabled x-ray computed tomography (XCT) to accelerate the qualification and certification of additively manufactured (AM) parts. At the core of this effort is Simurgh, an AI-powered XCT reconstruction framework jointly advanced by ORNL and ZEISS that integrates computer-aided design (CAD) models, physics-based simulations, and deep learning to overcome the long-standing challenges of metal artifact correction, long scan durations, and limited flaw detectability in dense and geometrically complex components. Simurgh enables high-throughput, high-quality 3D reconstruction from sparse and fast scans, which reduces XCT acquisition times by more than an order of magnitude and simultaneously improves defect detection limits by up to fourfold compared with industry-standard approaches. This capability reduces scan costs by more than 50%, lowers labor overhead, and makes XCT characterization economically viable for routine industrial use. By enabling reliable flaw detection in minutes rather than hours, Simurgh facilitates real-time feedback loops for process parameter optimization, which was highlighted in a recent npj Computational Materials (a Nature journal) issue. In the published study, more than 100 alloy coupons were characterized within a single day. This work represents a tenfold acceleration in the development of novel AM alloys and processes compared with conventional workflows. The ZEISS collaboration has also demonstrated the scalability of Simurgh to diverse application domains, including aerospace, nuclear, automotive, and biomedical components; in these applications, ensuring structural integrity is paramount. By drastically reducing barriers to XCT adoption, this partnership has laid the foundation for digital twins and data-driven certification pipelines and directly addressed bottlenecks in qualifying new materials and designs. Together, ORNL and ZEISS have shown that Simurgh advances the state of the art in nondestructive evaluation and aligns with the broader mission of enabling Industry 4.0 manufacturing ecosystems, in which intelligent, cost-effective, rapid quality assurance is integral to accelerating innovation and ensuring safety in critical applications.

36 MATERIALS SCIENCE↗

Computationally Guided and Experimentally Validated Design of Custom Chelators for Critical Mineral Recovery

Selective, high throughput separation of target critical metals from complex environments such as fly ash leachates and mining process streams presents a significant challenge for economical production. Custom chelators and sorbents are an attractive technology for selective metal extraction, however it can be difficult to predict their performance, and significant experimental efforts are often required to develop chelating technologies. Here, we present a computational strategy focused on modelling chelator-metal binding interactions and benchmark these results versus experimental data. A computational pipeline combining forcefield, semiempirical, and meta-GGA methods with a thermodynamic framework optimized for error cancellation has been developed to predict binding energies of chelator complexes towards critical mineral recovery applications. This approach, originally validated on [2.2.2] cryptates binding mono- and divalent cations, demonstrated robust predictive capabilities with an R2 of 0.850 against experimental aqueous binding energies. The workflow includes metadynamics for exploring high-dimensional potential energy surfaces and a cluster-continuum model for accurate yet computationally efficient solvation modeling. Error cancellation between solvation energies of free and chelator-coordinated ions enables faster convergence, even with finite cluster sizes. Initial studies on the cryptates revealed consistent metal-ligand coordination patterns, with systematic variations influenced by ion size and charge, highlighting key structural features linked to binding selectivity. Further studies of a proprietary chelator have resulted in identification of previously unreported selectivity towards economically significant metals, which in-house experiments have confirmed, demonstrating the feasibility of this approach. By applying this methodology to new chelators targeting critical minerals such as lithium, cobalt, nickel and other strategic metals, we aim to accelerate the discovery of next-generation chelators for efficient recovery, recycling, and separation processes. This computational framework serves as the backbone of a high-throughput design pipeline tailored for sustainable resource utilization and may be applied to a wide range of systems to meet experimental needs.

computational materials↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

UBW (USLCI-Brightway2) [SWR-25-169]

Life cycle inventory (LCI) data are critical for robust life cycle assessment (LCA), yet many widely used datasets such as the U.S. Life Cycle Inventory (USLCI) are not natively compatible with advanced modeling frameworks like Brightway2. This work presents an automated pipeline to transform USLCI data into a fully functional Brightway2 project. The workflow performs systematic data cleaning, resolves duplicate process and exchange identifiers, and applies allocation to multi-output processes. Technosphere and biosphere flows are harmonized through unit conversions and a bridge mapping to the biosphere3 database, with comprehensive logging of missing flows and cutoff issues. The resulting Brightway2 database is validated using matrix diagnostics to ensure consistency of the technosphere, and is benchmarked via life cycle impact assessment (LCIA) methods such as ReCiPe and IPCC GWP. Outputs include reproducible CSV exports of corrected processes, elementary flows, characterization factors, and LCIA results, alongside backup utilities for project sharing. This pipeline lowers barriers for integrating USLCI data into open-source LCA workflows, enabling reproducible, validated LCA inventories within the Brightway 2 framework.

Ghosh, Tapajyoti [National Laboratory of the Rocki↗

mphys-surrogate-model

This repository contains python scripts for building and studying reduced-order-modeling representations of droplet coalescence for eventual use in atmospheric models. The included data are generated from high-fidelity superdroplet methods and are utilized by machine learning pipelines to build data-driven models of droplet size distributions that evolve under coalescence. This repository further includes scripts to determine prediction (uncertainty) intervals on the data-driven model products based on conformal prediction.

Katona, JonasE [Lawrence Livermore National Labora↗

ICED: An Integrated CGRA Framework Enabling DFVS-Aware Acceleration

oarse-grained reconfigurable arrays (CGRAs) are a promising solution to enable energy-efficient acceleration of applications from different domains. By leveraging reconfiguration at the functional level, they can adapt to significantly different computational patterns. Existing CGRA mapping approaches extract instruction-level parallelism, exploit loop-pipelining opportunities, guarantee the data dependency, and target high throughput of a given loop. However, the recurrence data-dependency in the DFG and the mismatch between required and available computing/communication resources complicate the mapping, and might lead to significant unbalances in the utilization of the CGRA's tiles. This results in wasted power for tiles with low utilization. Applying dynamic voltage and frequency scaling (DVFS) can potentially solve this challenge and improve energy efficiency by adjusting voltage and frequency of different tiles independently. CGRAs have also been successful in accelerating data-dependent streaming applications. However, in these applications, the execution time of each kernel in the pipeline might dynamically vary depending on the characteristics of the input. This also leads to under-utilization of resources for the dynamically changing kernels that do not limit the application throughput. DVFS can also improve energy efficiency for these applications by dynamically changing the voltage and frequency levels of tiles that host non performance-constraining kernels. This paper proposes ICEDTEA -- an integrated DVFS-aware framework to map applications on CGRAs that support power islands. ICEDTEA proposes a CGRA architecture supporting DVFS islands at varying granularity (from a single tile to a group of tiles) and the related DVFS-aware compilation and mapping toolchain. ICEDTEA is the first work that introduces DVFS support for spatio-temporal CGRAs at power-island levels. The experimental evaluation shows that ICEDTEA improves average utilization by 2.3$\times$ and energy-efficiency by 1.32$\times$ over a conventional CGRA. With streaming applications, ICEDTEA improves energy efficiency by 1.12$\times$ over a state-of-the-art CGRA that introduces partial dynamic reconfiguration to adapt to variations in kernels' throughput.

Tan, Cheng↗

Untargeted, tandem mass spectrometry (LC/MS-MS) metaproteomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization. This package contains soil metaproteomics data in the context of site specific metagenomes from soil depth profiles in three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. These metaproteomes were collected in 2018 after 4.5 years of warming from five depth intervals (0-10 cm, 10-30 cm, 30-45 cm, 45-60 cm, 60-80 cm). For protein identification, the collected spectra were searched following a target-decoy search strategy against a database of metagenome predicted proteins (covering 96 samples from 2014 to 2021) representing the complete sequence diversity at the site. Data was searched with mass spectrometry database search tool (MS-GF+) using Pacific Northwest National Laboratory (PNNL)'s Data Management System (DMS) Processing pipeline. The metagenomes are published as part of another data package. Raw metaproteomic data and the data products from MS-GF+ are deposited in the Mass Spectrometry Interactive Virtual Environment (MassIVE) database under accession no. MSV000097826. Here we present a dataset that includes spectral counts for the detected proteins across samples (EMSL50964_BrodieAllMAGs_Globals_SC.txt), the sequences of the detected proteins, and sample metadata file that contains site information for the soil metaproteome samples.

Belowground Biogeochemistry Science Focus Area↗

Dark Energy Survey Year 6 Results: Weak Lensing and Galaxy Clustering Cosmological Analysis Framework

We present the methodology for the weak lensing and galaxy clustering analyses of the Dark Energy Survey (DES) Year 6 data set. In this work, we design and validate the analysis pipeline for the cosmic shear, galaxy clustering plus galaxy$-$galaxy lensing ($2 \times 2$pt), and the joint analysis in the $3 \times 2$pt. Our framework accounts for key theoretical uncertainties, such as baryonic feedback and galaxy bias, incorporating both linear and non-linear models. We apply scale cuts in regimes where theoretical modeling becomes unreliable. The robustness of the pipeline is validated using mock data and simulations, confirming unbiased cosmological constraints and highlighting the importance of posterior projection effects in the validation process. As a result, we deliver robust and validated analysis pipelines for cosmic shear, $2 \times 2$pt, and $3 \times 2$pt in $Λ$CDM and $w$CDM scenarios, including a well-defined set of scales suitable for real data analysis, a robust prescription for theoretical systematics, and the theoretical covariance of the signal. This comprehensive methodology also lays the groundwork for future galaxy surveys such as the Vera C. Rubin Observatory Legacy Survey of Space and Time.

Sanchez-Cid, D. [Zurich U.; Madrid, CIEMAT; Madrid↗

An analysis of parameter compression and Full-Modeling techniques with Velocileptors for DESI 2024 and beyond

In anticipation of forthcoming data releases of current and future spectroscopic surveys, we present the validation tests and analysis of systematic effects within velocileptors modeling pipeline when fitting mock data from the AbacusSummit N-body simulations. We compare the constraints obtained from parameter compression methods to the direct fitting (Full-Modeling) approaches of modeling the galaxy power spectra, and show that the ShapeFit extension to the traditional template method is consistent with the Full-Modeling method within the standard ΛCDM parameter space. We show the dependence on scale cuts when fitting the different redshift bins using the ShapeFit and Full-Modeling methods. We test the ability to jointly fit data from multiple redshift bins as well as joint analysis of the pre-reconstruction power spectrum with the post-reconstruction BAO correlation function signal. We further demonstrate the behavior of the model when opening up the parameter space beyond ΛCDM and also when combining likelihoods with external datasets, namely the Planck CMB priors. Finally, we describe different parametrization options for the galaxy bias, counterterm, and stochastic parameters, and employ the halo model in order to physically motivate suitable priors that are necessary to ensure the stability of the perturbation theory.

79 ASTRONOMY AND ASTROPHYSICS↗

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics↗

A portable application framework for energy management and information systems (EMIS) solutions using Brick semantic schema

This paper introduces a portable framework for developing, scaling and maintaining energy management and information systems (EMIS) applications using an ontology-based approach. Key contributions include an interoperable layer based on Brick schema, the formalization of application constraints pertaining metadata and data requirements, and a field demonstration. The framework allows for querying metadata models, fetching data, preprocessing, and analyzing data, thereby offering a modular and flexible workflow for application development. Its effectiveness is demonstrated through a case study involving the development and implementation of a data-driven anomaly detection tool for the photovoltaic systems installed at the Politecnico di Torino, Italy. During eight months of testing, the framework was used to tackle practical challenges including: (i) developing a machine learning-based anomaly detection pipeline, (ii) replacing data-driven models during operation, (iii) optimizing model deployment and retraining, (iv) handling critical changes in variable naming conventions and sensor availability (v) extending the pipeline from one system to additional ones.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗