Panel Segmentation: A Python Package for Automated Solar Array Metadata Extraction Using Satellite Imagery
Not Available
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
A computer-implemented method of deep packet inspection (DPI) in a network is provided. The method comprises collecting data packets comprising a number of traffic flows from a number of devices via a number of traffic taps and classifying each traffic flow according to data about network protocol layers of the packets comprising the traffic flow. Application layer metadata is extracted from the packets. Traffic flow classification data and the extracted metadata are ingested into a data cluster and normalized. The normalized classification data and extracted metadata is then correlated to other data sets.
Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.
The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest (disk) expressed as a latitude/longitude point and distance, and a set of SEC form types from which to extract entities and relations. There are three main components to this pipeline as currently implemented: Social Network Extraction, Critical Infrastructure Network Extraction, and Inference and Fusion. First, Social Network Extraction, implemented as the `organizations_sec` component of the workflow graph queries the SEC EDGAR webservice using the list of initial companies from the configuration file. Given this, it extracts metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Second, the Critical Network Extraction component extracts entities and relations for a critical infrastructure sector. Currently, we focus on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Third, the Inference and Fusion component relates the social network graph to the critical infrastructure graph in order to understand the impact of a company within a geographic region. Relations include ownership of the EV Charging Station asset as well as maintenance/ownership of the EV payment networks. The fused network can be represented in many ways and currently we emit a knowledge graph.
Abstract-Scientific computing communities increasingly run their experiments using complex data- and compute-intensive workflows that utilize distributed and heterogeneous architectures targeting numerical simulations and machine learning, often executed on the Department of Energy Leadership Computing Facilities (LCFs). We argue that a principled, systematic approach to implementing FAIR principles at scale, including fine-grained metadata extraction and organization, can help with the numerous challenges to performance reproducibility posed by such workflows. We extract workflow patterns, propose a set of tools to manage the entire life cycle of performance metadata, and aggregate them in an HPC-ready framework for reproducibility (RECUP). We describe the challenges in making these tools interoperable, preliminary work, and lessons learned from this experiment.
The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.
Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.
U.S. Department of Energy's National Renewable Energy Laboratory (NREL) hosts one of the world's most energy-efficient HPC data centers; this system uses component-level warm-water liquid cooling to efficiently remove heat from the data center and capture it for reuse in the building or rejection to the atmosphere. Given the complexity of this system, building data-driven tools for holistically monitoring and operating the entire data center is a priority for ensuring maximal efficiency and resiliency. In this advanced smart facility, over one million metrics are recorded per minute using state-of-the-art streaming data architecture and software to capture and process the state of the system in real time. Here we detail two efforts to effectively analyze, visualize, and interpret this large volume streaming data. We have developed a novel, flexible system for identifying and visualizing individual metric anomalies and component performance across the data center through automatic metadata extraction and physically-motivated visualization for quick interpretation. Additionally, to directly connect system maintenance to data stream processing we explore a physics informed multi-metric drift and anomaly detection application to detect scale-build up in heat exchangers.
Traditional filesystems organize data in directories. These directories are typically a collection of files whose grouping is based on a single criterion, e.g., the starting date of an experiment, experiment name, beamline ID, measurement device, or instrument. However, each file in a directory can belong to several logical groups, such as a special event type, experiment condition, or a part of a selected dataset. dCache is a storage system developed to store large amounts of scientific data, used by many HEP and Photon Science experiments. With recent developments in dCache, we have introduced a concept of file tagging, which dynamically groups files with the same label into virtual directories. The file labels can be added, removed, renamed, and deleted through the admin interface or via REST API. The files in virtual directories are exposed through all protocols supported by dCache. This contribution will describe the details of the implementation for file tagging in dCache and present our future development plans on automatic metadata extractions, a feature that will significantly simplify data management. Additionally, we are exploring the future use of virtual directories as a way to translate scientific data catalogs into filesystem views for direct data analysis.
This poster discusses AI and ML topics in PV reliability and system performance. In particular, automated metadata extraction and QA for fielded solar installations is covered for the PV Fleets Project. Additionally, statistical learning topics for the PVInsight Project are addressed, as well as development of the PV Validation Hub.
Rugged terrain distorts optical remote sensing observations and subsequently impacts land cover classification and biophysical and biochemical parameter retrieval over mountainous areas. Therefore, topographic correction (TC) is a prerequisite for many remote sensing applications. Although various TC methods have been explored over the past four decades to mitigate topographic effects, a systematic and global review of these studies is still lacking. Using a multicomponent bibliometric approach, we extracted bibliometric metadata from 426 publications identified by searching titles, keywords, and abstracts for research on “topographic correction” and “topographic effects” in Scopus and Web of Science (WoS) from 1980 to 2022. Here this systematic review revealed a rapid growth in the number of TC studies since the 1980s, primarily driven by the availability of decametric-resolution remote sensing observations and digital elevation models (DEMs). Most of the research has focused on relatively low-elevation regions, with increasing attention beyond American and European regions, particularly in China. The seasonal distribution of satellite acquisition for TC showed considerable imbalance, mainly concentrated in months with favorable solar illumination conditions (e.g., May to October). Important themes emerged from the keyword analysis, including satellite sensors, DEMs, TC methods, evaluation criteria, and applications.
A repository for extracting dehydrated metadata for distribution power grid model.
This report describes what is required in a data management plan for data that needs to be protected in some fashion. Here we provide an overview of what a data manager should consider, including the data ingestion and various extract-transform-load processes, metadata considerations and documentation, to what might need to be accounted for in the event of data loss. In addition, this report includes two appendices: forms that, when filled out, make the user compliant with DOE data management requirements as well as additional requirements for handling protected data at ORNL.
The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization.This package contains Fourier transform ion cyclotron resonance mass spectrometry (21 Tesla FTICR-MS) data measured in negative and positive ionization mode from water and methanol soil extracts. Soil samples were collected in 2014/06/03 and 2018/06/04 from 3 replicated paired plots that had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. The following files are included: (1) fticr_neg_h2oMeoh_data_raw.csv: raw data from combined water (H2O) and methanol (MeOH) extracts in negative ion mode, (2) fticr_neg_h2oMeoh_data_processed.csv: processed data from combined water (H2O) and methanol (MeOH) extracts in negative ion mode, (3) fticr_neg_metadata.csv: metadata for samples/measurements in negative ion mode, (4) fticr_pos_h2oMeoh_data_raw.csv: raw data from combined water (H2O) and methanol (MeOH) extracts in positive ion mode, (5) fticr_pos_h2oMeoh_data_processed.csv: processed data from combined water (H2O) and methanol (MeOH) extracts in positive ion mode, (6) fticr_pos_metadata.csv: metadata for samples/measurements in positive ion mode.
We are developing a technique to monitor microbiological activities referred to as zero resistance ammetry, which entails the deployment of graphite electrodes in sediments. Measurement of current between electrodes of contrasting redox regimes and/or predominant terminal electron accepting processes can be used as an indicator of the extents of microbiological activity. We deployed an electrode array at depths of 2 mm, 4 mm, 76 mm, 78 mm, 152 mm, 154 mm, 227 mm, and 229 mm below the wetland sediment water interface in the Old Woman Creek National Estuarine Research Center, Huron, OH, USA (Lat. = 41.380833, Long. = -82.508889). A core was collected from adjacent sediment and subsamples were collected from depth intervals of 0 – 25 mm, 25 – 127 mm, 127 – 128 mm, and below 178 mm. To determine if the microbial communities attached to the electrodes were reflective of the adjacent sediment-associated microbial community, we conducted a 16S rRNA gene-based (V4 region) survey of these respective materials. This data package contains the results of these surveys, including metadata on the depths from which samples were collected (samples.csv), DNA extraction and sequencing information (OWC_DEPTH_AMPLICON_SEQUENCING_METADATA), sequence processing information (OWC_DEPTH_BIOINFORMATIC_METADATA.csv), an operational taxonomic unit (OTU) table (OWC_DEPTH_97OTUS_TABLE.csv), and nucleotide sequences of OTUs (OWC_DEPTH_97OTUS_SEQS.fasta). All files can be opened using a text-editing application. The fasta file is compatible with bioinformatics applications.
SAND2026-16981O Bibcheck is designed to extract bibliographies from research papers and perform metadata searches to identify errors. It assists authors in checking their bibliographies for metadata errors during the writing process and helps reviewers identify errors in bibliographies of papers under review. The software uses large language models (LLMs) to extract bibliography entries from PDF documents, classifies the type of bibliography entry, and verifies referenced works. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.
This data set contains measurements of soil characteristics (aggregate size distribution and mean size, total aggregate associated carbon, extractable organic C, and microbial biomass C), microbial respiration, and soil metabolite concentrations from a transient and steady soil moisture incubation experiment using soils of different textures (sandy, loamy, and clayey). The study investigated mechanisms driving the Birch effect (increased carbon mineralization pulses with wetting following a drying period) in differing soil textures. Three different soils of distinctly different textures were collected from 0-15cm depth in Georgia (sandy, 2017-05-01), Missouri (loamy, 2017-06-14), and Texas (clayey, December 2017). Soils were incubated for 140 days with destructive harvests done on days 1, 29, 33, 56, 112, 116, and 140 in transient soil moisture incubation and on days 1, 33, 116, and 140 in steady state soil moisture incubation. This dataset contains six data files in comma separate (.csv) format. Additional metadata are provided: six data dictionaries and a file-level metadata file in comma separate (*.csv) format and a user guide in PDF (*.pdf) format.
This data contains data from 90-day long incubation study which aimed to look at the soil moisture-texture relationship on soil organic carbon (SOC) cycling. Soils were collected from three distinct soil textures from mixed forests in 2017: sandy (Georgia, 2017-05-01), loamy (Missouri, 2017-06-14) and clayey (Texas, December 2017) were incubated at different soil moisture levels (air-dried, 25% water holding capacity (WHC), 50% WHC, 100% WHC and 175% WHC) at room temperature for a period of 90 days. Files contain microbial respiration, active and slow SOC pools, and their respective mineralization rates, extractable organic carbon (C), and C-acquiring extracellular enzymes. Findings from these data were used in Singh et al. (2021). This study aimed to examine the interactive effect of soil moisture and texture on SOC mineralization. Soil samples of three distinct textures (sandy, loamy, and clayey) were collected from mixed forests of Georgia, Missouri, and Texas, respectively. Soil cores of 5 cm diameter were collected from numerous random locations at each site from 0-15 cm depth after scraping the litter layer and mixed thoroughly to obtain a composite sample per site. Three additional soil cores were collected to determine the WHC using pressure plate extractors. Soil samples were composited, and triplicate soil samples were incubated in mason jars for a period of 90 days at room temperature under different moisture regimes: air dried, 25% WHC, 50% WHC, at WHC and 100% saturation. Soil respiration was measured weekly, and destructive sampling was conducted at 1, 15, 60, and 90 days to determine extractable organic C, C acquiring enzyme activity, and active and slow SOC pools with their respective mineralization rates. The C acquiring enzyme activity was the total activity of α-glucosidase, β-glucosidase, cellobiohydrolase, and β-xylosidase enzymes. Gas samples for microbial respiration measurements were collected from headspace of incubation jars through the sampling ports on the lids and then analyzed using a Shimadzu Gas Chromatograph (GC-2014). Prior to sampling, the vials were evacuated. Blank correction was also done by collecting gas samples from empty incubation jars. Double pool exponential decay model was used in SigmaPlot to determine the active and slow SOC pools and their mineralization rates (Farrar et al., 2012; Jagadamma et al., 2014). The C-acquiring extracellular enzymes were measured using the microplate method by German et al., (2011). Microbial community structure was determined using the phospholipid fatty acid (PLFA) and neutral lipid fatty acid (NLFA) analyses (Buyer and Sasser, 2012). This dataset has seven data files provided in comma-separate (*.csv) format. Additional metadata are provided: seven data dictionaries and a file-level metadata file in comma separate (*.csv) format and a user guide in PDF (*.pdf) format.