Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Metadata Entry Optimization For NASA's Biological Institutional Scientific Collection (NBISC)

The NASA Biological Institutional Sample Collection (NBISC) at NASA’s Ames Research Center is a critical resource housing non-human samples collected from spaceflight missions and ground analog studies, primarily consisting of specimens from rats, mice, and select microbes. The primary objective of NBISC is to systematically receive, document, preserve, and facilitate access to these samples for the global scientific community. NBISC promotes international collaboration and maximizes the return on investment for precious tissues from spaceflight and analog experiments. Researchers can request physical samples through an online request form and subsequent written proposal review process. This study addresses two core research objectives: streamlining the NBISC sample lifecycle processes and strategizing for managing an influx of 50,000 tissue samples from a series of cosmic radiation analog experiments carried out at the NASA Space Radiation Laboratory (NSRL) by Drs. Eleanor Chang (Lawrence Berkeley Laboratory) and Polly Blakely (SRI). The Chang/Blakely studies investigated Harderian gland (HG) tumorigenesis in mice exposed to low dose and LET radiation comprising 8 different exposure protocols in over 4000 mice. NBISC sample metadata is stored in a Laboratory Information Management System (SLIMS). To streamline sample data entry, we customize python scripts using information extracted from the individual experimental protocols. The scripts automate entry into multiple SLIMS data fields including protocol name, unique sample barcode, tissue and sub-tissue information, freezer location, sample preservation method, etc. The semi-automated procedure significantly decreases the time spent on data entry by several orders of magnitude. Automation and data organization are essential, as they free up time for curation and promotion of the collection which, in turn, increase the accessibility of samples to the broader research community. NBISC benefits from streamlined data ingestion, and the methodologies developed here are applicable to other projects which use SLIMS including the NASA Biospecimen Sharing Program and GeneLab. As of Fall 2023, plans include transferring sample data from SLIMS to public facing repositories (OSDR and NLSP), expanding the reach of the Chang/Blakely sample collection. The Human Research Program Space Radiation Element plans to transfer non-human tissues from many more investigations to NBISC in the coming year.

Biospecimen↗

Maine Ecological Forecasting III: Utilizing Earth Observations to Monitor Federally Endangered Atlantic Salmon (Salmo salar) Habitat in Maine: An Interactive Workshop

Shifting patterns in land use and land cover (LULC), temperature, and precipitation have exacerbated a rapid decline in Federally Endangered wild Atlantic salmon (Salmo salar) populations. The team at NASA DEVELOP partnered with the Maine Department of Marine Resources (DMR) and the Downeast Salmon Federation (DSF) to create a comprehensive workshop designed to demonstrate the applicability of Earth observations in examining these threats using the Penobscot, Union, and Machias Rivers as case studies. This entailed curating tutorials for acquiring and analyzing satellite data using Google Earth Engine, EarthExplorer, and Earthdata. The team demonstrated how to classify LULC in ArcGIS Pro from 1985 until 2021 using Landsat 5 Thematic Mapper (TM), Landsat 8 Operation Land Imager (OLI), Sentinel-2 MultiSpectral Instrument (MSI), and datasets from the United Stated Geological Survey (USGS) National Land Cover Database (NLCD), showing an overall transition from coniferous forests to other LULC classes. The team also demonstrated how to use historical data from Terra Moderate Resolution Imaging Spectroradiometer (MODIS) and Integrated Multi-satellite Retrievals for Global Precipitation Measurement (GPM IMERG) to generate 2021 land surface temperature (LST) and precipitation maps, respectively, showing that Maine was abnormally dry during the summer in an increasingly warm region. These workshop materials will aid the partners in integrating NASA Earth observations into their future salmon habitat restoration initiatives.

Jonathan Falciani↗

SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images

The rapid advancement of generative models has made the detection of AI-generated images a critical challenge for both research and society. Recent works have shown that most state-of-the-art fake image detection methods overfit to their training data and catastrophically fail when evaluated on curated hard test sets with strong distribution shifts. In this work, we argue that it is more principled to learn a tight decision boundary around the real image distribution and treat the fake category as a sink class. To this end, we propose SimLBR, a simple and efficient framework for fake image detection with Latent Blending Regularization (LBR). Our method significantly improves cross-generator generalization, achieving up to +24.85% accuracy and +69.62% recall on the challenging Chameleon benchmark. SimLBR is also highly efficient, training orders of magnitude faster than existing approaches. Furthermore, we emphasize the need for reliability-oriented evaluation in fake image detection, introducing risk-adjusted metrics and worst-case estimates to better assess model robustness. All the code and models are availabe at: https://github.com/mvrl/SimLBR

Dhakal, Aayush [Washington University, St. Louis]↗

Antarctic ice sheet model comparison with uncurated geological constraints shows that higher spatial resolution improves deglacial reconstructions

Accurately reconstructing past changes to the shape and volume of the Antarctic ice sheet relies on the use of physically based and thus internally consistent ice sheet modeling, benchmarked against spatially limited geologic data. The challenge in model benchmarking against geologic data is diagnosing whether model-data misfits are the result of an inadequate model, inherently noisy or biased geologic data, and/or incorrect association between modeled quantities and geologic observations. In this work we address this challenge by (i) the development and use of a new model-data evaluation framework applied to an uncurated data set of geologic constraints, and (ii) nested high-spatial-resolution modeling designed to test the hypothesis that model resolution is an important limitation in matching geologic data. While previous approaches to model benchmarking employed highly curated datasets, our approach applies an automated screening and quality control algorithm to an uncurated public dataset of geochronological observations (specifically, cosmogenic-nuclide exposure-age measurements from glacial deposits in ice-free areas). This optimizes data utilization by including more geological constraints, reduces potential interpretive bias, and allows unsupervised assimilation of new data as they are collected. We also incorporate a nested model framework in which high-resolution domains are downscaled from a continent-wide ice sheet model. We highlight the application of this framework by applying these methods to a small ensemble of deglacial ice-sheet model simulations, and demonstrate that the nested approach improves the ability of model simulations to match exposure age data collected from areas of complex topography and ice flow. We develop a range of diagnostic model-data comparison metrics to provide more insight into model performance than possible from a single-valued misfit statistic, showing that different metrics capture different aspects of ice sheet deflation.

Geosciences↗

A total of 19 months of daily weather logging on the US east coast: the WFIP3 event log

The Third Wind Forecast Improvement Project (WFIP3) is a multi-institutional field campaign designed to advance the understanding and prediction of the offshore atmospheric boundary layer along the US east coast. Extending from February 2024 through August 2025, WFIP3 combines long-term coastal and offshore measurements with targeted modeling and forecasting efforts. This data paper presents the WFIP3 event log, a curated record of 578 d of meteorological phenomena and field observations that complements the campaign's extensive high-frequency datasets. The event log provides both manually documented daily weather discussions and automatically derived indicators of atmospheric processes – including low-level jets, wind ramps, extreme wind veer, and weak wind conditions – based on observations from scanning lidars deployed at three coastal and offshore sites. The dataset offers structured metadata, standardized time and site identifiers, and consistent terminology to facilitate its integration with WFIP3's observational and modeling data products. The log supports diverse applications, from model evaluation and forecast verification to the selection of case studies on offshore boundary-layer dynamics. The WFIP3 event log is publicly available through the US Department of Energy's Wind Data Hub, providing the research community with a transparent and enduring contextual reference for the interpretation and use of WFIP3 measurements.

17 WIND ENERGY↗

A-Train Datalist - A New GES DISC Service to Allow One-Stop Shopping for A-Train Data

The currently available services at the Goddard Earth Sciences Data Information Services Center (GES DISC) only allow users to select variables from a single data set at a time. Because entire variables from a data set are often displayed, user selection of variables of interest can be overwhelming. At the American Geophysical Union (AGU) 2016 Fall Meeting, GES DISC unveiled a new service called Datalist: a collection of predefined or user-defined data variables from one or more archived data sets. Our science support team has been curating Datalists and providing added value to the user community.Originally known as Afternoon Constellation, A-Train includes six currently on polar-orbiting Earth observation satellites: OCO-2, GCOM-W1, Aqua, CALIPSO, CloudSat, and Aura, which travel a few minutes apart from each other. This constellation arrangement has enabled coordinated science observations further forming comprehensive pictures of Earth weather and climate that are readily for use in crucial studies such as climate change.GES DISC Datalists are based on the software architecture of the new GES DISC website (also unveiled at the AGU 2016 Fall Meeting). The GES DISC science support team has created a Datalist to support the A-Train Data Depot (ATDD). Using pre-defined Datalist should hopefully save users significant effort in their data searches.

A-Train data ordering↗

Ames Life Science Data Archive: Translational Rodent Research at Ames

The Life Science Data Archive (LSDA) office at Ames is responsible for collecting, curating, distributing and maintaining information pertaining to animal and plant experiments conducted in low earth orbit aboard various space vehicles from 1965 to present. The LSDA will soon be archiving data and tissues samples collected on the next generation of commercial vehicles; e.g., SpaceX & Cygnus Commercial Cargo Craft. To date over 375 rodent flight experiments with translational application have been archived by the Ames LSDA office. This knowledge base of fundamental research can be used to understand mechanisms that affect higher organisms in microgravity and help define additional research whose results could lead the way to closing gaps identified by the Human Research Program (HRP). This poster will highlight Ames contribution to the existing knowledge base and how the LSDA can be a resource to help answer the questions surrounding human health in long duration space exploration. In addition, it will illustrate how this body of knowledge was utilized to further our understanding of how space flight affects the human system and the ability to develop countermeasures that negate the deleterious effects of space flight. The Ames Life Sciences Data Archive (ALSDA) includes current descriptions of over 700 experiments conducted aboard the Shuttle, International Space Station (ISS), NASA/MIR, Bion/Cosmos, Gemini, Biosatellites, Apollo, Skylab, Russian Foton, and ground bed rest studies. Research areas cover Behavior and Performance, Bone and Calcium Physiology, Cardiovascular Physiology, Cell and Molecular Biology, Chronobiology, Developmental Biology, Endocrinology, Environmental Monitoring, Gastrointestinal Physiology, Hematology, Immunology, Life Support System, Metabolism and Nutrition, Microbiology, Muscle Physiology, Neurophysiology, Pharmacology, Plant Biology, Pulmonary Physiology, Radiation Biology, Renal, Fluid and Electrolyte Physiology, and Toxicology. These experiment descriptions and data can be accessed online via the public LSDA website (http://lsda.jsc.nasa.gov) and information can be requested via the Data Request form at http://lsda.jsc.nasa.gov/common/dataRequest/dataRequest.aspx or by contacting the ALSDA Office at: Alison.J.French@nasa.gov

Life Sciences↗

Antarctic Meteorite Classification and Petrographic Database Enhancements

The Antarctic Meteorite collection, which is comprised of over 18,700 meteorites, is one of the largest collections of meteorites in the world. These meteorites have been collected since the late 1970 s as part of a three-agency agreement between NASA, the National Science Foundation, and the Smithsonian Institution [1]. Samples collected each season are analyzed at NASA s Meteorite Lab and the Smithsonian Institution and results are published twice a year in the Antarctic Meteorite Newsletter, which has been in publication since 1978. Each newsletter lists the samples collected and processed and provides more in-depth details on selected samples of importance to the scientific community. Data about these meteorites is also published on the NASA Curation website [2] and made available through the Meteorite Classification Database allowing scientists to search by a variety of parameters. This paper describes enhancements that have been made to the database and to the data and photo acquisition process to provide the meteorite community with faster access to meteorite data concurrent with the publication of the Antarctic Meteorite Newsletter twice a year.

Todd, N. S.↗

A Study of the Curation Protocol by Sample Analysis Working Team (SAWT) in Martian Moons eXploration (MMX) Project

Japan Aerospace Exploration Agency (JAXA) will launch a spacecraft in 2024 for a sample return mission from Phobos (Martian Moons eXploration: MMX). The major scientific goals of MMX are to constrain (1) the origin of Phobos and Deimos and (2) the evolution of the Mars-moon system [1]. The touchdown operations are planned to be performed twice at different landing sites on the Phobos surface to collect > 10 g of the surface materials [2]. After the return to the Earth, the Phobos samples will be collected from the individual sample canisters and introduced to the clean chamber installed at ISAS (Institute of Space and Astronautical Science). The Sample Analysis Working Team (SAWT) of MMX designed the procedure of Phobos sample analysis mainly conducted by the initial analysis teams [3]. For the next step, the SAWT will define the procedure of the curation process (mostly non-destructive analysis) of the Phobos samples, which will be presented here. The protocols of the Phobos sample curation is illustrated in figure 1. First, the headspace gas from the sample container will be collected during the Quick Analysis phase. The Quick Analysis will be operated by the sampler and curation teams in ISAS/JAXA. The terrestrial leak and contamination from the sampling systems will be tested using a quadrupole mass spectrometer equipped with a gas sampling system. Second, the bulk Phobos sample will be observed in the clean chamber under purified-N2 gas with an ambient condition (Pre-basic Characterization). This phase will be operated by the curation team in ISAS/JAXA and the instrument team of the MMX mission. The consistency between the data from the instruments in the clean chamber and the spacecraft will then be evaluated. Subsequently, the curation will distribute the small amount of Phobos samples to the Initial analysis team of MMX to conduct the "Preliminary Examination". The objectives of the preliminary examination are to provide (1) feedback on the subsequent sample allocation process, (2) preliminary scientific results that will address parts of MMX mission goals, and (3) evaluation of the sampling system and terrestrial alteration on Phobos samples. Because multiple models are proposed for the origin of Phobos [1] (e.g., giant impact, the capture of asteroids), the chemical and mineralogical characteristics of Phobos must be assessed before the allocation of the samples to the individual initial analysis teams. Simultaneously, the curation team in JAXA will observe the individual grains and aliquots of the samples in the clean chamber (Basic Characterization).

R Fukai↗

Enabling Space Biological Knowledge Discovery Through Image and Video Data Sharing

Increased biomedical risks and challenges associated with deep space missions and experiments (cis-Lunar, Mars transit/surface) require new knowledge discovery and development of novel ecosystems. Supporting distant and long-duration missions and experiments requires biological data (from yeast, microbes, fruit flies, C. elegans, plants, crops, rodents, humans) be findable, accessible, interoperable, reusable (FAIR), and maximally open-access. As data-intensive, bioinformatic, meta-analytical, and computer-assisted approaches continue to be a centerpiece of modern research, the NASA Biological and Physical Sciences division is expanding its Open Science capabilities beyond NASA GeneLab. The NASA Ames Life Sciences Data Archive (ALSDA) is a repository which is responsible for collecting and access to space biological imagery and video, alongside tabular and environmental data. In this presentation, we will discuss strategies dealing with archiving, curating, and accessibility of images from very distinct imaging modalities (e.g., micro-computed tomography, magnetic resonance imaging, photographic images of plants, fluorescence microscopy, behavioral videos, etc.). There are two main challenges: 1. Open-source data storage and 2. Metadata related to the imagery-video. Both have been solved by leveraging two existing open-source systems. For data storage, ALSDA is utilizing components through the Open Microscopy Environment (OME), which can read most imaging proprietary formats and display on a web interface complex multidimensional images (Z stack, multi-channel, temporal, spectral). Most technical metadata from imaging modalities are captured seamlessly. For metadata capturing experimental details, ALSDA (like GeneLab) uses the ISA-Tab specification which relies on the ISA data model to order and classify metadata. The ISA data model uses a tree structure with three files to capture the metadata: The top layer is the Investigations file, the second layer is the Study file(s), and the last layer is the Assay file(s). We believe such an approach may be useful for other types of image research data from other investigators in the AGU community.

imaging↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Addressing genome scale design tradeoffs in Pseudomonas putida for bioconversion of an aromatic carbon source

Genome-scale metabolic models (GSMM) are commonly used to identify gene deletion sets that result in growth coupling and pairing product formation with substrate utilization and can improve strain performance beyond levels typically accessible using traditional strain engineering approaches. However, sustainable feedstocks pose a challenge due to incomplete high-resolution metabolic data for non-canonical carbon sources required to curate GSMM and identify implementable designs. Here we address a four-gene deletion design in the Pseudomonas putida KT2440 strain for the lignin-derived non-sugar carbon source, p-coumarate (p-CA), that proved challenging to implement. We examine the performance of the fully implemented design for p-coumarate to glutamine, a useful biomanufacturing intermediate. In this study glutamine is then converted to indigoidine, an alternative sustainable pigment and a model heterologous product that is commonly used to colorimetrically quantify glutamine concentration. Through proteomics, promoter-variation, and growth characterization of a fully implemented gene deletion design, we provide evidence that aromatic catabolism in the completed design is rate-limited by fumarase hydratase (FUM) enzyme activity in the citrate cycle and requires careful optimization of another fumarate hydratase protein (PP_0897) expression to achieve growth and production. A double sensitivity analysis also confirmed a strict requirement for fumarate hydratase activity in the strain where all genes in the growth coupling design have been implemented. Metabolic cross-feeding experiments were used to examine the impact of complete removal of the fumarase hydratase reaction and revealed an unanticipated nutrient requirement, suggesting additional functions for this enzyme. While a complete implementation of the design was achieved, this study highlights the challenge of completely inactivating metabolic reactions encoded by under-characterized proteins, especially in the context of multi-gene edits.

59 BASIC BIOLOGICAL SCIENCES↗

A new version of the RDP (Ribosomal Database Project)

The Ribosomal Database Project (RDP-II), previously described by Maidak et al. [ Nucleic Acids Res. (1997), 25, 109-111], is now hosted by the Center for Microbial Ecology at Michigan State University. RDP-II is a curated database that offers ribosomal RNA (rRNA) nucleotide sequence data in aligned and unaligned forms, analysis services, and associated computer programs. During the past two years, data alignments have been updated and now include >9700 small subunit rRNA sequences. The recent development of an ObjectStore database will provide more rapid updating of data, better data accuracy and increased user access. RDP-II includes phylogenetically ordered alignments of rRNA sequences, derived phylogenetic trees, rRNA secondary structure diagrams, and various software programs for handling, analyzing and displaying alignments and trees. The data are available via anonymous ftp (ftp.cme.msu. edu) and WWW (http://www.cme.msu.edu/RDP). The WWW server provides ribosomal probe checking, approximate phylogenetic placement of user-submitted sequences, screening for possible chimeric rRNA sequences, automated alignment, and a suggested placement of an unknown sequence on an existing phylogenetic tree. Additional utilities also exist at RDP-II, including distance matrix, T-RFLP, and a Java-based viewer of the phylogenetic trees that can be used to create subtrees.

Non-NASA Center↗

Finding Atmospheric Composition (AC) Metadata

The Atmospheric Composition Portal (ACP) is an aggregator and curator of information related to remotely sensed atmospheric composition data and analysis. It uses existing tools and technologies and, where needed, enhances those capabilities to provide interoperable access, tools, and contextual guidance for scientists and value-adding organizations using remotely sensed atmospheric composition data. The initial focus is on Essential Climate Variables identified by the Global Climate Observing System CH4, CO, CO2, NO2, O3, SO2 and aerosols. This poster addresses our efforts in building the ACP Data Table, an interface to help discover and understand remotely sensed data that are related to atmospheric composition science and applications. We harvested GCMD, CWIC, GEOSS metadata catalogs using machine to machine technologies - OpenSearch, Web Services. We also manually investigated the plethora of CEOS data providers portals and other catalogs where that data might be aggregated. This poster is our experience of the excellence, variety, and challenges we encountered.Conclusions:1.The significant benefits that the major catalogs provide are their machine to machine tools like OpenSearch and Web Services rather than any GUI usability improvements due to the large amount of data in their catalog.2.There is a trend at the large catalogs towards simulating small data provider portals through advanced services. 3.Populating metadata catalogs using ISO19115 is too complex for users to do in a consistent way, difficult to parse visually or with XML libraries, and too complex for Java XML binders like CASTOR.4.The ability to search for Ids first and then for data (GCMD and ECHO) is better for machine to machine operations rather than the timeouts experienced when returning the entire metadata entry at once. 5.Metadata harvest and export activities between the major catalogs has led to a significant amount of duplication. (This is currently being addressed) 6.Most (if not all) Earth science atmospheric composition data providers store a reference to their data at GCMD.

metadata search↗

Wildfire Segmentation From Remotely Sensed Data Using Quantum-Compatible Conditional Vector Quantized-Variational Autoencoders

Wildfires represent a critical environmental hazard with multifaceted implications for ecosystems, communities, and public health [1]. The escalating frequency and intensity of wildfires globally have intensified the urgency for robust segmentation methodologies to facilitate effective mitigation, response, and recovery strategies [2]. Accurate wildfire segmentation is pivotal for delineating fire boundaries, assessing progression patterns, and prioritizing resource allocation during emergency scenarios. Furthermore, precise segmentation enables stakeholders, including policymakers, environmental scientists, and emergency responders, to formulate evidence-based strategies, thereby minimizing socio-economic disruptions and ecological degradation. Consequently, advancing wildfire segmentation techniques through innovative technological interventions remains a paramount research imperative. Although foundational in wildfire segmentation, traditional deterministic models exhibit inherent limitations that compromise their efficacy in dynamic and uncertain environments. These models often operate on rigid algorithms prioritizing deterministic classifications, thereby overlooking the inherent complexities and uncertainties associated with wildfire behavior and satellite data variability. Such deterministic frameworks tend to produce oversimplified representations that fail to capture the intricate nuances of evolving fire dynamics, spatial heterogeneity, and environmental interactions [1]. Consequently, the deterministic approach’s propensity for uncertainty collapsing [1, 3] hampers the accuracy, reliability, and applicability of segmentation outcomes in real-world scenarios. Contrastingly, stochastic models offer a more nuanced and adaptable framework for wildfire segmentation. By integrating probabilistic elements into the modeling paradigm, stochastic approaches, particularly probabilistic approaches such as variational auto encoders (VAEs) [4], facilitate comprehensive uncertainty assessment, enabling researchers to quantify and incorporate uncertainties into segmentation outcomes effectively. This probabilistic nature empowers stochastic models to encapsulate variability, account for data inconsistencies, and adapt to evolving environmental conditions, enhancing segmentation accuracy, reliability, and robustness. Embracing stochastic methodologies thus catalyzes advancements in wildfire science by fostering a more holistic, adaptive, and resilient segmentation framework. Despite VAEs demonstrating significant promise in various applications, they come with inherent limitations that have garnered attention within the machine learning community. One of the primary drawbacks lies in their reliance on static priors, which essentially assume a fixed distribution for latent variables, thereby limiting the model’s flexibility to capture complex data structures effectively [5]. This static nature leads to suboptimal representations, especially when dealing with complex and high-dimensional data. Additionally, VAEs often struggle with generating sharp and realistic samples, a phenomenon commonly referred to as mode collapse [5, 7, 6]. Furthermore, the optimization process in VAEs, which involves balancing the reconstruction loss and the regularization term, can sometimes be challenging to fine-tune [7]. In recent efforts to address these shortcomings, alternative approaches like Vector Quantized Variational Auto encoders(VQ-VAEs) [7], address the challenges by incorporating discrete latent variables and leveraging techniques that enhance the quality and diversity of generated samples while maintaining efficient training dynamics. VQ-VAEs propose a dynamic prior distribution generation mechanism that diverges from the static priors commonly associated with traditional VAEs. This dynamic approach allows for more adaptive and context-aware latent variable representations, thereby potentially capturing complex data structures more effectively. Unlike autoregressive prior models such as PixelCNN, which, despite their ability to model dependencies across data dimensions, suffer from significant computational inefficiencies and lack flexibility in handling diverse datasets. In our work, we propose to use a generative quantum-compatible approach to help alleviate the shortcomings of autoregressive prior model in VQ-VAEs. Restricted Boltzmann Machines (RBMs) are a viable alternative prior model that can learn prior distributions in a faster and more flexible manner. In this research endeavor, we meticulously curate a state-of-the-art dataset leveraging satellite MODIS data in conjunction with VIIRS fire masks, derived from Fire Radiative Power (FRP), thereby encapsulating diverse wildfire scenarios and environmental contexts. We developed a conditional VQ-VAE architecture with the RBM prior model that is trained in a supervised manner for segmenting wildfire masks. This innovative approach synergistically harnesses deep learning capabilities, enabling the generation of segmentation maps characterized by heightened precision, granularity, and contextual relevance. Furthermore, replacing the autoregressive prior learning method proposed by the original VQ-VAE with a prior density approximation via quantum-compatible RBM facilitates expedited inference processes, augments flexibility in prior sampling, optimizes computational efficiency and establishes a groundbreaking benchmark in wildfire segmentation methodologies.

quantum machine learning↗

Impacts of PV Module Connector Failures on Cost and Performance of Utility Scale Photovoltaic Systems

The reliability, cost and performance of electrical connectors are a concern in all types of electrical systems, and demands on connectors used on photovoltaic (PV) systems include that connectors maintain electrical conductivity and physical strength, endure ultraviolet sunlight and high ambient temperature, and resist moisture and chemical intrusion over a very long (>25 year) performance period. Connector failures increase operation and maintenance (O&M) costs and reduce plant production, but connector failure can also cause safety and liability problems, which are of greater concern. This work results from a three-year collaboration between Sandia National Laboratories (SNL), the Electric Power Research Institute (EPRI), and the National Renewable Energy Laboratory (NREL) and funded by the U.S. Department of Energy (DOE) Solar Energy Technology Office (SETO) under Agreements #39035 and #38531 "Connector Reliability Across the US Solar Sector." a multi-pronged investigation of PV connector health across the US (see https://energy.sandia.gov/pvconnectors/). This report presents derivation of a Techno-Economic Analysis (TEA) that models failure modes and frequencies (how often failure occurs), estimates O&M costs and lost production associated with connector failures, and then calculates the effect that PV module connectors can have on Levelized Cost of Energy (LCOE). The model is informed with initial data from quantitative assessment of failure rates, root causes and mechanisms, in-situ diagnostics and data collection, lab-based forensics, and interviews with PV connector manufacturers and plant operators. SNL conducted site inspections at multiple utility-scale sites in different climates and subjected field samples of new, used, and degraded connectors to visual and electrical characterization. EPRI conducted metallurgical analysis of the pin and sleeve conductors to study failure-induced morphological and compositional changes. There is in general a shortage of statistically valid data, but data from PVROM database maintained by SNL was sufficient to ascertain failure rates and lost production as well as provide qualitative insight in its curated maintenance records. This report details the structure of the mathematical model but the sources of data to inform the model will continue to evolve. Analysis of a 100 MW PV plant is provided as an example of the use of the model, with results indicating that connectors are responsible for Annualized O&M Costs of $\$$71,933/year; Annualized Unit O&M Costs of $\$$0.72/kW/year; that a Reserve Account of $\$$187,220 should be available to fund repairs related to connectors; that connectors add $\$$1,494,004 to the Net Present Value of the O&M Costs (project life); and that O&M related to connectors adds about $\$$0.00088/kWh to the Levelized Cost of Energy. The impact of this model is to provide a tool to make the US solar sector more robust by quantifying and monetizing the reliability risks to utility-scale PV systems posed by poorly installed, mismatched and/or poorly designed and manufactured connectors. The TEA provides a model incorporating failure statistics, O&M cost data, and lost production into a single figure of merit, informing decisions and enabling practitioners to optimize cost and performance trade-offs. Stakeholders include connector manufacturers, system designers and equipment specifiers, standards bodies, installers and O&M providers, investors and insurance underwriters. This report supports continued growth of PV predicated on assurances that properly installed and maintained PV system connectors are safe and reliable. The project team is proposing future work including accelerated testing of connectors and expanding the approach taken here to other PV system components, such as TEA for rapid shut-down devices.

14 SOLAR ENERGY↗

Antarctic Meteorite Classification and Petrographic Database

The Antarctic Meteorite collection, which is comprised of over 18,700 meteorites, is one of the largest collections of meteorites in the world. These meteorites have been collected since the late 1970's as part of a three-agency agreement between NASA, the National Science Foundation, and the Smithsonian Institution [1]. Samples collected each season are analyzed at NASA s Meteorite Lab and the Smithsonian Institution and results are published twice a year in the Antarctic Meteorite Newsletter, which has been in publication since 1978. Each newsletter lists the samples collected and processed and provides more in-depth details on selected samples of importance to the scientific community. Data about these meteorites is also published on the NASA Curation website [2] and made available through the Meteorite Classification Database allowing scientists to search by a variety of parameters

Todd, Nancy S.↗