Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Electricity Baseline 2021

The Electricity Baseline (2021) is a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data and was created using the ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0). The Python package used the "ELCI_2021" model configuration to set the facility and generation data sources and years that were used to create this life cycle inventory, which were taken from publicly accessible datasets and automatically curated into a local data store. An archive of the data stores used in this model is available online: https://doi.org/10.18141/2569576. This model is presented in GreenDelta's openLCA schema v2 JSON-LD format (https://greendelta.github.io/olca-schema/).

Electricity; LCA; LCI; Life Cycle↗

Electricity Baseline 2020

The Electricity Baseline (2020) is a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data and was created using the ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0). The Python package used the "ELCI_2020" model configuration to set the facility and generation data sources and years that were used to create this life cycle inventory, which were taken from publicly accessible datasets and automatically curated into a local data store. An archive of the data stores used in this model is available online: https://doi.org/10.18141/2569605. This model is presented in GreenDelta's openLCA schema v2 JSON-LD format (https://greendelta.github.io/olca-schema/).

Electricity; LCA; LCI; data inventory↗

Spaceflight Environmental-Telemetry Data for Biological Science

There is a critical need for better access and visualization of spaceflight environmental telemetry and mission hardware data from sensors including relative humidity, carbon dioxide, oxygen, radiation, airflow, temperature, acceleration, and acoustics. Under the stewardship of the Ames Life Sciences Data Archive (ALSDA) and GeneLab, an effort is underway to consolidate, normalize and provide accessibility of archived mission environmental data and hardware information, with the purpose of providing important context to biological data. This effort is necessary to provide scientific context of its impact upon biological and biomedical data from spaceflight missions and experiments (genomic, metagenomic, gene expression, proteomic, metabolomic, physiological, phenomics, behavioral; tabular, imaging, video). Environmental spaceflight data is derived from dozens of sources, with various formats, and in the past year a pipeline is in development to collect, curate and present this data efficiently. In the upcoming year, a new Data Visualization Portal will utilize the standardized pipeline data to provide easy user access to compare parameters and environmental conditions between missions, locations, subjects, and durations. Environmental and hardware data enables broad accessibility and analytics, without the need for advanced data informatic expertise. Familiarity with the capabilities and limitations of a variety of existing hardware/tools is a strength that could be applied to creation of improved hardware for future ecosystems on the Moon and Mars. The intention is to make biological and environmental telemetry data maximally open-access and FAIR (findable, accessible, interoperable, reusable) for data mining-informatic approaches to support knowledge discovery necessary for low Earth orbit, cis-Lunar, Mars transit, and Mars surface missions.

Danielle K. Lopez↗

NASA Open Science Data Repository: Maximizing Spaceflight Bioscience Data

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for data re-analysis and re-use via Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). To address the challenges posed by gaining new knowledge from a vast and diverse amount of biological, health and environmental data in space, the NASA Open Science Data Repository (OSDR - osdr.nasa.gov/bio) plays a crucial role in curating and openly publishing biological data from space-related experiments. Its design incorporates successes and lessons from NASA GeneLab, encompassing not only high-throughput sequencing data but also physiological, phenotypic, and telemetry data. The OSDR makes space biological data FAIR (findable, accessible, interoperable, reusable), and facilitates effective data ingestion, dissemination, and Open Science collaborations. The OSDR also has the capability to integrate human astronaut data with state-of-the-art security and accessibility procedures. We will discuss here several strategies that NASA’s Biological and Physical Science Division have put in place to maximize the return on investment for spaceflight bioscience data.

space biology↗

Genesis Solar Wind Sample Curation Documentation

Introduction: A scientist with experience as a sample science analyst, provider of flight hardware for multiple missions, and senior engineer in an ISO 2000-rated manufacturing plant has described the timeline of key participants in any PI-led sample return mission, the breadth of the organizations involved [1,2], and, of interest to this meeting, choosing the types of data to preserve and issues of future data accessibility. This work broadens that perspective by giving similar lessons from Genesis sample curation point-of-view. Curation participation regarding data gathering was part of the mission review process from the beginning. Genesis’ story illustrates outcome of several choices about types of data to record and preserve. Precision analysis of solar wind atoms captured in pure, ultraclean substrates is the driving science goal; therefore, detailed documentation was captured from all mission and curation phases and from investigator laboratories because these processes affect the final analytical results [3]. Pre-flight: Design and fabrication of the spacecraft. Like many modern small sample return missions, Genesis was a tightly managed team integrated across science, engineering and curation. Communication across the team was excellent, and, for the most part, the hands-on engineering technicians understood the impacts of “small choices” they routinely make, and the eyes-on oversight of manufacturing processes by scientists was mindful of details. The payload was designed by the Jet Propulsion Laboratory and the spacecraft by Lockheed Martin. Solar wind collectors and instruments were fabricated by multiple vendors and laboratories. The main portion of the payload was assembled at JSC. Fabrication procedures and contamination-control data (with witness coupons) were stored primarily at JSC. The original composition, dimensions and configuration of components, results of thermal testing, etc. are still needed for interpretation of analytical data. At times, these must be estimated from secondary information acquired pre-flight. Moreover, some files (e.g., original 3-D models and early Powerpoint) cannot be opened using software. Archived curation data includes 2-D drawings, material usage lists, QA documentation and analyses of consumables used during fabrication. Important chemical information still resides in archived hardware, paints and lubricants, material coupons, cleaning coupons, environmental witness plates and reference materials from manufacturing facilities. Purity and cleanliness of collector substrates. Semi-conductor vendors provided surface cleanliness data and some purity data. Purity for specific elements of interest was verified by science team members in their laboratories [4]. Curation archived procurement and shipping records, analysis reports, and non-proprietary fabrication data. A physical archive of flight collector reference materials is maintained for future use so additional data can be collected as analytical techniques improve. These are of increased value due to the hard landing upon re-entry. Cleaning and cleanliness assessments of flight hardware. Cleaning of the science canister payload was performed at JSC in a dedicated ISO 4 cleanroom using ultrapure water (UPW). The cleanliness of this UPW was monitored throughout processing. The archive for the clean lab also includes airborne particle counts, airborne molecular and inorganic contamination measurements as well as cleanroom construction material coupons and witness coupons. Hardware cleanliness was assessed by particle counts in rinse water batches. This information is recorded in batch cleaning forms and logbooks, and are, perhaps, of decreased value due to the hard landing. Post-flight: Curation-generated data. The curation handling history of each Genesis sample is documented in a typical astromaterials sample database which captures sample location, physical description and characterization data. Samples have a “shelf life”. Crucial to the preservation of samples is ongoing documentation of the sample environment, initially under curatorial control but is now a separate facility function with requires coordination. PI-generated data. Data on sample characterization and cleaning techniques continues to be generated by sample users [5]. These are often captured in LPSC abstracts, but these “engineering” results often are not publishable as stand-alone papers. We are actively looking for ways to make this information more accessible to users. Ion implants into samples have aided science return and can be shared among investigators. These (and similar) materials should be added to the curatorial collection with appropriate process and characterization data generated externally. Summary: Complete data archives for returned astromaterial samples must be broad in types and formats, and inclusive of environmental monitoring.

Genesis↗

Expanding Biological Repository Data Available for Sharing and Knowledge Discovery

Biology has developed next-generation data science and alternative analytical approaches with methodologies which require principal investigator (PI) experimental assay data be re-used. This new approach involves mining multiple datasets at once from various hierarchical organizations of biological complexity, while concurrently evaluating how experimental factors affect endpoints of standard assays. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make findable, accessible, interoperable, and reusable (FAIR) all non-human space-relevant biological data. These data include mission metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery, and subject-experienced telemetry (radiation, temperature, humidity, acoustics, vibrations). ALSDA has transformed to bring current biological repository data and all future collected data into this new scientific data mining reality. It has integrated into the ‘NASA Open Science’ group of projects to facilitate a suite of new tools and workflows to improve data accessibility and reusability by implementing data management plans, automating data submission agreements, and adopting the single-point-of-entry data submission portal, originally developed by NASA GeneLab. These systems required ALSDA to develop science assay configurations for the submission portal, capturing essential assay parameters according to established norms in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. ALSDA datasets are curated to maintain rich metadata, accuracy of datasets, data transparency, provenance, and additionally ensure data are machine-readable (e.g., R and Python languages). ALSDA integration with GeneLab and its analysis portals enable higher-order physiological-level datasets be mined in conjunction with -omics datasets. As ALSDA physiological-level datasets are published (micro-computed tomography, histology, intraocular pressure, hormonal assays, immunostaining, ultrasonography), the merging of hierarchical organizations of biological complexity from spaceflight will enable new knowledge discovery approaches.

Ryan T Scott↗

ARES Biennial Report 2012 Final

Since the return of the first lunar samples, what is now the Astromaterials Research and Exploration Science (ARES) Directorate has had curatorial responsibility for all NASA-held extraterrestrial materials. Originating during the Apollo Program (1960s), this capability at Johnson Space Center (JSC) included scientists who were responsible for the science planning and training of astronauts for lunar surface activities as well as experts in the analysis and preservation of the precious returned samples. Today, ARES conducts research in basic and applied space and planetary science, and its scientific staff represents a broad diversity of expertise in the physical sciences (physics, chemistry, geology, astronomy), mathematics, and engineering organized into three offices (figure 1): Astromaterials Research (KR), Astromaterials Acquisition and Curation (KT), and Human Exploration Science (KX). Scientists within the Astromaterials Acquisition and Curation Office preserve, protect, document, and distribute samples of the current astromaterials collections. Since the return of the first lunar samples, ARES has been assigned curatorial responsibility for all NASA-held extraterrestrial materials (Apollo lunar samples, Antarctic meteorites - some of which have been confirmed to have originated on the Moon and on Mars - cosmic dust, solar wind samples, comet and interstellar dust particles, and space-exposed hardware). The responsibilities of curation consist not only of the longterm care of the samples, but also the support and planning for future sample collection missions and research and technology to enable new sample types. Curation provides the foundation for research into the samples. The Lunar Sample Facility and other curation clean rooms, the data center, laboratories, and associated instrumentation are unique NASA resources that, together with our staff's fundamental understanding of the entire collection, provide a service to the external research community, which relies on access to the samples. The curation efforts are greatly enhanced by a strong group of planetary scientists who conduct peerreviewed astromaterials research. Astromaterials Research Office scientists conduct peer-reviewed research as Principal or Co-Investigators in planetary science (e. g., cosmochemistry, origins of solar systems, Mars fundamental research, planetary geology and geophysics) and participate as Co-Investigators or Participating Scientists in many of NASA's robotic planetary missions. Since the last report, ARES has achieved several noteworthy milestones, some of which are documented in detail in the sections that follow. Within the Human Exploration Science Office, ARES is a world leader in orbital debris research, modeling and monitoring the debris environment, designing debris shielding, and developing policy to control and mitigate the orbital debris population. ARES has aggressively pursued refinements in knowledge of the debris environment and the hazard it presents to spacecraft. Additionally, the ARES Image Science and Analysis Group has been recognized as world class as a result of the high quality of near-real-time analysis of ascent and on-orbit inspection imagery to identify debris shedding, anomalies, and associated potential damage during Space Shuttle missions. ARES Earth scientists manage and continuously update the database of astronaut photography that is predominantly from Shuttle and ISS missions, but also includes the results of 40 years of human spaceflight. The Crew Earth Observations Web site (http://eol.jsc.nasa.gov/Education/ESS/crew.htm) continues to receive several million hits per month. ARES scientists are also influencing decisions in the development of the next generation of human and robotic spacecraft and missions through laboratory tests on the optical qualities of materials for windows, micrometeoroid/orbital debris shielding technology, and analog activities to assess surface science operations. ARES serves as host to numerous students and visiting scientists as part of the services provided to the research community and conducts a robust education and outreach program. ARES scientists are recognized nationally and internationally by virtue of their success in publishing in peer-reviewed journals and winning competitive research proposals. ARES scientists have won every major award presented by the Meteoritical Society, including the Leonard Medal, the most prestigious award in planetary science and cosmochemistry; the Barringer Medal, recognizing outstanding work in the field of impact cratering; the Nier Prize for outstanding research by a young scientist; and several recipients of the Nininger Meteorite Award. One of our scientists received the Department of Defense (DoD) Joint Meritorious Civilian Service Award (the highest civilian honor given by the DoD). ARES has established numerous partnerships with other NASA Centers, universities, and national laboratories. ARES scientists serve as journal editors, members of advisory panels and review committees, and society officers, and several scientists have been elected as Fellows in their professional societies. This biennial report summarizes a subset of the accomplishments made by each of the ARES offices and highlights participation in ongoing human and robotic missions, development of new missions, and planning for future human and robotic exploration of the solar system beyond low Earth orbit.

Stansbery, Eileen↗

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration↗

Decoding substrate specificity determining factors in glycosyltransferase-B enzymes – insights from machine learning models

Substrate specificity is an essential characteristic of any enzyme's function and an understanding of the factors that determine this specificity is crucial for enzyme engineering. Unlike the structure of an enzyme which is directly impacted by its sequence, substrate specificity as an enzyme attribute involves a rather indirect relationship with sequence as it also depends on structural aspects that dictate substrate accessibility and active site dynamics. In this study, we explore the performance of classifier-based machine learning models trained on curated sequence and structural data for a class of glycosyltransferases (GTs), namely GT-Bs, to understand their substrate specificity determining factors. GTs enable the transfer of sugar moieties to other biomolecules such as oligosaccharides or proteins and are found in all kingdoms of life. In plants, GTs participate in the biosynthesis of plant cell wall biopolymers (e.g.: hemicelluloses and pectins) and are an integral part of the enzymatic machinery that enables the storage of carbon and energy as plant biomass. To elucidate the substrate specificity of uncharacterized GT-Bs, we constructed multi-label machine learning models (Support Vector Classifier, K-Nearest Neighbors, Gaussian Naïve-Bayes, Random Forest) that incorporate both sequence and structural features. These models achieve good predictive accuracies on test datasets. However, despite our use of structural information, we highlight that there is further scope for improvement in training these models to draw interpretable relationships between sequence, structure and substrate specificity determining motifs in GT-Bs.

97 MATHEMATICS AND COMPUTING↗

Multi-strain analysis of Pseudomonas putida reveals the metabolic and genetic diversity of the species

Pseudomonas putida is a gram-negative bacterial species increasingly utilized in biotechnology due to its robust growth, ability to degrade aromatic compounds, solvent tolerance, and genetic tractability. In this study, we report a comprehensive multi-strain analysis of 164 P. putida strains based on the reconstruction of a pan-putida metabolic network and the formulation of strain-specific genome-scale metabolic models (GEMs). We performed whole-genome sequencing and hybrid assembly for 40 strains, contributing a ~8% increase to the available genomic data for P. putida . Furthermore, high-throughput phenotypic profiling using the Biolog phenotype microarray system for 24 strains on 190 unique carbon sources, along with 15 aromatic compounds not present on Biolog plates, yielded 4,920 unique strain-phenotype measurements. These data were leveraged to curate GEMs for 24 representative strains, including a refined model for strain KT2440, which comprised 1,480 genes and 2,191 metabolites, achieving a prediction accuracy of 91.2% in carbon utilization. Systematic comparison of genomes and GEMs revealed both conserved core pathways and significant allelic and functional divergence across strains, highlighting strain-specific variation in aromatic degradation. While pathways for protocatechuate and phenylacetate degradation were widely conserved, metabolic capabilities for compounds such as ferulate, phenol, and cresols varied markedly, suggesting adaptation to distinct ecological niches. Alleleome analysis of enzymes, such as PcaI and PcaJ, revealed distinct, functionally similar clades, indicating possible convergent evolution or horizontal gene transfer. These results provide computable resources and informative models for selecting P. putida strains with desired traits for biomanufacturing and bioremediation and offer insights into the evolution and phylogeny of the P. putida species.

aromatics utilization↗

Climate Data Initiative: A Geocuration Effort to Support Climate Resilience

Curation is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest and typically occurs in museums, art galleries, and libraries. The task of organizing data around specific topics or themes is a vibrant and growing effort in the biological sciences but to date this effort has not been actively pursued in the Earth sciences. In this paper, we introduce the concept of geocuration and define it as the act of searching, selecting, and synthesizing Earth science data/metadata and information from across disciplines and repositories into a single, cohesive, and useful collection. We present the Climate Data Initiative (CDI) project as a prototypical example. The CDI project is a systematic effort to manually curate and share openly available climate data from various federal agencies. CDI is a broad multi-agency effort of the U.S. government and seeks to leverage the extensive existing federal climate-relevant data to stimulate innovation and private-sector entrepreneurship to support national climate-change preparedness. We describe the geocuration process used in the CDI project, lessons learned, and suggestions to improve similar geocuration efforts in the future.

Metada↗

Climate Data Initiative: A Geocuration Effort to Support Climate Resilience

Curation is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest and typically occurs in museums, art galleries, and libraries. The task of organizing data around specific topics or themes is a vibrant and growing effort in the biological sciences but to date this effort has not been actively pursued in the Earth sciences. In this paper, we introduce the concept of geocuration and define it as the act of searching, selecting, and synthesizing Earth science data/metadata and information from across disciplines and repositories into a single, cohesive, and useful compendium We present the Climate Data Initiative (CDI) project as an exemplar example. The CDI project is a systematic effort to manually curate and share openly available climate data from various federal agencies. CDI is a broad multi-agency effort of the U.S. government and seeks to leverage the extensive existing federal climate-relevant data to stimulate innovation and private-sector entrepreneurship to support national climate-change preparedness. We describe the geocuration process used in CDI project, lessons learned, and suggestions to improve similar geocuration efforts in the future.

virtual collections↗

Murine Host-gut Microbiota Interactions are Modulated During Spaceflight

The rodent habitat on the International Space Station has provided critical insight into the impact of spaceflight on mammalian physiology. These effects include dysfunction of carbohydrate, steroid and lipid metabolism, and immune response, as well as induction of symptoms characteristic of liver disease, insulin resistance, osteopenia and myopathy, which are anticipated to intensify over long-duration spaceflight. Although these physiological responses can involve the microbiome, the host-microorganism interactions during spaceflight are still largely unknown. NASA GeneLab curates a wide range of space research data and the current work harnesses GeneLab multi’omic data from recent Rodent Research studies to explore changes to gut microbiota during spaceflight and their associations with host physiology when compared to ground controls. Using a hybrid analysis of DNA barcoding and whole genome shotgun data, an array of bacteria, fungi and nematodes could be identified at species level, and significant differences in relative abundances associated with spaceflight. Functional prediction based on differential abundance of species and metagenome gene inventories as well as metatranscriptomic gene expression at the host-gut microbiome interface implicate microbiota interactions could contribute to spaceflight pathology. Harnessing carefully curated publicly available data, such as from Genelab, to generate multi‘omic space science discoveries can help decipher the complex host-microbiome interactions that influence both health on Earth and the feasibility of long-duration spaceflight.

Microbiome↗

Overview of the Digitization Workflow Post Image Acquisition of Apollo Lunar and Antarctic Meteorite Samples Using Agisoft Photoscan for the NASA 3D Astromaterials Virtual Samples Collection

The 3D Virtual Astromaterials Samples (3DVAS) collection is a multi-year funded project to create a digital database of sixty Apollo Lunar and Antarctic Meteorite samples following non-destructive documentation conservation protocols. After initial image processing, the photos are evaluated and processed using unique structure-from-motion photogrammetric techniques in a high performance modelling software designed to create a 3D model from 2D images: Agisoft Photoscan Pro. Agisoft Photoscan Pro uses image processing algorithms and techniques originating in computer vision to resolve 3D models for accurate and detailed visualization of a subject. The software provides a stepwise process that is tailored per model based on spatial and specular reflectance properties, for example. The process includes: photo alignment, creation of a dense point cloud, mesh, and finally texture. Photo alignment is dependent on model properties. The 3DVAS process requires a special rotation platform with calibrated photogrammetric targets, specific distance rotation protocols, and a contrasting background for alignment and scale accuracy. As a result of the photographic process, alignment will complete with two mirrored hemispheres that, in a sense, represent the 2D images overlapping to create a 3D model. Each dense point cloud is analyzed with provided statistical measures in a gradual selection process to eliminate outliers. The point cloud is reduced to include only data valuable to the final model. When a precise dense point cloud is achieved, a mesh and texture are applied. Each model is scaled with scale bar accuracies within 100 microns. Each sample has its own intimate process for modelling; there is no standard for the parameters required in the final creation of a high resolution model. By processing multiple samples, a skill is gained in practice to allow a close definition of the original sample and will result in the most detailed version of the sample shell. This process completes one-fifth of the 3DVAS protocol for providing accurate digital documentation. Each model shell is merged with X-ray Computed Tomography data to create a full volumetric sample. All 3DVAS data will be served on NASA's Astromaterials Acquisition and Curation website with an early subset of data available in 2019 and the 3D Virtual Astromaterials Samples Collection launch in 2020.

Thomas, Andi B.↗

Carbon Storage Core Characterization Efforts at NETL

The multi-scale Computed Tomography (CT) and core flow facility in the Geocharacterization Laboratory at NETL, Morgantown yields porosity, permeability, and fracture properties of rock core samples obtained from the subsurface while maintaining the integrity of the sample. Additionally, geophysical bulk rock properties are analyzed with the laboratory’s GeoTEK multi-sensor core logger in a comparable fashion to downhole methods. NETL researchers collaborate with stakeholders within the carbon storage, oil and gas, and critical minerals sectors. Since 2017, over 1.88 miles of core have been analyzed within the laboratory and all data is publicly available through the Technical Report Series (TRS) on the Energy Data eXchange (EDX). Additionally, the website, RokBase, was curated to extrapolate and visualize the high-resolution data from field operations. The characterization of the Lively Grove #1 (LG#1) well provides a case study into the full capabilities of the Geocharacterization Laboratory. During the comprehensive study of LG#1 ~1-2 mm in diameter, vertical to bedding, cylindrical structures were identified throughout the St. Peter Formation. These structures are pervasive throughout the St. Peter Formation at depth and are characterized as the trace fossil, Skolithos.

Isom, Shelby L↗

Curating AI-Ready Datasets for Equity and Environmental Justice: A Data-Centric AI Case Study

An equitable and environmentally just community is essentialin order to avoid disproportionate burden borne by vulnerablecommunities. This need becomes pressing in the aftermathof an extreme event such as disaster or hazard when it is diffi-cult for the governing bodies to implement resource allocationas per the need. Artificial Intelligence (AI) algorithms canhelp surface Equity and Environmental Justice (EEJ) issueswhen trained on EEJ datasets. However, curating AI-readyEEJ training datasets is challenging due to differences in fac-tors such as heterogeneity, resolution, modality, and level ofexpertise in labeling. Additionally, EEJ issues involve sensi-tive information where uncertainties and errors could degradethe performance of AI algorithms. For eg. Error in seasonalcrop yield information can highly affect the prediction of an-nual crop yield. To address these challenges, Data-centricAI (DCAI) methods are employed, which enhance AI algo-rithm performance even with limited training samples. DCAIprioritizes data quality, thereby reducing the adverse effectsof uncertainties and errors during the model training process.This research proposes a novel dataset and benchmark for an-alyzing the effect of the Maui Wildfire of 2023 for Equityand Environmental Justice (EEJ) issues. The proposed datasetaligns with the concepts of DCAI such as annotation quality,data preprocessing, privacy, feature engineering, governanceand provenance. We firmly believe that the proposed datasetwould lay a foundation to implement robust and reliable mod-ern AI algorithms for addressing EEJ issues.

Paridhi Parajuli↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

NASA's GeneLab Phase II: Federated Search and Data Discovery

GeneLab is currently being developed by NASA to accelerate 'open science' biomedical research in support of the human exploration of space and the improvement of life on earth. Phase I of the four-phase GeneLab Data Systems (GLDS) project emphasized capabilities for submission, curation, search, and retrieval of genomics, transcriptomics and proteomics ('omics') data from biomedical research of space environments. The focus of development of the GLDS for Phase II has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

exobiology↗