Search NASA⌕ Search

SEARCH · Search NASA

Results for “metadata architecture”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

An Update on the CDDIS

The Crustal Dynamics Data Inforn1ation System (CoorS) supports data archiving and distribution activities for the space geodesy and geodynamics community. The main objectives of the system are to store space geodesy and geodynamics related data products in a central data bank, to maintain infom1ation about the archival of these data, and to disseminate these data and information in a timely mam1er to a global scientific research community. The archive consists of GNSS, laser ranging, VLBI, and OORIS data sets and products derived from these data. The coors is one of NASA's Earth Observing System Oata and Infom1ation System (EOSorS) distributed data centers; EOSOIS data centers serve a diverse user community and are tasked to provide facilities to search and access science data and products. The coors data system and its archive have become increasingly important to many national and international science communities, in pal1icular several of the operational services within the International Association of Geodesy (lAG) and its project the Global Geodetic Observing System (GGOS), including the International OORIS Service (IDS), the International GNSS Service (IGS), the International Laser Ranging Service (ILRS), the International VLBI Service for Geodesy and Astrometry (IVS), and the International Earth Rotation Service (IERS). The coors has recently expanded its archive to supp011 the IGS Multi-GNSS Experiment (MGEX). The archive now contains daily and hourly 3D-second and subhourly I-second data from an additional 35+ stations in RINEX V3 fOm1at. The coors will soon install an Ntrip broadcast relay to support the activities of the IGS Real-Time Pilot Project (RTPP) and the future Real-Time IGS Service. The coors has also developed a new web-based application to aid users in data discovery, both within the current community and beyond. To enable this data discovery application, the CDDIS is currently implementing modifications to the metadata extracted from incoming data and product files pushed to its archive. This poster will include background information about the system and its user communities, archive contents and updates, enhancements for data discovery, new system architecture, and future plans.

Noll, Carey↗

Managing Sustainable Data Infrastructures: The Gestalt of EOSDIS

EOSDIS epitomizes a System of Systems, whose many varied and distributed parts are integrated into a single, highly functional organized science data system. A distributed architecture was adopted to ensure discipline-specific support for the science data, while also leveraging standards and establishing policies and tools to enable interdisciplinary research, and analysis across multiple scientific instruments. The EOSDIS is composed of system elements such as geographically distributed archive centers used to manage the stewardship of data. The infrastructure consists of underlying capabilities connections that enable the primary system elements to function together. For example, one key infrastructure component is the common metadata repository, which enables discovery of all data within the EOSDIS system. EOSDIS employs processes and standards to ensure partners can work together effectively, and provide coherent services to users.

remote sensing↗

ATAT: Astronomical Transformer for time series and Tabular data

Context. The advent of next-generation survey instruments, such as theVera C. RubinObservatory and its Legacy Survey of Space and Time (LSST), is opening a window for new research in time-domain astronomy. The Extended LSST Astronomical Time-Series Classification Challenge (ELAsTiCC) was created to test the capacity of brokers to deal with a simulated LSST stream. Aims. Our aim is to develop a next-generation model for the classification of variable astronomical objects. We describe ATAT, the Astronomical Transformer for time series And Tabular data, a classification model conceived by the ALeRCE alert broker to classify light curves from next-generation alert streams. ATAT was tested in production during the first round of the ELAsTiCC campaigns. Methods. ATAT consists of two transformer models that encode light curves and features using novel time modulation and quantile feature tokenizer mechanisms, respectively. ATAT was trained on different combinations of light curves, metadata, and features calculated over the light curves. We compare ATAT against the current ALeRCE classifier, a balanced hierarchical random forest (BHRF) trained on human-engineered features derived from light curves and metadata. Results. When trained on light curves and metadata, ATAT achieves a macro F1 score of 82.9 ± 0.4 in 20 classes, outperforming the BHRF model trained on 429 features, which achieves a macro F1 score of 79.4 ± 0.1. Conclusions. The use of transformer multimodal architectures, combining light curves and tabular data, opens new possibilities for classifying alerts from a new generation of large etendue telescopes, such as theVera C. RubinObservatory, in real-world brokering scenarios.

Astronomy & Astrophysics↗

GGOS Bureau of Networks and Observations: Network Infrastructure and Related Activities

The GGOS Bureau of Networks and Observations works with the IAG Services (IVS, ILRS, IGS, IDS, IGFS, and PSMSL) to advocate for the expansion and upgrade of space geodesy networks for the maintenance and improvement of the reference frame and other applications, as well as for the integration with other techniques, including absolute gravity and sea level measurements from tide gauges. New sites are being established following the GGOS concept of “core” and co-location sites, and new technologies are being implemented to enhance performance in data yield as well as accuracy. The Bureau continues to meet with organizations to discuss possibilities, including partnerships, for new and expanded participation. The GGOS Network continues to grow as new stations join every year. The Bureau holds meetings frequently, providing the opportunity for representatives from the services to meet and share progress and plans, and to discuss issues of common interest. It also monitors the status and projects the evolution of the network based on information from the current and expected future participants. Of particular interest at the moment is the integration of gravity and tide gauge networks and the forthcoming establishment of the new absolute gravity reference frame. The IAG Committees and Joint Working Groups play an essential role in the Bureau activity. The Standing Committee on Performance Simulations and Architectural Trade-offs (PLATO) uses simulation and analysis techniques to project future network capability and to examine trade-off options. The Committee on Data and Information is working on a strategy for a GGOS metadata system for data products and a more comprehensive long-term plan for an all-inclusive system. The Committee on Satellite Missions is working to enhance communication with the space missions, to advocate for missions that support GGOS goals and to enhance ground systems support. The IERS Working Group on Site Survey and Co-location (also participating in the Bureau) is working to enhance standardization in procedures, outreach and to encourage new survey groups to participate and improve procedures to determine systems’ reference points, a crucial aid in the detection of technique-specific systematic errors. We will give a brief update on the status and projection of the network infrastructure of the next several years, and the progress and plans of the Committees/Working Group in their critical role in enhancing data product quality and accessibility to the users.

Carey Noll↗

Earth Science Datacasting v2.0

The Datacasting software, which consists of a server and a client, has been developed as part of the Earth Science (ES) Datacasting project. The goal of ES Datacasting is to provide scientists the ability to automatically and continuously download Earth science data that meets a precise, predefined need, and then to instantaneously visualize it on a local computer. This is achieved by applying the concept of podcasting to deliver science data over the Internet using RSS (Really Simple Syndication) XML feeds. By extending the RSS specification, scientists can filter a feed and only download the files that are required for a particular application (for example, only files that contain information about a particular event, such as a hurricane or flood). The extension also provides the ability for the client to understand the format of the data and visualize the information locally. The server part enables a data provider to create and serve basic Datacasting (RSS-based) feeds. The user can subscribe to any number of feeds, view the information related to each item contained within a feed (including browse pre-made images), manually download files associated with items, and place these files in a local store. The client-server architecture enables users to: a) Subscribe and interpret multiple Datacasting feeds (same look and feel as a typical mail client), b) Maintain a list of all items within each feed, c) Enable filtering on the lists based on different metadata attributes contained within the feed (list will reference only data files of interest), d) Visualize the reference data and associated metadata, e) Download files referenced within the list, and f) Automatically download files as new items become available.

Bingham, Andrew W.↗

Intelligent Systems Technologies and Utilization of Earth Observation Data

The addition of raw data and derived geophysical parameters from several Earth observing satellites over the last decade to the data held by NASA data centers has created a data rich environment for the Earth science research and applications communities. The data products are being distributed to a large and diverse community of users. Due to advances in computational hardware, networks and communications, information management and software technologies, significant progress has been made in the last decade in archiving and providing data to users. However, to realize the full potential of the growing data archives, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications. Sponsored by NASA s Intelligent Systems Project within the Computing, Information and Communication Technology (CICT) Program, a conceptual architecture study has been conducted to examine ideas to improve data utilization through the addition of intelligence into the archives in the context of an overall knowledge building system (KBS). Potential Intelligent Archive concepts include: 1) Mining archived data holdings to improve metadata to facilitate data access and usability; 2) Building intelligence about transformations on data, information, knowledge, and accompanying services; 3) Recognizing the value of results, indexing and formatting them for easy access; 4) Interacting as a cooperative node in a web of distributed systems to perform knowledge building; and 5) Being aware of other nodes in the KBS, participating in open systems interfaces and protocols for virtualization, and achieving collaborative interoperability.

Ramapriyan, H. K.↗

JPL Project Information Management: A Continuum Back to the Future

This slide presentation reviews the practices and architecture that support information management at JPL. This practice has allowed concurrent use and reuse of information by primary and secondary users. The use of this practice is illustrated in the evolution of the Mars Rovers from the Mars Pathfinder to the development of the Mars Science Laboratory. The recognition of the importance of information management during all phases of a project life cycle has resulted in the design of an information system that includes metadata, has reduced the risk of information loss through the use of an in-process appraisal, shaping of project's appreciation for capturing and managing the information on one project for re-use by future projects as a natural outgrowth of the process. This process has also assisted in connection of geographically disbursed partners into a team through sharing information, common tools and collaboration.

Information Management↗

Integrating Ideas for International Data Collaborations Through The Committee on Earth Observation Satellites (CEOS) International Directory Network (IDN)

The capabilities of the International Directory Network's (IDN) version MD9.5, along with a new version of the metadata authoring tool, "docBUILDER", will be presented during the Technology and Services Subgroup session of the Working Group on Information Systems and Services (WGISS). Feedback provided through the international community has proven instrumental in positively influencing the direction of the IDN s development. The international community was instrumental in encouraging support for using the IS0 international character set that is now available through the directory. Supporting metadata descriptions in additional languages encourages extended use of the IDN. Temporal and spatial attributes often prove pivotal in the search for data. Prior to the new software release, the IDN s geospatial and temporal searches suffered from browser incompatibilities and often resulted in unreliable performance for users attempting to initiate a spatial search using a map based on aging Java applet technology. The IDN now offers an integrated Google map and date search that replaces that technology. In addition, one of the most defining characteristics in the search for data relates to the temporal and spatial resolution of the data. The ability to refine the search for data sets meeting defined resolution requirements is now possible. Data set authors are encouraged to indicate the precise resolution values for their data sets and subsequently bin these into one of the pre-selected resolution ranges. New metadata authoring tools have been well received. In response to requests for a standalone metadata authoring tool, a new shareable software package called "docBUILDER solo" will soon be released to the public. This tool permits researchers to document their data during experiments and observational periods in the field. interoperability has been enhanced through the use of the Open Archives Initiative s (OAI) Protocol for Metadata Harvesting (PMH). Harvesting of XML content through OAI-MPH has been successfully tested with several organizations. The protocol appears to be a prime candidate for sharing metadata throughout the international community. Data services for visualizing and analyzing data have become valuable assets in facilitating the use of data. Data providers are offering many of their data-related services through the directory. The IDN plans to develop a service-based architecture to further promote the use of web services. During the IDN Task Team session, ideas for further enhancements will be discussed.

Olsen, Lola M.↗

Intelligent Systems Technologies to Assist in Utilization of Earth Observation Data

With the launch of several Earth observing satellites over the last decade, we are now in a data rich environment. From NASA's Earth Observing System (EOS) satellites alone, we are accumulating more than 3 TB per day of raw data and derived geophysical parameters. The data products are being distributed to a large user community comprising scientific researchers, educators and operational government agencies. Notable progress has been made in the last decade in facilitating access to data. However, to realize the full potential of the growing archives of valuable scientific data, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications. Sponsored by NASA s Intelligent Systems Project within the Computing, Information and Communication Technology (CICT) Program, a conceptual architecture study has been conducted to examine ideas to improve data utilization through the addition of intelligence into the archives in the context of an overall knowledge building system. Potential Intelligent Archive concepts include: 1) Mining archived data holdings using Intelligent Data Understanding algorithms to improve metadata to facilitate data access and usability; 2) Building intelligence about transformations on data, information, knowledge, and accompanying services involved in a scientific enterprise; 3) Recognizing the value of results, indexing and formatting them for easy access, and delivering them to concerned individuals; 4) Interacting as a cooperative node in a web of distributed systems to perform knowledge building (i.e., the transformations from data to information to knowledge) instead of just data pipelining; and 5) Being aware of other nodes in the knowledge building system, participating in open systems interfaces and protocols for virtualization, and collaborative interoperability. This paper presents some of these concepts and identifies issues to be addressed by research in future intelligent systems technology.

Ramapriyan, Hampapuram K.↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences↗

Using Open and Interoperable Ways to Publish and Access LANCE AIRS Near-Real Time Data

The Atmospheric Infrared Sounder (AIRS) Near-Real Time (NRT) data from the Land Atmosphere Near real-time Capability for EOS (LANCE) element at the Goddard Earth Sciences Data and Information Services Center (GES DISC) provides information on the global and regional atmospheric state, with very low temporal latency, to support climate research and improve weather forecasting. An open and interoperable platform is useful to facilitate access to, and integration of, LANCE AIRS NRT data. As Web services technology has matured in recent years, a new scalable Service-Oriented Architecture (SOA) is emerging as the basic platform for distributed computing and large networks of interoperable applications. Following the provide-register-discover-consume SOA paradigm, this presentation discusses how to use open-source geospatial software components to build Web services for publishing and accessing AIRS NRT data, explore the metadata relevant to registering and discovering data and services in the catalogue systems, and implement a Web portal to facilitate users' consumption of the data and services.

Zhao, Peisheng↗

FAIR Ecosystems for Science at Scale

High Performance Computing (HPC) centers provide resources to users who require greater scale to “get science done”. They deploy infrastructure with singular hardware architectures, cutting-edge software environments, and stricter security measures as compared with users’ own resources. As a result, users often create and configure digital artifacts in ways that are specialized for the unique infrastructure at a given HPC center. Each user of that center will face similar challenges as they develop specialized solutions to take full advantages of the center’s resources, potentially resulting in significant duplication of effort. Much duplicated effort could be avoided, however, if users of these centers found it easier to discover others’ solutions and artifacts as well as share their own. The FAIR principles address this problem by presenting guidelines focused around metadata practices to be implemented by vaguely defined “communities”; in practice, these tend to gather by domain (e.g. bioinformatics, geosciences, agriculture). Domain-based communities can unfortunately end up functioning as silos that tend both to inhibit sharing of solutions and best practices as well as to encourage fragile and unsustainable improvised solutions in the absence of best-practice guidance. We propose that these communities pursuing “science at scale” be nurtured both individually and collectively by HPC centers so that users can take advantage of shared challenges across disciplines and potentially across HPC centers. We describe an architecture based on the EOSC-Life FAIR Workflows Collaboratory, specialized for use with and inside HPC centers such as the Oak Ridge Leadership Computing Facility (OLCF), and we speculate on user incentives to encourage adoption. We note that a focus on FAIR workflow components rather than FAIR workflows is more likely to benefit the users of HPC centers.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

Digital Archive Issues from the Perspective of an Earth Science Data Producer

Contents include the following: Introduction. A Producer Perspective on Earth Science Data. Data Producers as Members of a Scientific Community. Some Unique Characteristics of Scientific Data. Spatial and Temporal Sampling for Earth (or Space) Science Data. The Influence of the Data Production System Architecture. The Spatial and Temporal Structures Underlying Earth Science Data. Earth Science Data File (or Relation) Schemas. Data Producer Configuration Management Complexities. The Topology of Earth Science Data Inventories. Some Thoughts on the User Perspective. Science Data User Communities. Spatial and Temporal Structure Needs of Different Users. User Spatial Objects. Data Search Services. Inventory Search. Parameter (Keyword) Search. Metadata Searches. Documentation Search. Secondary Index Search. Print Technology and Hypertext. Inter-Data Collection Configuration Management Issues. An Archive View. Producer Data Ingest and Production. User Data Searching and Distribution. Subsetting and Supersetting. Semantic Requirements for Data Interchange. Tentative Conclusions. An Object Oriented View of Archive Information Evolution. Scientific Data Archival Issues. A Perspective on the Future of Digital Archives for Scientific Data. References Index for this paper.

Barkstrom, Bruce R.↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

infrastore [SWR-26-077]

Infrastore is time-series storage for energy-systems simulations, backed by HDF5 + SQLite, with Rust, Python, Julia, gRPC, and CLI bindings. It is a Rust library for managing time-series data in power-systems and energy simulations. Numerical arrays are persisted in HDF5, and the metadata associating each array with its owning component lives in SQLite. Identical arrays are stored once and shared through content addressing. It ships native Rust, Python (PyO3), and Julia (C ABI) interfaces, the infrastore command-line tool, and a read-only gRPC server with a Rust client. Documentation: https://natlabrockies.github.io/infrastore/latest/ — start with the Quick Start or the Architecture.

Thom, Daniel [National Laboratory of the Rockies (↗

Normalizing Resource Identifiers using Lexicons in the Global Change Information System: Linking Earth Science Identifiers, Concepts, and Communities

Earth Science informatics involves collaboration between multiple groups of people with diverse specializations and goals,often using variations in terminology to refer to common resources. The uniformity of the resource identifiers often does not cross organizational boundaries. Because of this, permanent, widely used, unambiguous identifiers for resources are elusive. We examine real world cases of changing and inconsistent identifiers which inherently work against persistence and uniformity. We also present a solution which mediates factors in these situations; namely the creation of lexicons:mappings of sets of terms to URIs which are curated within the Global Change Information System (GCIS). We discuss aspects of the GCIS which facilitate the use of lexicons: an information model which disambiguates resources, a RESTful API which provides metadata through content-negotiation, and a strategy for long term curation of URIs, including mechanisms for handling changes to URIs and variations in terms used by different communities while providing persistent URIs and preserving relationships between resources We provide working definitions of terms,contexts, and lexicons, and relate them to the practical challenges of disambiguation and curation. We also discuss the mechanisms employed and architecture of the GCIS, and how these choices facilitate representation of persistent identifiers and mappings of them to identifiers used colloquially within various earth science communities of practice.

Linkded Data↗

Biomass Harmonization and SAR Analysis with the Multi-mission Algorithm and Analysis Platform (MAAP)

The Multi‐mission Algorithm and Analysis Platform (MAAP) is a collaborative effort between NASA and the European Space Agency (ESA) to support above ground biomass (AGB) research in an open science framework. MAAP brings together relevant data, algorithms, and computing capabilities in a common cloud environment to address the challenges of sharing and processing data from field, airborne and satellite measurements. MAAP was publicly released in October 2021, providing computing capabilities co-located with the data, a collaborative coding and analysis environment, and a set of interoperable tools and algorithms developed to support the estimation and visualization of data. MAAP has allowed scientists from both North America and Europe to collaborate on the generation and analysis/visualization of data derived from multiple, discipline-adjacent missions in an open, collaborative environment that has reached beyond traditional scientific investigation. MAAP has been used to support multiple scientific activities. To date, existing LiDAR data from multiple platforms has been calibrated with field measurements and combined for more comprehensive and accurate estimates of above ground biomass AGB; these LiDAR platforms include airborne (e.g. LVIS), the International Space Station (NASA’s Global Ecosystem Dynamics Investigation (GEDI), and satellites (e.g. ICESat-2). The current challenge is to effectively and seamlessly combine the aforementioned LiDAR-based data with new data sources such as P-band RADAR from ESA’s upcoming BIOMASS mission, existing ESA Sentinel-1 C-band SAR, and the 30 PB/yr of high cadence global coverage L-band SAR data from the upcoming NASA-ISRO SAR (NISAR) mission. Recent analysis using MAAP merged ICESat-2 and optical data (Harmonized Landsat Sentinel) produced the most comprehensively precise estimate of boreal-wide AGB to date. Another effort using MAAP is the production and open distribution of global comparisons of AGB map estimates, including from ICESat-2 and GEDI, to bolster stakeholder uptake for policy applications. These map estimates will feed into the Intergovernmental Panel on Climate Change (IPCC) database, likely aiding the next Global Carbon Stocktake of the UNFCCC. Furthermore, the biomass retrieval intercomparison exercise BRIX-2 could benefit from the MAAP providing standardized test cases (based on airborne campaign and spaceborne data) allowing the community to develop and apply retrieval algorithms based on these test cases, while forthcoming SAR data training curricula could also use the MAAP as a teaching and learning platform. The MAAP is meeting the challenges inherent in international, open science collaboration and large scale computing with a platform that is entirely open source and cloud native, using open standards for data access, manipulation, protocols, and formats. The MAAP data system consists of a dedicated data store whose data is indexed in an online catalog conforming to established metadata, application programmatic interfaces (APIs), and service interface standards, using an implementation of the open sourced NASA Common Metadata Repository. Federation of user identities allows users from either NASA or ESA to access and consume services from the other using a unified metadata catalog for the data utilized across the ESA and NASA MAAP platforms. Similarly, we are exploring how to increase interoperability to achieve a common approach to packaging, orchestrating and executing algorithms, with interoperable access to data for subsetting, fast browse, and cloud-optimized access, all using interoperable standards such as those from the Open Geospatial Consortium (OGC). Designed for interoperability, ESA and NASA utilize a common architecture for the software platform. It provides a cloud-based algorithm development environment (ADE) that enables scientists to develop algorithms collaboratively with access to the MAAP data catalog as well as other data archives. MAAP provides an Eclipse Che-based ADE supporting both Python and R languages, popular in this biomass community. Algorithms developed and containerized within the ADE can be deployed to run to thousands of computational nodes in the MAAP’s data processing system (DPS), dramatically speeding up processing and giving scientists a rapid, iterative turnaround of results. NASA’s implementation of the DPS is based on the Hybrid Science Data System (HySDS) framework, used by NASA flight projects to produce Earth science standard products.

cloud computing↗