Search NASA⌕ Search

Engineering topics

Sandeep D Shetye

Publications and source records attributed to Sandeep D Shetye.

DRF: A Software Architecture for a Data Marketplace to Support Advanced Air Mobility

Advanced Air Mobility is a new aviation vision, where unmanned aerial systems will trans- port passengers and cargo across urban and rural areas. Critical to the realization of this vision is the development of a digital marketplace, which allows service providers and consumers operating in the airspace ecosystem to securely exchange data and reasoning insights. In this paper, we present the architecture of a decentralized data marketplace that connects data and reasoning service providers to vehicles and other service consumers along the cloud-to-edge continuum. We also present two example use cases to demonstrate the value of our approach.

Autonomy↗

LSKnowledge: Nexus for Transformative Scientific Discoveries and Enhanced Information Retrieval in NASA Life Sciences Portal

We stand at the brink of an extraordinary transformation in the field of AI, driven by the convergence of generative AI and semantic technologies (e.g., knowledge graphs). This fusion holds immense potential and could redefine the future of scientific exploration, particularly in the realm of life sciences research. In this context, we shed light on the pivotal roles that Large Language Models (LLMs) and semantic technologies will play in advancing research, unearthing and comprehending life sciences information through innovative approaches, and empowering researchers to extract insights from NASA's extensive Life Sciences Data Archive. Within the NASA Life Sciences Portal (NLSP), the integration of LLMs and semantic technologies unlocks several advanced capabilities. First and foremost, it equips scientists with sophisticated tools to manage the ever-expanding wealth of scientific literature and data. Furthermore, it facilitates the creation of knowledge graphs that visually represent intricate relationships among biological entities, enabling comprehensive systems-level analysis. Additionally, the fusion of generative AI (including LLMs) and semantic technology can significantly benefit NASA's life sciences research by enhancing information retrieval and hypothesis generation. These tools enhance natural language understanding, facilitating knowledge discovery within NLSP. The overarching vision is to establish a cohesive knowledge ecosystem within NLSP, harnessing the power of LLMs and semantic technologies to synthesize and cross-reference data from diverse missions, disciplines, and research domains. This holistic approach ultimately deepens our understanding of how space environments impact life sciences data. To advance this initiative, we have launched LSKnowledge, aimed at enhancing the information retrieval capabilities of NLSP. In the short term, our primary goal is to develop a robust semantic search system. This system will empower HRP (Human Research Program) researchers to navigate NLSP data repositories more efficiently and precisely, catalyzing the process of hypothesis formation and scientific breakthroughs. To achieve this, we have employed pre-trained LLMs as part of a semantic search tool that can rank and highlight the most relevant records for user queries. To assess the tool's performance, we have curated a set of approximately 200 queries from subject matter experts (SMEs) and manually ranked the top records retrieved by both the current search system and the new semantic search, using SME judgments as the gold standard for relevancy. Herein, we present the results of our comparative analysis and illustrate how these findings have informed the fine-tuning of the system for enhanced performance. In the long term, our objectives include 1) retrieving publicly available information and integrating it with NLSP data to provide more precise answers to user queries, and 2) incorporating non-textual information from the NLSP database into our approach. In conclusion, the fusion of LLMs and semantic technologies within NLSP represents a pioneering stride towards reshaping the landscape of scientific discovery. This synergy not only equips researchers with powerful tools to navigate the burgeoning sea of information but also facilitates a deeper understanding of complex biological relationships, all while accelerating hypothesis generation and knowledge discovery. Through our initiative, LSKnowledge, we are committed to continually refining and expanding these capabilities, with the aim of not only enhancing information retrieval but also integrating diverse data sources to provide more precise insights. In the grand vision, NLSP strives to become the cornerstone of a comprehensive knowledge ecosystem, unraveling the enigmatic intricacies of life sciences phenomena in the context of space environments.

Life Sciences↗

Transformation of the NASA Life Sciences Portal to a FAIR Data Point

The FAIR principles emphasize optimizing metadata, the vast majority of which are textual in nature, and often organized into attribute name-value pairs. This uniformity has led to the development of guidelines and best practices for providing programmatic access to scientific data through their metadata, yielding the first iteration of the FAIR Data Point Specifications (FDPS). A key feature of the FDPS is its support for automated agents seeking and fetching data without first needing to learn a plethora of different application programming interfaces. These software agents can interrogate metadata catalogs that adhere to FDPS in a uniform manner because each catalog describes itself and its metadata schema consistently. This approach enhances the sustainability of data retrieval support, allowing systems to refine and update their metadata schemas as needed and without requiring data-seeking software agents to change how they interrogate FDPS catalogs. An essential aspect of the FDPS is the standardization of data catalog semantics, which formalizes concepts such as “metadata” and “metadata service” and links them to other concepts specifications including the Data Catalog Vocabulary (DCAT), a W3C standard that is also the basis of NASA-STD-2831 “Metadata Standard for Data Discoverability,” authored by NASA’s Office of the Chief Information Officer. The FDPS references DCAT (version 2) elements which focus on the distribution of datasets and support the goal of stream-lined catalog integration across repositories for improved data discovery. Additionally, the FDPS also prescribe the use of Linked Data Platform elements for data catalog-metadata record containment descriptions, allowing users to ascertain which data and metadata belong to which catalogs. NASA’s Life Sciences Portal is implementing the FDPS while formalizing its metadata schema to support the accelerated synthesis of knowledge from space life sciences investigations.

platform↗

Governing Data Findability, Accessibility, Interoperability and Reusability (FAIR) Compliance

The most recent data strategy documents at both the federal and NASA levels stipulate that systems should strive for the data they manage to be Findable, Accessible, Interoperable, and Reusable (FAIR). The NASA Life Sciences Portal (NLSP) has already begun leading efforts in this area for HRP, initiating efforts to comply with the FAIR principles. The broad interpretation of the FAIR principles has led to a plethora of tools that use a splay of metrics specifically but variably developed to judge how compliant data and systems are with the principles. A recent review [3] identified and studied 1,180 metrics across 20 publicly available tools for checking FAIR compliance of data and systems. Because of their very recent development, many organizations and data systems managers and developers have not yet had adequate time or resources to understand these FAIR compliance tools and metrics, their variations in design, accuracy or ease of application to their specific data sets and systems. Thus, it would be best for larger organizations like NASA to approach formulating a strategy for governance of FAIR compliance that can be flexibly applied and is adaptable to an evolving awareness knowledge of FAIR compliance methods and tools. In September 2024, the NASA Science Mission Directorate(SMD) organized a workshop on NASA science data repositories, including the topics of implementing FAIR and governing FAIR compliance across SMD. The initial part of these FAIR discussions focused on developing consensus around required science metadata fields. This is challenging given the diverse nature of NASA’s scientific data portfolio, the variety of metadata models and vocabularies used, and variable level of resources available to curate these data. Later discussion focused on three possible approaches to governing FAIR compliance: distributed, in which various programs, projects or systems define their own methods for assessing FAIR compliance, reporting results up appropriate management lines; centralized, in which higher-level organization(s) specify compliance tools or methods for the various data systems; and multi-level, in which a group comprised of individuals with expertise from multiple levels with organizations is formed to provide guidance and/or specifications for governing FAIR compliance. We report on the recommendations this session yielded, and how these might be shaped specifically to help implement and govern the compliance with FAIR of Human Research Program data and systems.

governance↗