Search NASASearch

SEARCH · Search NASA

Results for “open data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]

Open-Source Data for MAC-POSTS: Mobility Data Analytics Center - Prediction, Optimization, and Simulation Toolkit for Transportation Systems

MAC-POSTS (Mobility Data Analytics Center - Prediction, Optimization, and Simulation toolkit for Transportation Systems) is a toolkit for dynamic transportation network modeling. Developed by the Mobility Data Analytics Center (MAC) at Carnegie Mellon University, this package implements many classic dynamic transportation network models, as well as new models proposed by MAC members. It has served as one building block for many other models and research projects. As such, this package used to be treated as an internal research project of the MAC lab, and admittedly, the code base is messy, and the interface is hard to use. However, we are working hard to make it a generally usable and useful toolkit for dynamic transportation network modeling. We would really appreciate any feedback, comments, suggestions, or criticisms.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

TPSAS-NF1676L-32493-DND

The Committee on Earth Observation Satellites (CEOS) System Engineering Office (SEO) has supported the Open Data Cube (ODC) initiative to provide a data architecture solution that has value to its global users and increases the impact of EO satellite data. ODC is an open-source platform for processing satellite data. We have developed software products and tools around the core ODC that would help users perform machine learning on EO satellite data. The recent United Nations (UN) Sustainable Development Agenda provides a shared blueprint for peace and prosperity for people and for the planet, considering our current situation and helping to create a plan. The core of this agenda is a set of seventeen Sustainable Development Goals (SDGs), which represent an urgent call for action by all countries - both developed and developing - in a global partnership. The CEOS SEO team has recently developed and released a set of innovative Jupyter notebooks addressing UN SDGs 6.6.1 (spatial extents of water-related ecosystems), 11.3.1 (ratio of land consumption rate to population growth rate), and 15.3.1 (proportion of land that is degraded over total land area). These notebooks empower users by providing features that will assist with streamlining analysis ready data retrieval, processing, and visualization. We have recently incorporated several machine learning techniques in these notebooks. In this paper, we present the lessons learned from our experience on classifying land using supervised and unsupervised machine learning techniques using ODC framework for UN SDGs. We identify the current limitations of ODC to seamlessly support machine learning techniques. We propose features that would help machine learning, specifically within the ODC framework. We propose a thematic indexing/loading of data for both unsupervised learning as well as data annotation/labeling pipeline. Currently, ODC supports machine learning by separating data-management from the analysis process. It works as a mechanism to load cubes of data. ODC does not natively support features that are vital in machine learning such as validation splits, fair/balanced sampling, establishing load size constraints, etc. We believe that our proposed features will empower users by providing features that bring machine learning techniques closed to ODC. Enhancements to ODC to better accommodate machine learning techniques can assist in fulfilling UN SDGs such as 6.3.2, 6.4.2, 6.6.1, 11.3.1, 14.1.1, 15.1.1, 15.3.1, and 15.4.2.

Syed R Rizvi

Advancing Open Science Through Innovative Data System Solutions: The Joint ESA-NASA Multi-Mission Algorithm and Analysis Platform (MAAP)'s Data Ecosystem

Collaborative open science practices are changing the way research is conducted. These changes affect how scientists work together on data, code and information. Data systems enhance open science by offering forward thinking technological solutions, such as providing data and computation on the cloud, to enable collaboration, sharing and analysis. In this paper, we present our vision for a conceptual data system on the cloud that enables open science. We also present our work on the Multi-Mission Algorithm and Analysis Platform (MAAP)which has served as a pathfinder data system for this conceptual approach.

Kaylin Bugbee

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics

TRUST, Trustworthiness and EOSDIS

In recent years there has been considerable attention by the international scientific research and applications community to ensure high quality of data and information management. The terms FAIR (Findable, Accessible, Interoperable, Reusable) data, TRUST (Transparency, Responsibility, User Community, Sustainability, and Technology) principles, and CARE (Collective Benefit, Authority to Control, Responsibility, and Ethics) principles have come into vogue during the last decade. NASA has been managing data and information for over 60 years. NASA’s Earth Observing System Data and Information System (EOSDIS) has been in operation for over 25 years, managing most of NASA’s Earth science data. Trustworthiness is a goal that NASA has always strived to achieve or exceed, because it: enables the success of any NASA science mission; inspires general science research and applications; justifies the cost of operations; contributes to the value of NASA’s Open Data Policy; and influences the long term, historical view for the data collection. Given the recent growth of interest in TRUST principles, it is useful to assess and show how NASA’s attention to trustworthiness maps into those principles. This presentation addresses shows how the various steps that have been taken by the Earth Science Data and Information System (ESDIS) Project in the implementation and evolution of EOSDIS map into the TRUST principles.

Remote Sensing

Rhode Island Ecological Conservation: Methods for Monitoring Rhode Island Habitats: Contributing to a Framework for Targeted Conservation and Management

Global avian population decline since the 1970s is largely attributable to habitat loss and degradation from anthropogenic disturbances. NASA DEVELOP’s Rhode Island Ecological Conservation team partnered with the Audubon Society of Rhode Island to compute land use land cover (LULC) maps of Rhode Island to aid in the conservation of the state’s 140 bird species. This project aimed to support the partner’s land acquisition strategies with updated and specific LULC classifications showing potential bird-habitat locations across the state. We incorporated remotely sensed data from Landsat 8 and 9 Operational Land Imager (OLI) into LULC maps using unsupervised classification techniques in ArcGIS Pro and supervised classification in Google Earth Engine. We generated six land classifications for 2023, which showed land cover dominated by upland habitats (forests, scrub/shrub, and grasslands), followed by development. We used TerrSet’s Land Change Modeler to forecast LULC change through 2043, using 2011 and 2021 National Land Cover Database (NLCD) land cover maps derived from Landsat 8 and 9 imagery. Project results suggest that non-urban upland and wetland habitats will decrease over time, while development will continue to encroach on non-urban avian habitats. Our maps and associated data will allow for more efficient land acquisition and management efforts to support avian habitat conservation across Rhode Island. Our study shows that data acquisition and processing from open data sources is feasible and further analysis can be done through GIS classification tools. More analysis is needed beyond this study to obtain more detailed land cover maps, though Audubon can aid its targeted conservation efforts with our current, historic, and forecasted LULC maps.

Remote sensing

High School Citizen Scientists Use AI/ML to Predict Intra-Ocular Pressure From Gene Expression Data for Spaceflown Mice

Artificial Intelligence (AI) and Machine Learning (ML) have increasingly become pivotal in biological and biomedical research, largely due to the culture of open data sharing and its associated benefits. The methodologies inherent in AI/ML are particularly adept at identifying and forecasting biological phenotypes from the vast amounts of data generated by next-generation sequencing technologies. These techniques offer substantial promise for advancing research in space biosciences and for the development of automated systems for monitoring space health. Nevertheless, there are crucial aspects to consider when training, validating, and testing machine learning models in both biological research and clinical contexts. It is essential that Open Science principles, including data sharing and the availability of open-source code, are complemented by high-quality, publicly accessible training resources. These resources should focus on best practices and include modules based on real-world scientific cases and data to ensure that future AI/ML practitioners gain practical experience with genuine problems. Addressing this knowledge gap, we have designed, developed, and delivered both interactive and self-paced training programs for citizen scientists worldwide, enabling them to utilize AI/ML for space biology research. This initiative was made possible through generous funding from a Transformation to Open Science Training grant. The interactive training sessions, conducted this summer, utilized AI/ML techniques to analyze data from the Open Science Data Repository, specifically targeting the effects of spaceflight on ocular structure and function. The dataset OSD-583, from the Rodent Research 9 mission, provides experimental data detailing the ocular responses of mice subjected to a 35-day spaceflight, compared with ground control counterparts. Using OSD-583 as observational data, our summer training participants applied AI/ML methods to predict intraocular pressure from RNA-seq data and identify the genes most predictive of the observed responses. Further analysis through pathway enrichment and gene set enrichment revealed that these genes are involved in molecular and cellular processes contributing to retinal degeneration.

James Casaletto

Circularity Futures Workshop Series: Summary Report

The aim of this report is to synthesize key feedback received from the three-part Circularity Futures workshop series held in Spring 2024. The workshop series was conducted by the National Renewable Energy Laboratory (NREL) on behalf of U.S. Department of Energy, Office Energy Efficiency and Renewable Energy (EERE), and was broken into three workshops: Workshop 1 - Circularity Analysis Needs and Priorities; Workshop 2 - Circularity Metrics and Indicators; and Workshop 3 - Circularity Data. Together, the workshops focused on identifying the existing priorities and gaps in the circularity modeling space, understanding different stakeholders' use and interpretation of circularity metrics and indicators, identifying common data gaps and data quality challenges, and assessing the robustness of available solutions. The workshop series brought a diverse group of stakeholders - including representatives from U.S. government offices, national labs, nonprofit organizations, industry, and academia - to collect first-hand feedback on needs, priorities, challenges and opportunities in the circularity modeling and analysis space. The workshop discussions highlighted numerous common needs, priorities and challenges among the interviewed groups. Several topics were frequently discussed, including: 1) Circularity as a pathway for sustainable economic growth: While circularity is generally defined in terms of resource conservation and reducing wasteful disposal of materials, participants agreed that circular strategies should serve broader economic, environmental, and social goals. It is therefore crucial for circularity analysis to look beyond waste reduction and instead evaluate a variety of impact metrics such as cost savings, job creation, air quality, and pollutant emissions. Mutli-criteria decision-making frameworks may be useful for making sense of disparate metrics and evaluating tradeoffs between impact categories.; 2) Economic and social factors are not well understood: Underdevelopment of existing end-of-life (EOL) management infrastructure, inconsistent standardization codes and policy space in reusing recycled content, and suboptimal collection and sorting strategies collectively contribute to uncertainty about the economic potential of circular pathways. The latter observation is consistent among all technologies but more emphasized for renewable energy systems. Social impacts of circularity practices are less understood and less researched than other sustainability aspects.; 3) Inconsistent methods for assessing emerging technologies: LCA and TEA results vary widely depending on the assumptions made with regards to market adoption of new technologies. Emerging technologies suffer limited availability of data needed to conduct a robust circularity analysis. Yet, understanding projected impacts of proposed nascent technology is a key need for different stakeholder groups.; and 4) Lack of temporally and geospatially explicit data: There is a need for open data that represents variations in circularity technologies over time and location. The lack thereof leads to aggregated and potentially misrepresented results in circularity analysis. Sensitivity analyses should be included to verify whether options perceived as more sustainable align with real-world practices.

29 ENERGY PLANNING, POLICY, AND ECONOMY

NASA's Earth Science Data Systems: A "Bit of History" and Observations

NASA has significantly improved its Earth Science Data Systems over the last two decades. Open data policy and inexpensive (or free) availability of data has promoted data usage by broad research and applications communities. Flexibility, accommodation of diversity, evolvability, responsiveness to community feedback are key to success.

Ramapriyan, H. K.