Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Experimental Data on Open Circuit Voltage Characterization for Li-Ion Batteries

In this article, we present the datasets collected from nine different Li-ion batteries. These datasets contain voltage, current and time measurements during a full charge-discharge cycle of a battery at very low current (that is nearly at �/30 rate). Such low current rate data is suitable for open circuit voltage characterization. The collection of this data was done through the use of an Arbin battery cycler and a thermal chamber was used to control the test temperature. Data were collected over a wide range of temperatures from −25∘C to 50∘C.

Battery management system↗

Evaluation of remote sensing-based evapotranspiration products at low-latitude eddy covariance sites

Remote sensing-based evapotranspiration (ET) products have been evaluated primarily using data from northern middle latitudes; therefore, little is known about their performance at low latitudes. To address this bias, an evaluation dataset was compiled using eddy covariance data from 40 sites between latitudes 30° S and 30° N. The flux data were obtained from the emerging network in Mexico (MexFlux) and from openly available databases of FLUXNET, AsiaFlux, and OzFlux. This unique reference dataset was then used to evaluate remote sensing-based ET products in environments that have been underrepresented in earlier studies. The evaluated products were: MODIS ET (MOD16, both the discontinued collection 5 (C5) and the latest collection (C6)), Global Land Evaporation Amsterdam Model (GLEAM) ET, and Atmosphere-Land Exchange Inverse (ALEXI) ET. Products were compared with unadjusted fluxes (ETorig) and with fluxes corrected for the lack of energy balance closure (ETebc). Three common statistical metrics were used: coefficient of determination (R2), root mean square error (RMSE), and percent bias (PBIAS). The effect of a vegetation mismatch between pixel and site on product evaluation results was investigated by examining the relationship between the statistical metrics and product-specific vegetation match indexes. Evaluation results of this study and those published in the literature were used to examine the performance of the products across latitudes. Differences between the MOD16 collection 5 and 6 datasets were generally smaller than differences with the other products. Performance and ranking of the evaluated products depended on whether ETorig or ETebc was used. When using ETorig, GLEAM generally had the highest R2, smallest PBIAS, and best RMSE values across the studied land cover types and climate zones. Neither MOD16 nor ALEXI performed consistently better than the other. When using ETebc, none of the products stood out in terms of both low bias and strong correlations. The use of ETebc instead of ETorig affected the biases more than the correlations. The product evaluation results showed no significant relationship with the degree of match between the vegetation at the pixel and site scale. The latitudinal comparison showed tendencies of lower R2 (all products) but better PBIAS and normalized RMSE values (MOD16 and GLEAM) for forests at low latitudes than for forests at northern middle latitudes. For non-forest vegetation, the products showed no clear latitudinal differences in performance.

Diego Salazar-Martínez↗

genomeocean: a pretrained microbial genome foundational model (genomeoceanLLM) v1.0

We present Genomeocean, a foundational genome language model that represents the microbial genome sequences from complex environmental samples. By training on a large, diverse metagenomic dataset, Genomeocean learns species-specific sequence composition and can generate long, realistic open reading frames (ORFs). Our model employs a Byte-pair-encoding (BPE) tokenization strategy, allowing it to efficiently process large genomic datasets and generate long sequences up to 50kb. We demonstrate that fine-tuning Genomeocean can generate novel gene clusters encoding biosynthetic pathways, showcasing its ability to model both fundamental and complex biological processes. Our work establishes Genomeocean as a powerful tool for understanding microbial genome biology and paves the way for its application in a range of fields, from synthetic biology to microbiome research.

Wang, Zhong [Lawrence Berkeley National Laboratory↗

NeMO-Net - The Neural Multi-Modal Observation & Training Network for Global Coral Reef Assessment

In the past decade, coral reefs worldwide have experienced unprecedented stresses due to climate change, ocean acidification, and anthropomorphic pressures, instigating massive bleaching and die-off of these fragile and diverse ecosystems. Furthermore, remote sensing of these shallow marine habitats is hindered by ocean wave distortion, refraction and optical attenuation, leading invariably to data products that are often of low resolution and signal-to-noise (SNR) ratio. However, recent advances in UAV and Fluid Lensing technology have allowed us to capture multispectral 3D imagery of these systems at sub-cm scales from above the water surface, giving us an unprecedented view of their growth and decay. By combining spatial and spectral information from varying resolutions, we seek to augment and improve the classification accuracy of previously low-resolution datasets at large temporal scales.NeMO-Net, the first open-source deep convolutional neural network (CNN) and interactive learning and training software, currently being developed at NASA Ames, is aimed at assessing the present and past dynamics of coral reef ecosystems through determination of percent living cover and morphology. The latest iteration uses fully convolutional networks to segment and identify coral imagery taken by UAVs and satellites, including WorldView-2 and Sentinel. We present results taken from the Indian Ocean where classification accuracy has exceeded 91% for 24 geomorphological classes given ample training data. In addition, we utilize deep Laplacian Pyramid Super-Resolution Networks (LapSRN) to reconstruct high resolution information from low resolution imagery, trained from various UAV and satellite datasets. Finally, in the case of insufficient training data, we have developed an interactive online platform that allows users to easily segment and submit their classifications, which has been integrated with the current NeMO-Net workflow. Specifically, we present results from the Fiji islands in which preliminary user data has allowed for the accurate identification of 9 separate classes, despite issues such as cloud shadowing and spectral variation. The project is being supported by NASA's Earth Science Technology Office (ESTO) Advanced Information Systems Technology (AIST-16) Program.

Neural↗

ORBITaL-Net: A labeled training library for large-scale building feature extraction

Over the course of several years, nearly 1.5 million building outlines have been created from approximately 128,000 training tiles covering roughly 7,000 km 2 of very high-resolution multispectral overhead imagery, primarily dated between 2010 and 2020. This dataset, dubbed the Oak Ridge Building Image and TrAining Label Net (ORBITaL-Net), is designed for machine learning applications and is global in scope, with samples drawn from 72 countries across North America, South America, Africa, Europe, and Asia. ORBITaL-Net captures a great diversity in geographic setting, structural characteristics, land use (urban and rural), terrain, and imagery conditions. While the labeled building outlines are themselves valuable, the dataset’s true strength lies in the pairing of these labels with corresponding reference imagery, which is being released for open source use. Similar to SpaceNet and Replicable AI For Microplanning (ramp), this building outline dataset will allow the larger computer vision community from academia, government, and industry the opportunity to develop robust, scalable, and generalizable geospatial machine learning techniques. Unlike SpaceNet and ramp, which offer high resolution labels and imagery primarily for large urban cities, ORBITaL-Net is not focused on training samples from heavily populated areas but instead aims to capture the innate variability of conditions present in both the physical environment and imagery collections.

Geography↗

Data for Zheng et al. (2025), "AquaMEND: Reconciling multiple impacts of salinization on soil carbon biogeochemistry"

Soil salinization, exacerbated by climate change, poses a global threat to coastal ecosystems and soil function. Salinity affects soil carbon cycling by directly impacting microbial activity and indirectly altering soil physicochemical properties, but current models inadequately represent these complexities. This dataset contains the observational and modeling data from Zheng et al. (2025), which described a process-based modeling framework that couples soil solution chemistry with microbial carbon cycling reactions to study the impacts of soil salinization. This conceptual model is implemented numerically into the open-source geochemical program PHREEQC 3.0 (Parkhurst and Appelo, 2013). This dataset consists of: - Figure2_AquaMEND_salinity_buffer: Contains model simulation outputs to assess the impact of three different cation exchange and surface complexation processes on salinity buffering (Fig. 2 from Zheng et al. 2025). - Figure3_Salinity_function: Contains salinity function fitting for literature data (Fig. 3 from Zheng et al. 2025). - Figure4_AquaMEND_microbial_mechanisms: Contains model simulation outputs for testing various microbial process-based hypotheses related to soil salinization, including microbial mortality, carbon use efficiency (CUE), extracellular enzyme activity, and other microbial mechanisms (Fig. 4 from Zheng et al. 2025). - Figure5_AquaMEND_Redox: Contains on model simulation outputs to evaluate shifts among key redox processes, such as aerobic respiration, sulfate reduction, and methanogenesis (Fig.5 from Zheng et al. 2025). - Figure6_AquaMEND_sorption: Contains on model simulation outputs for investigating the effects of salinity on dissolved organic matter (DOM) sorption and desorption processes (Fig. 6 from Zheng et al. 2025). - Figure7_AquaMEND_process_couple: Contains on model simulation outputs for exploring coupled biotic-abiotic processes and their interactions (Fig. 7 from Zheng et al. 2025). - data: Includes datasets used to develop salinity response functions and evaluate salinity buffering capacity. Datasets for MEND model calibration. - database: Contains the `.dat` file required by PHREEQC for model execution. - README.md: A Markdown plain text file describing the computational tools and directories. Files are a mixture of plain text CSV (comma-separated value) and plain text *.dat files written by the model; no special software is required to read them.

EARTH SCIENCE > AGRICULTURE > SOILS > SOIL SALINIT↗

Utilizing Google Scholar as a Bibliographic Resource for Publications Search

This study focuses on the automated search for publication citations for the Earth Observing System Data and Information System (EOSDIS) datasets. The research investigates the feasibility of using automated search methods to gather published works from various bibliometric databases. A comparison is presented, highlighting the differences in citation counts obtained from Google Scholar compared to established bibliographic databases. The study also introduces a methodology and an open-source tool for getting publication citations from Google Scholar, utilizing dataset DOIs and keyword searches. The findings contribute to understanding the reliability and effectiveness of Google Scholar as a source for dataset citation retrieval and provide researchers with a valuable resource for obtaining comprehensive citation data.

Infometrics↗

Fast and Invertible Simplicial Approximation of Magnetic‐Following Interpolation for Visualizing Fusion Plasma Simulation Data

We introduce a fast and invertible approximation for fusion plasma simulation data represented as 2D planar meshes with connectivities approximating magnetic field lines along the toroidal dimension in deformed 3D toroidal spaces. Scientific variables (e.g., density and temperature) in these fusion data are interpolated following a complex magnetic-field-line-following scheme in the toroidal space represented by a cylindrical coordinate system. This deformation in the 3D space poses challenges for root-finding and interpolation. To this end, we propose a novel paradigm for visualizing and analyzing such data based on a newly developed algorithm for constructing a 3D simplicial mesh within the deformed 3D space. Our algorithm generates a tetrahedral mesh that connects the 2D meshes using tetrahedra while adhering to the constraints on node connectivities imposed by the magnetic field-line scheme. Specifically, we first divide the space into smaller partitions to reduce complexity based on the input geometries and constraints on connectivities. Then, we independently search for a feasible tetrahedralization of each partition, considering nonconvexity. We demonstrate our method with two X-Point Gyrokinetic Code (XGC) simulation datasets on the International Thermonuclear Experimental Reactor (ITER) and Wendelstein 7-X (W7-X), and use an ocean simulation dataset to substantiate broader applicability of our method. An open source implementation of our algorithm is available at https://github.com/rcrcarissa/DeformedSpaceTet.

Ren, Congrong [The Ohio State Univ., Columbus, OH ↗

Smart Meter Data: A Gateway for Reducing Solar Soft Costs with Model-Free Hosting Capacity Maps

Public-facing solar hosting capacity (HC) maps, which show the maximum amount of solar energy that can be installed at a location without adverse effects, have proven to be a key driver of solar soft cost reductions through a variety of pathways (e.g., streamlining interconnection, siting, and customer acquisition processes). However, current methods for generating HC maps require detailed grid models and time-consuming simulations that limit both their accuracy and scalability—today, only a handful out of almost 2,000 utilities provide these maps. This project developed and validated data-driven algorithms for calculating solar HC using data from AMI without the need of detailed grid models or simulations. The algorithms were validated on utility datasets and incorporated as an application into NRECA’s Open Modeling Framework (OMF.coop) for the over 260 coops and vendors throughout the US to use. The OMF is free and open-source for everyone.

14 SOLAR ENERGY↗

Storage of Physical Sample Metadata in the Astrobiology Habitable Environments Database (AHED)

The National Aeronautics and Space Administration has begun an effort to store, curate, and publish information about physical samples collected and analyzed in conjunction with NASA-funded astrobiology research. Astrobiology is a multidisciplinary area of scientific research being conducted by collaborating teams of biologists, chemists, geologists, atmospheric scientists, oceanographers, astrophysicists, astronomers, and other specialists. Astrobiology studies the origin, evolution, and distribution of life in the Universe. NASA uses the results of astrobiology research to focus its future missions on targets of opportunity for the discovery of life off Earth. Astrobiology researchers conduct both field-based and laboratory-based research, during which physical samples are collected, processed, and catalogued. The cataloguing practices employed by different teams of astrobiologists vary widely, and there are no specific standards available to guide the collection and recording of astrobiology sample data. The disparity in data collection approaches and the lack of a centralized sample repository makes it difficult for astrobiology teams to share data and benefit from resultant synergies.To facilitate data sharing within the astrobiology community, NASA is developing a prototype database the Astrobiology Habitable Environments Database (AHED) and an associated set of data collection templates. The database will store information about samples, along with associated measurements and analyses, including information about biological cultures enriched or isolated from samples, and the results of analyses performed on the samples (e.g., via spectrography, microscopy, etc.). In addition, the system will store contextual information about field sites where samples were collected, the instruments or equipment used for analysis, and people and institutions involved in their collection. AHED is being implemented on top of Open Data Repository's Data Publisher [1], an open source software platform for the publication of scientific datasets. The data collection templates under development represent an initial attempt to propose a set of metadata for capture and storage within AHED. The design of these templates is being conducted by a consolidated group of astrobiologists from active research teams at NASA Ames Research Center, assisted by data science and software engineering specialists. These initial templates must be vetted with the broader astrobiology community through a defined process to ensure that they meet community needs. Each template captures a different type of data collection record. For each template, we are developing a list of fields to be captured, including a set of required entry fields, a set of recommended but optional fields, and a set of discretionary fields. A datatype selected from a variety of text and numeric types is specified for each field. Included is a 'choice' type that restricts user input to an enumerated list of values. Many of the fields and field values capture information of particular interest to the astrobiology community, and are intended to facilitate search and retrieval of relevant data across multiple datasets.

Keller, Rich↗

Investigating the Relationship between the Cell Wall Integrity Pathway and Unfolded Protein Response

Plants have made significant contributions to astronaut health in spaceflight missions. To further spaceflight research in optimizing plant viability, this study aims to understand the factors involved in maintaining cell wall integrity, which is vital to plant morphology and structural stability. Spaceflight can negatively impact the cell wall; thus, it is crucial to investigate how to mitigate spaceflight stressors to maintain the integrity of the cell wall. The structural integrity of plants’ cell walls depends on secondary cell wall biogenesis, which enables the repair and architectural support of plants like A. thaliana. This biogenesis is triggered by a signal transduction cascade: first initiated by cell wall stress, the CWI (cell wall integrity) pathway is activated, followed by the UPR (unfolded protein response), then the cell wall’s integrity is maintained through secondary cell wall biogenesis. Through a re-analysis of GeneLab Dataset 321 (GLDS-321), a study from NASA’s Open Science Data Repository that investigates the effects of spaceflight on the UPR, several genes were found to be associated with the cell wall. This proposal postulates a relationship between the UPR and the CWI pathway and their direct effect on secondary cell wall biogenesis by investigating IRX7, a gene associated with secondary cell wall biogenesis. The predicted outcome of overexpressing IRX7 is increased resilience of the cell wall by upregulating both the UPR and the CWI pathway, while silencing IRX7 is predicted to compromise the cell wall integrity by downregulating the UPR and the CWI pathway. This study will give insight into the needed measures to increase cell wall resilience in stressful environments: As spaceflight durations increase and uncertain climate change events progress on Earth, understanding how to optimize cell wall resilience – a fundamental pillar of plant health – can effectively enhance mass crop production and quality and ensure the physical and psychological health of astronauts in long-term space missions.

GL4HS↗

Livewire User Guide

The Livewire Data Platform houses a catalog of transportation- and mobility-related project data, as well as a publications database, making it easy to search and share data. It allows transportation researchers, industry, and academic partners to increase the visibility of their projects within the research community, securely share and preserve data, and leverage datasets from other projects. Public data on Livewire are open to anyone with a Livewire account. This guide will help Livewire users understand how to store project data as a data steward, as well as access data as a data consumer.

33 ADVANCED PROPULSION SYSTEMS↗

NeMO-Net: The Neural Multi-Modal Observation and Training Network for Global Coral Reef Assessment

In the past decade, coral reefs worldwide have experienced unprecedented stresses due to climate change, ocean acidification, and anthropomorphic pressures, instigating massive bleaching and die-off of these fragile and diverse ecosystems. Furthermore, remote sensing of these shallow marine habitats is hindered by ocean wave distortion, refraction and optical attenuation, leading invariably to data products that are often of low resolution and signal-to-noise (SNR) ratio. However, recent advances in UAV and Fluid Lensing technology have allowed us to capture multispectral 3D imagery of these systems at sub-cm scales from above the water surface, giving us an unprecedented view of their growth and decay. Exploiting the fine-scaled features of these datasets, machine learning methods such as MAP, PCA, and SVM can not only accurately classify the living cover and morphology of these reef systems (below 8 percent error), but are also able to map the spectral space between airborne and satellite imagery, augmenting and improving the classification accuracy of previously low-resolution datasets. We are currently implementing NeMO-Net, the first open-source deep convolutional neural network (CNN) and interactive active learning and training software to accurately assess the present and past dynamics of coral reef ecosystems through determination of percent living cover and morphology. NeMO-Net will be built upon the QGIS platform to ingest UAV, airborne and satellite datasets from various sources and sensor capabilities, and through data-fusion determine the coral reef ecosystem makeup globally at unprecedented spatial and temporal scales. To achieve this, we will exploit virtual data augmentation, the use of semi-supervised learning, and active learning through a tablet platform allowing for users to manually train uncertain or difficult to classify datasets. The project will make use of Pythons extensive libraries for machine learning, as well as extending integration to GPU and High-End Computing Capability (HECC) on the Pleiades supercomputing cluster, located at NASA Ames. The project is being supported by NASAs Earth Science Technology Office (ESTO) Advanced Information Systems Technology (AIST-16) Program.

NeMO-Net↗

Silica Retention and Enrichment in Open-System Chemical Weathering on Mars

Chemical signatures of weathering are evident in the Alpha Particle X-ray Spectrometer (APXS) datasets from Gusev Crater, Meridiani Planum, and Gale Crater. Comparisons across the landing sites show consistent patterns indicating silica retention and/or enrichment in open-system aqueous alteration.

Yen, A. S.↗

Subhourly Clipping Correction Model Comparison

This work will compare the Allen method and Walker method of accounting for subhourly inverter clipping power losses in hourly PV performance models. The Allen method uses a matrix lookup based on DNI clearness and clipping potential to assign a clipping correction loss at each simulation timestep. The Walker method models the PV DC power input to the inverter as a distribution over the hourly timestep and uses integration over the timestep to determine the amount of clipping that occurs within the timestep. Both these models have been recently implemented in the System Advisor Model's (SAM) open-source code, and will be applied to hourly SURFRAD datasets to analyze the subhourly clipping loss predicted by each model for different system designs and inverter loading conditions. Both models will be compared to "true" 1-minute SURFRAD data simulations to see their accuracy against more accurate 1-minute clipping correction loss predictions. This model comparisons will be investigated in more detail at the PVSC conference in Seattle, Washington June 2024.

ENERGY PLANNING, POLICY, AND ECONOMY,MATHEMATICS A↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

Design Thinking for the Applied Sciences: Developing a Novel Approach to Encourage the Use of Synthetic Aperture Radar (SAR) and Open Source Tools for Forest Monitoring

Earth observations from Synthetic Aperture Radar, or SAR, have yet to be fully leveraged for forest monitoring applications. While SAR sensors are uniquely able to capture components of forest structure over optical imagery, especially in cloud-heavy regions, there is a shortage of freely-available applied training materials and related case studies. With the wealth of available datasets from Sentinel-1 and other missions, such as ALOS-Palsar open historical archive, and in preparation for upcoming opendata policy SAR missions (e.g. NISAR and BIOMASS), the applied forestry community would benefit from increased access to relevant, understandable SAR training materials. This work documents lessons learned and best practices for creating EO capacity building/training materials gleaned from the SAR Handbook project. Strategies for increasing legibility for both print and online applications, illustration and editing guidelines for original and modified figures, and the development of quick-reference guides will be shared. Additionally, the conception and use of companion “explainer” videos, using cartoon characters and humor to outline relevant SAR concepts will be explored. Preliminary results indicate the SAR Handbook and supplemental project materials are already having an impact in training sessions. Increased uptake of SAR technologies in SERVIR Hub regions, where Hubs are leading follow-on SAR trainings, has also been noted. In addition, a review of download statistics from the SERVIR global website indicates widespread worldwide access. We conclude similar holistic approaches integrating design concepts into future content development would help increase uptake of EO applications by the earth science community.

Kucera, Leah M.↗

2019 Meander C and Meander Z floodplain groundwater chemistry from the East River Watershed, CO, USA

This dataset includes groundwater geochemistry data from floodplain piezometers collected as a part of the Watershed Function Scientific Focus Area (SFA) located in the Upper Colorado River Basin. The data were collected in order to investigate the role of hyporheic exchange and other river corridor processes on riverine export of solutes. Data includes samples from two intra-meander zones: Meander C, in the Pumphouse vicinity, and Meander Z, just upstream of the confluence with Brush Creek. Floodplain piezometers installed along two transects across Meander C (MCP and MCB wells) and Meander Z (MZA and MZB wells) were sampled on daily to weekly time scales during summer-fall 2019. Some river water grab samples are also included. Data includes in-field measurements (pH, electrical conductivity [EC], oxidation reduction potential [ORP], dissolved oxygen [DO], and groundwater level) along with laboratory measurements (dissolved inorganic carbon [DIC], dissolved organic carbon [DOC], metals and major cations, anions [chloride, sulfate, nitrate], and dissolved ammonium). Files are included in this dataset include: sample locations and depths in both a kmz file which can be opened in Google Earth and a csv file, aqueous geochemistry data in a csv files for Meander C and Meander Z, and analytical detection limits in a csv file. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗