Search NASA⌕ Search

SEARCH · Search NASA

Results for “database for machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Non-Electricity Based Renewable Fuels: Theory and Computation for Solar Thermochemical Hydrogen

Dominated by photovoltaics and wind, current renewable energy sources generate mostly electricity, but 80% of the global final energy consumption occurs in form of fuels. Therefore, direct solar fuel generation would be a major breakthrough for the energy transition. Solar thermochemical hydrogen (STCH) is one of the very few potential routes towards scalable renewable fuels, but currently suffers from lack of an oxide working material that could optimally perform energy conversion within the thermodynamic boundary conditions. Theory and computation can contribute in two distinct ways, through materials search and discovery, but also by providing detailed mechanistic models for specific systems so to advance our understanding of possible design strategies. To enable high-throughput materials screening, we developed a defect graph neural network (dGNN) machine learning approach,[1] which accelerates the prediction of defect formation energies by replacing the tedious density functional theory (DFT) supercell calculations for all possible defect sites. This approach enables high-throughput database screening of oxides, which was integrated with thermodynamic modeling to extract the reduction entropies as additional selection criterion for STCH. Once potential candidate materials are identified, detailed models can guide materials design by predicting performance characteristics. One challenge is to quantitatively predict thermochemical equilibria at high concentrations when the redox active defects start to interact with each other, thereby impeding the formation of additional defects. Introducing a model for the free energy of defect interaction, parametrized on the basis of DFT data, we simulated the complete STCH redox cycle for (Sr,Ce)MnO3 alloys, achieving near-quantitative agreement with experimental data.[2] The analysis of these simulations reveals how defect interactions diminish the reduction entropy and H2 yield, suggesting to include these interactions in design considerations. Finally, we revisit the popular van't Hoff method for analyzing reduction enthalpies and entropies. This method is not ideal, as it involves a temperature-dependent convolution of gas-phase and solid-state entropies, causing uncertainties in the same order of magnitude as the physical quantities of interest. To avoid this problem, we suggest a simple alternative approach which can be applied to experimental and simulated data alike.

first-principles calculations↗

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES↗

Artificial Intelligence and Machine Learning Applications in Modern Power Systems

Machine learning (ML) and artificial intelligence (AI) algorithms offer valuable tools for the analysis and interpretation of large datasets. These tools have the capability to uncover insights that may not be readily apparent within these datasets. In recent years, the integration of ML and AI has become increasingly prevalent in various applications within the power system domain. One of the earliest instances of machine learning in power systems can be traced back to demand forecasting, where artificial neural networks were employed for short-term load forecasting. In contemporary power systems, an abundance of high-resolution geospatial and temporal data is generated at various time intervals, ranging from sub-seconds (Phasor Measurement Units or PMUs) to seconds (Supervisory Control and Data Acquisition or SCADA), minutes (Process Information or PI), and extending to days, months, and years. These datasets contain valuable information concerning system reliability and performance. This information holds the potential to offer critical insights into system operations, as well as solutions for predicting and mitigating contingencies to prevent cascading outages. Despite the immense power of machine learning tools, system operators, planners, and utilities often exhibit hesitancy in fully embracing AI-enabled system operations and planning. This cautious approach persists, even as numerous diverse applications of machine learning continue to emerge in the realm of power systems. In this chapter, our focus will delve deep into ML and AI applications tailored for power systems. These applications aim to furnish system operators with enhanced situational awareness and augment their decision-making capabilities, especially during challenging operating conditions. Specific areas of interest encompass root cause analyses of electricity market datasets and the strategic selection of representative samples from vast power system databases for training ML/AI models. Finally, the chapter will conclude with a short discussion on the future of ML/AI in power systems and possible directions that the industry is moving towards.

power system applications, machine learning (ML), ↗

Temperature‐Dependent Crystallization in Two‐Step Perovskite Deposition Revealed by In Situ GIWAXS and Machine Learning‐Guided Analysis

The performance and stability of perovskite solar cells are strongly governed by the crystallization behavior of their active layer. In two-step sequential deposition, early-stage film formation plays a decisive role in determining final phase purity and device quality. Guided by a data-driven analysis of nearly 39 000 devices in the FAIR perovskite database, we identified solvent-mediated quenching and thermal processing as key variables affecting power conversion efficiency (PCE), particularly in two-step fabrication. Here, to investigate these effects in real time, we designed and implemented a custom-built, temperature-controlled spin-coating system, enabling precise thermal modulation during precursor deposition. Using this platform, we performed in situ GIWAXS measurements to study the crystallization dynamics of FA 0.5 MA 0.5 PbI 3 films over a temperature range of 30°C–90°C. Our results reveal a non-monotonic relationship between spin-coating temperature and α-phase formation, governed by the interplay between precursor interdiffusion, PbI 2 crystallinity, and δ-phase suppression. The custom thermal control enabled us to isolate and quantify these competing effects during the earliest stages of film formation, providing mechanistic insight into how spin-coating temperature governs both phase purity and kinetic pathways in two-step perovskite systems. Temperature-dependent SEM and photovoltaic device measurements further demonstrate that early-stage crystallization pathways directly translate into differences in morphology, charge-transport continuity, and device performance. These findings inform targeted strategies for optimizing deposition protocols to balance rapid nucleation, phase stability, and device performance.

Saadawy, Ahmed [King Fahd University of Petroleum ↗

A machine learning estimator trained on synthetic data for real-time earthquake ground-shaking predictions in Southern California

Abstract After large-magnitude earthquakes, a crucial task for impact assessment is to rapidly and accurately estimate the ground shaking in the affected region. To satisfy real-time constraints, intensity measures are traditionally evaluated with empirical Ground Motion Models that can drastically limit the accuracy of the estimated values. As an alternative, here we present Machine Learning strategies trained on physics-based simulations that require similar evaluation times. We trained and validated the proposed Machine Learning-based Estimator for ground shaking maps with one of the largest existing datasets (<100M simulated seismograms) from CyberShake developed by the Southern California Earthquake Center covering the Los Angeles basin. For a well-tailored synthetic database, our predictions outperform empirical Ground Motion Models provided that the events considered are compatible with the training data. Using the proposed strategy we show significant error reductions not only for synthetic, but also for five real historical earthquakes, relative to empirical Ground Motion Models.

Environmental Sciences & Ecology↗

Atomic cluster expansion potential for large scale simulations of hydrocarbons under shock compression

We present an Atomic Cluster Expansion (ACE) machine learned potential developed for high-fidelity atomistic simulations of hydrocarbons, targeting pressures and temperatures near and above supercritical fluid regimes for molecular fluids. A diverse set of stoichiometries were covered in training, including 1:0 (pure carbon), 1:4 (methane), and 1:1 (benzene), and rich bonding environments sampled at supercritical temperatures, hydrogen rich, reactive mixtures where metastable stoichiometries arise, including 1:2 (ethylene) and 1:3 (ethane). A high-fidelity training database was constructed by performing large-scale quantum molecular dynamic simulations [density functional theory (DFT) MD] of diamond, graphite, methane, and benzene. A novel approach to selecting structures from DFT MD is also presented, which allows for the rapid selection of unique DFT MD frames from complex trajectories. Comparisons to DFT and experimental data demonstrate that the presented ACE potential accurately reproduces isotherms, carbon melting curves, radial distribution functions, and shock Hugoniots for carbon and hydrocarbon systems for pressures up to 100 GPa and temperatures up to 6000 K for hydrocarbon systems and up to 9000 K for pure carbon systems. This work delivers a potential that can be used for accurate, large-scale simulations of shocked hydrocarbons and demonstrates a methodology for fitting and validating machine learning interatomic potentials to complex molecular environments, which can be applied to energetic materials in future works.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Prototype crystal structures for exa-AMD framework

This dataset provides a curated collection of initial crystal structure prototypes for the exa-AMD framework, serving as templates for generating hypothetical candidates in multinary materials discovery through elemental substitution and scaling. It includes 36553 ternary and 5254 quaternary unique structures sourced from public databases such as Materials Project, GNoME, OQMD, AFLOW, and NovoMag, formatted in CIF.

exa-AMD↗

Advancing AI-Driven Analysis in X-ray Absorption Spectroscopy: Spectral Domain Mapping and Universal Models

In recent years, rapid progress has been made in developing artificial intelligence (AI) and machine learning (ML) methods for X-ray absorption spectroscopy (XAS) analysis. Compared to traditional XAS analysis methods, AI/ML approaches offer dramatic improvements in efficiency and help eliminate human bias. To advance this field, we advocate an AI-driven XAS analysis pipeline that features several interconnected key building blocks: benchmarks, workflows, databases, and AI/ML models. Specifically, we present two case studies for XAS ML. In the first study, we demonstrate the importance of reconciling the discrepancies between simulation and experiment using spectral domain mapping (SDM). Our ML model, which is trained solely on simulated spectra, predicts an incorrect oxidation state trend for Ti atoms in a combinatorial zinc titanate film. After transforming the experimental spectra into a simulation-like representation using SDM, the same model successfully recovers the correct oxidation state trend. In the second study, we explore the development of universal XAS ML models that are trained on the entire periodic table, which enables them to leverage common trends across elements. Looking ahead, we envision that an AI-driven pipeline can unlock the potential of real-time XAS analysis to accelerate scientific discovery.

36 MATERIALS SCIENCE↗

Life Cycle Inventory Availability: Status and Prospects for Leveraging New Technologies

The demand for life cycle assessments (LCA) is growing rapidly, which leads to an increasing demand of life cycle inventory (LCI) data. While the LCA community has made significant progress in developing LCI databases for diverse applications, challenges still need to be addressed. This perspective summarizes the current data gaps, transparency, and uncertainty aspects of existing LCI databases. Additionally, we survey and discuss novel techniques for LCI data generation, dissemination, and validation. We propose key future directions for LCI development efforts to address these challenges, including leveraging scientific and technical advances such as the Internet of Things (IoT), machine learning, and blockchain/cloud platforms. Adopting these advanced technologies can significantly improve the quality and accessibility of LCI data, thereby facilitating more accurate and reliable LCA studies.

blockchain platforms↗

Metrics and extrapolation of resonant magnetic perturbation thresholds for ELM suppression

This large database study of resonant magnetic perturbation (RMP) edge localized mode (ELM) suppression thresholds in the AUG, DIII-D, EAST, and KSTAR tokamaks details the key strengths and weaknesses of RMP metrics. The RMP ELM suppression database used for this work contains plasma information at the time of transition from ELMing to ELM suppressed states where a clear experimental threshold is identified. The experimental threshold distributions are compared for five metrics: (1) the island overlap width, (2) pedestal top Chirikov overlap, (3) peeling edge displacement, (4) pedestal top resonant drive, and (5) edge dominant mode overlap. The distributions, the regularity of the dependence on RMP coil currents, and the sensitivities of a given metric to equilibrium reconstruction details are compared. The overlap metric proves to be a good compromise between including the appropriate plasma response physics and maintaining a numerical robustness. This quantity does not exhibit clear power-law scalings for projection, but machine learning can assist in predicting thresholds within the existing parameter ranges and providing uncertainty quantification of those predictions. Two new first-principles models, one utilizing a threshold from the non-linear Modified Rutherford equation evaluated at the pedestal top and one utilizing the SLAYER code to calculate the linear tearing threshold from torque balance, offer possible paths to extrapolation beyond the existing database parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗

Computational toolkit for predicting thickness of 2D materials using machine learning and autogenerated dataset by large language model

The thickness of 2D materials not only plays a crucial role in determining the performance of nanoelectronic and optoelectronic devices but also introduces complexities in predicting volume-dependent properties, such as energy storage capacity, due to the intrinsic vacuum within these materials. Although a plethora of experimental techniques, including but not limited to optical contrast, Raman spectroscopy, nonlinear optical spectroscopy, near-field optical imaging, and hyperspectral imaging, facilitate the measurement of 2D material thickness, comprehensive data for many materials remain elusive. Over the past decade, the exponential proliferation of 2D materials and their heterostructures has outstripped the capabilities of conventional experimental and computational approaches. In this evolving landscape, machine learning (ML) has emerged as an indispensable tool, offering a scalable approach to augment these traditional methodologies. Addressing the critical gap, we introduce THICK2D—Thickness Hierarchy Inference and Calculation Kit for 2D Materials. This Python-based computational framework harnesses an autogenerated thickness database, developed using large language models, and advanced ML algorithms to facilitate the rapid and scalable estimation of material thickness, relying solely on crystallographic data. To demonstrate the utility and robustness of THICK2D, we successfully used the toolkit to predict the thickness of more than 8000 2D-based materials, sourced from two extensive 2D materials databases. THICK2D is disseminated as an open-source utility, accessible on GitHub at https://github.com/gmp007/THICK2D, and archived on Zenodo at https://10.5281/zenodo.11216648.

Ekuma, Chinedu E. (ORCID:0000000258527556)↗

Mapping wall-to-wall fractional cover of Arctic tundra plant functional types in Alaska using 20-m spatial resolution satellite imagery and harmonized plot observations

Estimates of fractional cover (fCover) across given land surfaces are used to assess, and often model, vegetation composition and diversity, which are crucial for understanding the health and functioning of terrestrial ecosystems. Remote sensing provides a useful means for scaling local, plot-measured fCover estimates to regional scales. Leveraging a recently synthesized and harmonized plot database, this study generated wall-to-wall maps of fCover for six Alaskan-Arctic plant functional types (PFT), including non-vascular plants, forbs, graminoids, and deciduous and evergreen shrubs, using 20-m satellite data (Sentinel-1, Sentinel-2, ArcticDEM) using a machine learning regression approach, specifically the random forest (RF) algorithm, which is well-suited for handling nonlinear relationships and high-dimensional satellite datasets. This study additionally addressed the spatio-temporal inconsistencies e.g., sampling scale, plot size, and collection year in plot measured fCover by adopting a multivariate outlier detection approach—Cook’s distance—to identify high-quality plots for model training and validation. Our approach achieves high accuracy (R 2 = 0.59–0.93, root mean squared errors = 0.02–0.10 for all PFTs) between plot-observed and satellite-derived fCover when using high-quality plot samples. The mapped fCover characterizes the spatial patterns of different PFTs across the tundra biome at a 20-m resolution, providing key information needed for improved representation of Arctic tundra vegetation in terrestrial biosphere models to better understand climate-vegetation feedback across the Arctic tundra.

Arctic tundra↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

Autonomous Synthesis and Inverse Design of Electrochromic Polymers with High Efficiency and Accuracy

Here, the design and synthesis of functional polymers, aimed at targeted properties through specific structures, have long been challenged by their complex and often nonlinear structure–property relationships. Key processes, including knowledge accumulation for predictive design and experimental refinement and validation, are traditionally labor-insensitive and time-consuming, making it difficult to balance accuracy and efficiency. Here, we introduce an accelerated, autonomous system for the on-demand synthesis of electronic polymers that achieves the desired electrochromic functionality with high accuracy and efficiency. Our approach leverages large language model-assisted data mining, a physics-informed copolymer machine learning model, and an AI-driven autonomous robotic workflow in the Polybot lab. Within 72 h, Polybot autonomously synthesized electrochromic polymers (ECPs) with targeted, previously-unreported color values, including green polymers with specific absorption profiles, precisely fine-tuning copolymer structures with a 5% step size in comonomer composition within a three-monomer system. A publicly accessible ECP informatics database has also been created to foster knowledge exchange.

AI-driven Robotic Lab↗

Data Science-Driven Discovery of Multimetallic Oxygen-cycle Electrocatalysts for Enhanced Energy Conversion

The overarching objective of this effort has been to combine state-of-the-art data science techniques, first principles analyses, and molecular-level characterization of electrocatalyst structure and reactivity to identify both in-situ mechanisms for degradation and transformation of electrocatalysts with highly complex catalytic structures and the impact of these transformations on catalytic activity. The primary catalysts of interest have been multielemental alloys, including high entropy alloys (HEA’s), which are characterized by a high degree of disorder and up to 20 different elements within a single nanoparticle. We have applied these strategies primarily to energy-critical oxygen cycle electrocatalytic reactions, including oxygen reduction (ORR), but we have also considered extensions to non-electrochemical chemistries such as ammonia synthesis and decomposition. We have made strong progress in the development of computational methods on both the level of machine learning methods development as well as first principles-based treatments of HEA’s, and we have leveraged these insights to propose promising HEA catalysts for the ORR. On the experimental side, we developed new HEA synthesis and characterization protocols relevant to these reactions and developed a database combining our experimental results with corresponding computational tools.

36 MATERIALS SCIENCE↗

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),↗

deadtrees.earth — An open-access and interactive database for centimeter-scale aerial imagery to uncover global tree mortality dynamics

Excessive tree mortality is a global concern and remains poorly understood as it is a complex phenomenon. We lack global and temporally continuous coverage on tree mortality data. Ground-based observations on tree mortality, e.g., derived from national inventories, are very sparse, and may not be standardized or spatially explicit. Earth observation data, combined with supervised machine learning, offer a promising approach to map overstory tree mortality in a consistent manner over space and time. However, global-scale machine learning requires broad training data covering a wide range of environmental settings and forest types. Low altitude observation platforms (e.g., drones or airplanes) provide a cost-effective source of training data by capturing high-resolution orthophotos of overstory tree mortality events at centimeter-scale resolution. Here, we introduce deadtrees.earth, an open-access platform hosting more than two thousand centimeter-resolution orthophotos, covering more than 1,000,000 ha, of which more than 58,000 ha are manually annotated with live/dead tree classifications. This community-sourced and rigorously curated dataset can serve as a comprehensive reference dataset to uncover tree mortality patterns from local to global scales using space-based Earth observation data and machine learning models. This will provide the basis to attribute tree mortality patterns to environmental changes or project tree mortality dynamics to the future. The open nature of deadtrees.earth, together with its curation of high-quality, spatially representative, and ecologically diverse data will continuously increase our capacity to uncover and understand tree mortality dynamics.

Citizen science↗