Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data-driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Mapping Global Forest Age from Forest Inventories, Biomass and Climate Data

Forest age can determine the capacity of a forest to uptake carbon from the atmosphere. However, a lack of global diagnostics that reflect the forest stage and associated disturbance regimes hampers the quantification of age-related differences in forest carbon dynamics. This study provides a new global distribution of forest age circa 2010, estimated using a machine learning approach trained with more than 40 000 plots using forest inventory, biomass and climate data. First, an evaluation against the plot-level measurements of forest age reveals that the data-driven method has a relatively good predictive capacity of classifying old-growth vs. non-old-growth (precision = 0.81 and 0.99 for old-growth and non-old-growth, respectively) forests and estimating corresponding forest age estimates (NSE = 0.6 – Nash–Sutcliffe efficiency – and RMSE = 50 years – root-mean-square error). However, there are systematic biases of overestimation in young- and underestimation in old-forest stands, respectively. Globally, we find a large variability in forest age with the old-growth forests in the tropical regions of Amazon and Congo, young forests in China, and intermediate stands in Europe. Furthermore, we find that the regions with high rates of deforestation or forest degradation (e.g. the arc of deforestation in the Amazon) are composed mainly of younger stands. Assessment of forest age in the climate space shows that the old forests are either in cold and dry regions or warm and wet regions, while young–intermediate forests span a large climatic gradient. Finally, comparing the presented forest age estimates with a series of regional products reveals differences rooted in different approaches and different in situ observations and global-scale products. Despite showing robustness in cross-validation results, additional methodological insights on further developments should as much as possible harmonize data across the different approaches. The forest age dataset presented here provides additional insights into the global distribution of forest age to better understand the global dynamics in the forest water and carbon cycles. The forest age datasets are openly available at https://doi.org/10.17871/ForestAgeBGI.2021 (Besnard et al., 2021).

Simon Besnard↗

Advancing Open Science through Public-Private Partnerships

Rapid technology developments are changing the way data-driven research is performed within the science community. With the emergence of cloud computing, this has quickly become a viable approach for enabling “science at scale”. Researchers are no longer hindered by obstacles of data management and data wrangling, allowing them to quickly discover, access and perform analysis on extremely large datasets. Infrastructures that move data out of institutional silos and into a computational platform, will ensure that data and tools are accessible to all users. NASA’s Interagency Implementation and Advanced Concepts Team (IMPACT) seeks to address these rapid technology developments by establishing Space Act Agreements with selected partners from the public-private sector working in the area of cloud computing. These agreements aim to explore new opportunities with commercial cloud providers to accelerate open science and enable discovery, access and use of data sets on the cloud. In addition, they will also help establish training workshops for the science community to help researchers utilize the cloud for science. In this talk, we will present an overview of current and new partnerships we are developing to support open science and open data initiatives.

Elizabeth Fancher↗

Natural Language Processing Methods for Air Traffic Management Text and Speech Data

This presentation discusses two efforts of the NARI AI/ML Intern team during the Fall 2021 OSTEM Internship term. For Letters of Agreement (LoA), we have studied how LoAs are structured and explored the question ‘What is an LoA constraint?’ To do this, our approach is data-driven, iterative, and assisted by machine learning when available. In this presentation, we will walk through our tasks of manually scanning through documents, performing a preliminary entity labelling task, and our unsupervised analysis on LoA procedures sections. After this research phase, we define the smallest constraint unit in an LoA, and start to perform entity extraction. Looking towards constraint extraction, we are also exploring the use of a one-class support vector machine (OneClassSVM) model to identify patterns within the data. The second effort of our team this term is focused on Air Traffic Control System Command Center (ATCSCC) advisory meetings, and the subsequent advisory documents that get published from their content. These advisory documents are important to give readily accessible summaries of daily operations, so that data centers, airline officials, and other stakeholders can easily understand the context of these meetings in real time. In applying machine learning to this scenario, two natural language processing tasks are used. First is developing machine learning models to convert the meeting speech data into text. With this text, use of extractive and abstractive text summarization models are used to automatically generate preliminary versions of the advisory documents.

Natural Language Processing↗

Machine Vision based Sample-Tube Localization for Mars Sample Return

A potential Mars Sample Return (MSR) architecture is being jointly studied by NASA and ESA. As currently envisioned, the MSR campaign consists of a series of 3 missions: sample cache, fetch and return to Earth. In this paper, we focus on the fetch part of the MSR, and more specifically the problem of autonomously detecting and localizing sample tubes deposited on the Martian surface. Towards this end, we study two machine-vision based approaches: First, a geometrydriven approach based on template matching that uses hardcoded filters and a 3D shape model of the tube; and second, a data-driven approach based on convolutional neural networks (CNNs) and learned features. Furthermore, we present a large benchmark dataset of sample-tube images, collected in representative outdoor environments and annotated with ground truth segmentation masks and locations. The dataset was acquired systematically across different terrain, illumination conditions and dust-coverage; and benchmarking was performed to study the feasibility of each approach, their relative strengths and weaknesses, and robustness in the presence of adverse environmental conditions.

Detry, R.↗

Experiences in the Practice of Design of Experiments at NASA

Statistical design of experiments (DOE) has gained recognition within NASA by spurring innovation and efficiency, and it has demonstrated significant impact. Aerospace research and development benefits DOE techniques by accelerating learning, maximizing knowledge, ensuring strategic resource investment, and informing data-driven decisions. In practice, DOE relies on multidisciplinary collaboration to develop solution strategies that integrate statistical methods with subject-matter expertise to meet challenging research objectives. This presentation shares experiences in the practice of design of experiments at NASA in aeronautics, space exploration, and atmospheric science.

Peter A. Parker↗

Revisiting the Solar Research Cyberinfrastructure Needs: A White Paper of Findings and Recommendations

Solar and Heliosphere physics are areas of remarkable data-driven discoveries. Recent advances in high cadence, high-resolution multiwavelength observations, growing amounts of data from realistic modeling, and operational needs for uninterrupted science-quality data coverage generate the demand for a solar metadata standardization and overall healthy data infrastructure. This white paper is prepared as an effort of the working group “Uniform Semantics and Syntax of Solar Observations and Events” created within the “Towards Integration of Heliophysics Data, Modeling, and Analysis Tools” EarthCube Research Coordination Network (@HDMIEC RCN), with primary objectives to discuss current advances and identify future needs for the solar research cyberinfrastructure. The white paper summarizes presentations and discussions held during the special working group session at the EarthCube Annual Meeting on June 19th, 2020, as well as community contribution gathered during a series of preceding workshops and subsequent RCN working group sessions. The authors provide examples of the current standing of the solar research cyberinfrastructure, and describe the problems related to current data handling approaches. The list of the top-level recommendations agreed by the authors of the current white paper is presented at the beginning of the paper.

SMD↗

Sustainable Aviation Operations and the Role of Information Technology and Data Science: Background, Current Status and Future Directions

This paper reviews the achievements of the international community towards environmentally friendly aviation operations, also referred to as Sustainable Aviation Operations in the last 25 years and the aspirations and goals to limit the impact of aviation and climate in the future. The framework for achieving global progress is provided by the International Civil Aviation Organization. NASA and FAA supported research and development to advance ATM concepts, and implemented the technology, concepts, and procedures that were responsible for creating fuel efficient flights. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. Future developments in aviation operations require new concepts, procedure, modeling, and analysis techniques. There is an increasing interest in applying methods based on Machine Learning Techniques to problems in Air Traffic Management. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and the availability of a rich historical database provide opportunities to exploit the richness of data-driven methods. The promises and challenges in applying Machine Learning Techniques to Air Traffic Management are discussed in the paper along with the testing and trustworthiness required for adoption of the techniques in operations.

Sustainable Aviation, Data Science, Machine Learni↗

IMPACT - Information for the Development of Human Health and Performance Systems

IMPACT is a suite of computational and systems engineering tools. Currently being developed for NASA Exploration Medical Capabilities (ExMC) and could be extended to other Human Research Program (HRP) elements. The purpose is to provide a data-driven means to inform human health and performance risk mitigation during exploration missions considering resource constraints in the medical system. IMPACT enables systematic trade studies to evaluate options. IMPACT uses Probabilistic Risk Assessment (PRA) as a systematic methodology to evaluate risks. IMPACT is currently under development to inform decision-makers the risks involved with trade-options that have likelihoods and consequences backed by scientific research in a medical evidence library. IMPACT supports multi-segment analysis to determine probabilities and risks in different mission segments. There are currently 122 medical conditions and 900 resources in the IMPACT Medical Database. IMPACT modeling is using mission segments and multiple vehicles for the NASA Artemis campaign and could support future complex missions in deep space.

IMPACT PRA↗

Scaling a Smart Sub-metered Electrical Data Solution for the Center

Currently, there is no single Center-level point solution that enables facility systems data to be collected and stored from buildings on-site, on-demand, and through a faceted search capability that would enable the cultivation of deeper insights into incipient facility-related issues that would otherwise be overlooked. Such insights would represent a powerful and very valuable tool to aid in data-driven decision making and informing actionable steps for remediation and resolution of these problems. Access to electrically sub-metered data via hardware upgrades is in the process of being restored for Building N232, Sustainability Base, but this is only the first step. The native software capabilities that were originally part of a proposed Agencywide Smart Center initiative project can unlock the ability to access actionable insights. This work would also nicely complement many initiatives that are actively being pursued, under consideration, or pending submission to this same call at the Center. This project is poised to support an NRSAA with Verdigris Technologies, as well as in support of an Agencywide Smart Center initiative. Monitoring electrical consumption usage at the most granular level in Bldg. N232 will be instrumental in accurately quantifying potential utility costs and investments, and monitoring usage trends, as it is being converted into hoteling spaces, to help. This will also help with determining how the capability can scale more broadly across the Center and Agency at large.

Rodney Alexander Martin↗

Statistical Engineering

This webinar provides an overview of the International Statistical Engineering Association (ISEA), and it illustrates the practice of statistical engineering at NASA. ISEA was formed to promote the study of how successful data-based problem-solving methods are leveraged to realize innovative opportunities and solve problems sustainably. ISEA is comprised of statisticians, engineers, scientists, and other professionals that exchange ideas and experiences in the development and application of statistical engineering theories. ISEA is building the body of knowledge of the statistical engineering discipline with a particular focus on improving academic preparation for tackling complex problems. Over the past 15 years, the practice of statistical engineering has gained recognition within NASA by spurring innovation and efficiency, and it has demonstrated significant impact. Aerospace research and development benefits from an application-focused statistical engineering perspective to accelerate learning, maximize knowledge, ensure strategic resource investment, and inform data-driven decisions. The second portion of this presentation provides an overview of infusing statistical engineering at NASA through pioneering case studies in aeronautics, space exploration, and atmospheric science.

Peter A Parker↗

Albuquerque Urban Development: Enhancing Urban Cooling Interventions by Modeling Urban Forestry through NASA Earth Observations in Albuquerque, New Mexico

The City of Albuquerque, New Mexico is experiencing increasing urban heat island(UHI) effects, which impact the health, safety, and comfort of the community. In partnership with the City of Albuquerque Department of Environmental Health, Department of Parks and Recreation, and Let’s Plant Albuquerque!, this project used satellite Earth observations from April 2016–August 2022 to model increases in tree canopy within the City of Albuquerque to help combat the urban heat island in the city’s warmer areas over the next decade. Using Landsat 8’s Thermal Infrared Sensor (TIRS) and the Ecosystem Spaceborne Thermal Radiometer Experiment on the International Space Station (ECOSTRESS), along with the Integrated Valuation of Ecosystem Services and Tradeoffs (InVEST) Urban Cooling and ENVI-Met models, the team modeled tree cover interventions and created land surface temperature maps. These outputs will help the city make data-driven decisions for their tree planting goal in a targeted approach.

Max Stewart↗

A Dynamic Landslide Hazard Monitoring Framework for the Lower Mekong Region

The Lower Mekong region is one of the most landslide-prone areas of the world. Despite the need for dynamic characterization of landslide hazard zones within the region, it is largely understudied for several reasons. Dynamic and integrated understanding of landslide processes requires landslide inventories across the region, which have not been available previously. Computational limitations also hamper regional landslide hazard assessment, including accessing and processing remotely sensed information. Finally, open-source software and modelling packages are required to address regional landslide hazard analysis. Leveraging an open-source data-driven global Landslide Hazard Assessment for Situational Awareness model framework, this study develops a region-specific dynamic landslide hazard system leveraging satellite-based Earth observation data to assess landslide hazards across the lower Mekong region. A set of landslide inventories were prepared from high-resolution optical imagery using advanced image-processing techniques. Several static and dynamic explanatory variables (i.e., rainfall, soil moisture, slope, relief, distance to roads, distance to faults, distance to rivers) were considered during the model development phase. An extreme gradient boosting decision tree model was trained for the monsoon period of 2015–2019 and the model was evaluated with independent inventory information for the 2020 monsoon period. The model performance demonstrated considerable skill using receiver operating characteristic curve statistics, with Area Under the Curve values exceeding 0.95. The model architecture was designed to use near-real-time data, and it can be implemented in a cloud computing environment (i.e., Google Cloud Platform) for the routine assessment of landslide hazards in the Lower Mekong region. This work was developed in collaboration with scientists at the Asian Disaster Preparedness Center as part of the NASA SERVIR Program’s Mekong hub. The goal of this work is to develop a suite of tools and services on accessible open-source platforms that support and enable stakeholder communities to better assess landslide hazard and exposure at local to regional scales for decision making and planning.

Nishan Kumar Biswas↗

Validation of Machine Learning Algorithms for Hyperspectral Inversion of Common Water Quality Indicators

The upcoming transition to a diverse suite hyperspectral airborne and orbiting optical sensors will provide an unprecedented opportunity to measure inland water quality characteristics at a fidelity not previously achievable. This presentation will assess prototype deep learning models trained on synthetic hyperspectral data and validated with collocated in-situ measurements. Synthesized data is becoming increasingly popular for use in data-driven approaches to complex problems, and can compliment real data to increase performance on complex and unusual phenomenon, reduce or test bias, and experiment to demonstrate explainability. We will present insights from hyperspectral inversions of Chlorophyl-a, Phycocyanin, and concentration of non-algal particles using selected orbiting and airborne sensors over diverse, optically complex aquatic scenarios. We analyze how various optical water types affect fidelity of results and where improvements can be made as we prototype for globally operational water quality algorithms which can be leveraged by upcoming hyperspectral missions such as the Surface Biology and Geology (SBG) mission.

Surface Biology and Geology (SBG)↗

Continuing Global SO2 Data Record from OMI and SNPP/OMPS to JPSS-1/NOAA-20/OMPS

Since 2004, the Ozone Monitoring Instrument (OMI) aboard NASA's Earth Observing System (EOS) Aura spacecraft has been providing global observations that help to constrain the sources, transport, and environmental impacts of anthropogonic and volcanic SO2. The OMI SO2 data record is now being continued with the NASA/NOAA Suomi National Polar-orbiting Partnership (SNPP)/Ozone Mapping and Profiler Suite (OMPS) launched in 2011. Both OMI and SNPP/OMPS SO2 products are produced with the Goddard principal component analysis (PCA) based spectral fitting algorithm. This data-driven technique inherently accounts for various instrumental factors and geophysical interferences, leading to high-quality, consistent SO2 retrievals between OMI and SNPP/OMPS, despite coarser spectral (~0.5 nm vs. ~1 nm) and spatial (13  24 km2 vs. 50  50 km2 at nadir) resolution for the latter. In this presentation, we describe our effort to continue the long-term SO2 climate data record using measurements from the Joint Polar Satellite System (JPSS)-1/NOAA-20 (N20)/OMPS. Launched in 2017, the N20/OMPS is a follow-on for SNPP/OMPS but features a spatial resolution (17  13 km2) that is comparable with OMI. We will discuss our progress implementing the PCA SO2 algorithm with N20/OMPS, especially algorithmic improvements to further reduce retrieval noise and bias for large volcanic eruptions. We will present examples for both continuously emitting sources (e.g., power plants in India and oil/gas fields in the Middle East) and volcanic eruptions (e.g., Raikoke in 2019). We will also compare N20/OMPS SO2 retrievals with OMI and SNPP/OMPS, as well as other instruments such as the ESA Copernicus Sentinel-5 Precursor (S5P)/TROPOspheric Monitoring Instrument (TROPOMI). To assess the ability of N20/OMPS to monitor and quantify SO2 sources, we will run the level 2 retrievals through a top-down emission algorithm to estimate the SO2 emission strengths for a number of point sources. Finally, we will outline our plan for further algorithm refinement and public data release.

SO2↗

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

The Role in the Virtual Astronomical Observatory in the Era of Massive Data Sets

The Virtual Observatory (VO) is realizing global electronic integration of astronomy data. One of the long-term goals of the U.S. VO project, the Virtual Astronomical Observatory (VAO), is development of services and protocols that respond to the growing size and complexity of astronomy data sets. This paper describes how VAO staff are active in such development efforts, especially in innovative strategies and techniques that recognize the limited operating budgets likely available to astronomers even as demand increases. The project has a program of professional outreach whereby new services and protocols are evaluated.

data-driven science↗