Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data-driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Landslide Susceptibility and Runout Assessment for the Arun Hydropower Cascade in Nepal Based on Multitemporal Landslide Inventory Developed Using Planet Imagery

The transboundary Pumqu/Arun River basin spreads across Nepal and Tibet. Nearly 95% of the basin lies in Tibet through which the Pumqu River flows forming the Arun River once it enters Nepal. The Arun River valley has about 3,163 MW in five large hydropower projects undergoing construction or planned for the future. Rainfall and earthquake-induced landslides, landslide dammed lakes and landslide-induced glacial lake outburst floods pose major risks to smooth operation of these projects. To safeguard upcoming hydropower projects, areas susceptible to landslides in the Arun River valley must be identified. We generated a multitemporal landslide inventory (2010-2020), modeled landslide susceptibility and runouts around the hydropower sites using Earth Observation (EO) techniques. We used high-resolution satellite imagery from Planet and open-source tool to establish a detailed and comprehensive multitemporal landslide inventory for the basin. The rigorous quality-controlled inventory represents yearly record of landslides from 2010 to 2020. To our knowledge this is one of the most comprehensive landslide inventories for the basin with yearly records of landslides.A data-driven approach was used to map areas susceptible to landslides within the Pumqu/Arun River basin. The multitemporal landslide inventory combined with other readily available EO-based variables were used to create a landslide susceptibility map. The highly susceptible areas delineated from the map was further used to characterize runouts near the hydropower sites. The susceptibility and runout analysis provide a valuable initial estimate of where landslides are likely to trigger and where the mobilized material is likely to end up.

Pukar Amatya↗

Statistical Engineering

This webinar provides an overview of the International Statistical Engineering Association (ISEA), and it illustrates the practice of statistical engineering at NASA. ISEA was formed to promote the study of how successful data-based problem-solving methods are leveraged to realize innovative opportunities and solve problems sustainably. ISEA is comprised of statisticians, engineers, scientists, and other professionals that exchange ideas and experiences in the development and application of statistical engineering theories. ISEA is building the body of knowledge of the statistical engineering discipline with a particular focus on improving academic preparation for tackling complex problems. Over the past 15 years, the practice of statistical engineering has gained recognition within NASA by spurring innovation and efficiency, and it has demonstrated significant impact. Aerospace research and development benefits from an application-focused statistical engineering perspective to accelerate learning, maximize knowledge, ensure strategic resource investment, and inform data-driven decisions. The second portion of this presentation provides an overview of infusing statistical engineering at NASA through pioneering case studies in aeronautics, space exploration, and atmospheric science.

Peter A Parker↗

A New Machine Learning Based Analysis for Improving Satellite Retrieved Atmospheric Composition Data: OMI SO2 as an Example

Despite recent progress, satellite retrievals of anthropogenic SO2 still suffer from relatively low signal-tonoise ratios. In this study, we demonstrate a new machine learning data analysis method to improve the quality of satellite SO2 products. In the absence of large ground-truth datasets for SO2, we start from SO2 slant column densities (SCDs) retrieved from the Ozone Monitoring Instrument (OMI) using a data-driven, physically based algorithm and calculate the ratio between the SCD and the root mean square (rms) of the fitting residuals for each pixel. To build the training data, we select presumably clean pixels with small SCD / rms ratios (SRRs) and set their target SCDs to zero. For polluted pixels with relatively large SRRs, we set the target to the original retrieved SCDs. We then train neural networks (NNs) to reproduce the target SCDs using predictors including SRRs for individual pixels, solar zenith, viewing zenith and phase angles, scene reflectivity, and O3 column amounts, as well as the monthly mean SRRs. For data analysis, we employ two NNs: (1) one trained daily to produce analyzed SO2 SCDs for polluted pixels each day and (2) the other trained once every month to produce analyzed SCDs for less polluted pixels for the entire month. Test results for 2005 show that our method can significantly reduce noise and artifacts over background regions. Over polluted areas, the monthly mean NN-analyzed and original SCDs generally agree to within ±15 %, indicating that our method can retain SO2 signals in the original retrievals except for large volcanic eruptions. This is further confirmed by running both the NN-analyzed and original SCDs through a topdown emission algorithm to estimate the annual SO2 emissions for ∼ 500 anthropogenic sources, with the two datasets yielding similar results. We also explore two alternative approaches to the NN-based analysis method. In one, we employ a simple linear interpolation model to analyze the original SCD retrievals. In the other, we develop a PCA–NN algorithm that uses OMI measured radiances, transformed and dimension-reduced with a principal component analysis (PCA) technique, as inputs to NNs for SO2 SCD retrievals. While the linear model and the PCA–NN algorithm can reduce retrieval noise, they both underestimate SO2 over polluted areas. Overall, the results presented here demonstrate that our new data analysis method can significantly improve the quality of existing OMI SO2 retrievals. The method can potentially be adapted for other sensors and/or species and enhance the value of satellite data in air quality research and applications.

Can Li↗

Albuquerque Urban Development: Enhancing Urban Cooling Interventions by Modeling Urban Forestry through NASA Earth Observations in Albuquerque, New Mexico

The City of Albuquerque, New Mexico is experiencing increasing urban heat island(UHI) effects, which impact the health, safety, and comfort of the community. In partnership with the City of Albuquerque Department of Environmental Health, Department of Parks and Recreation, and Let’s Plant Albuquerque!, this project used satellite Earth observations from April 2016–August 2022 to model increases in tree canopy within the City of Albuquerque to help combat the urban heat island in the city’s warmer areas over the next decade. Using Landsat 8’s Thermal Infrared Sensor (TIRS) and the Ecosystem Spaceborne Thermal Radiometer Experiment on the International Space Station (ECOSTRESS), along with the Integrated Valuation of Ecosystem Services and Tradeoffs (InVEST) Urban Cooling and ENVI-Met models, the team modeled tree cover interventions and created land surface temperature maps. These outputs will help the city make data-driven decisions for their tree planting goal in a targeted approach.

Max Stewart↗

NASA Symposium on Turbulence Modeling: Roadblocks, and the Potential for Machine Learning

A three-day symposium sponsored by NASA was held in July 2022 in Suffolk, Virginia on the subject of Turbulence Modeling: Roadblocks, and the Potential for Machine Learning. This meeting brought together over 80 experts from academia, government, and industry to discuss critical issues for Reynolds-averaged Navier-Stokes turbulence and transition models, as well as to evaluate the results from a collaborative testing challenge based on data-driven methods and machine learning technology. This report puts this symposium in context with an earlier similar meeting and summarizes many of the questions, discussions, and conclusions that arose from it. Next steps are suggested.

machine learning↗

Natural Language Processing Techniques for Intelligent Knowledge Management of Safety Reports

Safety, failure, and incident reports are common artifacts across various domains, including aviation and wildfire response. These reports are often mandatory to submit, resulting in the culmination of large repositories of text-based documents. Simultaneously, these reports and corresponding repositories are often only manually analyzed and queried by users via out-of-date search engines. As a consequence, we have been developing the Manager for Intelligent Knowledge Access (MIKA) toolkit, which uses natural language processing to improve information access and reuse. In this presentation, we discuss natural language processing techniques for knowledge discovery and apply these methods to a repository of aerial wildfire mishap reports. Two methods are used for knowledge discovery: topic modeling and named-entity recognition. We use topic modeling to identify hazards and perform a trend analysis to produce a data-driven risk matrix. A custom named-entity recognition model, build from fine tuning a pre-trained language model, is used to identify failure modes, failure causes, failure effects, control processes, and recommendations to aid in failure modes and effects analysis (FMEA). Throughout the presentation, we discuss and apply natural language processing techniques to better leverage the vast amount of information contained in report repositories.

Machine learning↗

A Dynamic Landslide Hazard Monitoring Framework for the Lower Mekong Region

The Lower Mekong region is one of the most landslide-prone areas of the world. Despite the need for dynamic characterization of landslide hazard zones within the region, it is largely understudied for several reasons. Dynamic and integrated understanding of landslide processes requires landslide inventories across the region, which have not been available previously. Computational limitations also hamper regional landslide hazard assessment, including accessing and processing remotely sensed information. Finally, open-source software and modelling packages are required to address regional landslide hazard analysis. Leveraging an open-source data-driven global Landslide Hazard Assessment for Situational Awareness model framework, this study develops a region-specific dynamic landslide hazard system leveraging satellite-based Earth observation data to assess landslide hazards across the lower Mekong region. A set of landslide inventories were prepared from high-resolution optical imagery using advanced image-processing techniques. Several static and dynamic explanatory variables (i.e., rainfall, soil moisture, slope, relief, distance to roads, distance to faults, distance to rivers) were considered during the model development phase. An extreme gradient boosting decision tree model was trained for the monsoon period of 2015–2019 and the model was evaluated with independent inventory information for the 2020 monsoon period. The model performance demonstrated considerable skill using receiver operating characteristic curve statistics, with Area Under the Curve values exceeding 0.95. The model architecture was designed to use near-real-time data, and it can be implemented in a cloud computing environment (i.e., Google Cloud Platform) for the routine assessment of landslide hazards in the Lower Mekong region. This work was developed in collaboration with scientists at the Asian Disaster Preparedness Center as part of the NASA SERVIR Program’s Mekong hub. The goal of this work is to develop a suite of tools and services on accessible open-source platforms that support and enable stakeholder communities to better assess landslide hazard and exposure at local to regional scales for decision making and planning.

Nishan Kumar Biswas↗

A Robust Machine Learning Schema for Developing, Maintaining, and Disseminating Machine Learning Models

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of ML models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based modeling of material behavior at various length scales and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using ML techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus, effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train ML models and the defining model parameters and architectures within the Granta MI Platform. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in the prediction of material behavior, while following outlined best practices for effective data management. An effective schema for ML data and models can help prevent the recreation of virtual/real training data and surrogate models, help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Brandon L. Hearley↗

Human Contribution to Safety: Human Performance is Not Just About Error

- When the only data that are available are about human failure, then data-driven designs only consider that humans fail. - Designs intended to "protect" the system from "error-prone" humans can design-out the capability for humans to effectively intervene/adapt, which is a far more common behavior.

Jon Holbrook↗

Validation of Machine Learning Algorithms for Hyperspectral Inversion of Common Water Quality Indicators

The upcoming transition to a diverse suite hyperspectral airborne and orbiting optical sensors will provide an unprecedented opportunity to measure inland water quality characteristics at a fidelity not previously achievable. This presentation will assess prototype deep learning models trained on synthetic hyperspectral data and validated with collocated in-situ measurements. Synthesized data is becoming increasingly popular for use in data-driven approaches to complex problems, and can compliment real data to increase performance on complex and unusual phenomenon, reduce or test bias, and experiment to demonstrate explainability. We will present insights from hyperspectral inversions of Chlorophyl-a, Phycocyanin, and concentration of non-algal particles using selected orbiting and airborne sensors over diverse, optically complex aquatic scenarios. We analyze how various optical water types affect fidelity of results and where improvements can be made as we prototype for globally operational water quality algorithms which can be leveraged by upcoming hyperspectral missions such as the Surface Biology and Geology (SBG) mission.

Surface Biology and Geology (SBG)↗

Continuing Global SO2 Data Record from OMI and SNPP/OMPS to JPSS-1/NOAA-20/OMPS

Since 2004, the Ozone Monitoring Instrument (OMI) aboard NASA's Earth Observing System (EOS) Aura spacecraft has been providing global observations that help to constrain the sources, transport, and environmental impacts of anthropogonic and volcanic SO2. The OMI SO2 data record is now being continued with the NASA/NOAA Suomi National Polar-orbiting Partnership (SNPP)/Ozone Mapping and Profiler Suite (OMPS) launched in 2011. Both OMI and SNPP/OMPS SO2 products are produced with the Goddard principal component analysis (PCA) based spectral fitting algorithm. This data-driven technique inherently accounts for various instrumental factors and geophysical interferences, leading to high-quality, consistent SO2 retrievals between OMI and SNPP/OMPS, despite coarser spectral (~0.5 nm vs. ~1 nm) and spatial (13  24 km2 vs. 50  50 km2 at nadir) resolution for the latter. In this presentation, we describe our effort to continue the long-term SO2 climate data record using measurements from the Joint Polar Satellite System (JPSS)-1/NOAA-20 (N20)/OMPS. Launched in 2017, the N20/OMPS is a follow-on for SNPP/OMPS but features a spatial resolution (17  13 km2) that is comparable with OMI. We will discuss our progress implementing the PCA SO2 algorithm with N20/OMPS, especially algorithmic improvements to further reduce retrieval noise and bias for large volcanic eruptions. We will present examples for both continuously emitting sources (e.g., power plants in India and oil/gas fields in the Middle East) and volcanic eruptions (e.g., Raikoke in 2019). We will also compare N20/OMPS SO2 retrievals with OMI and SNPP/OMPS, as well as other instruments such as the ESA Copernicus Sentinel-5 Precursor (S5P)/TROPOspheric Monitoring Instrument (TROPOMI). To assess the ability of N20/OMPS to monitor and quantify SO2 sources, we will run the level 2 retrievals through a top-down emission algorithm to estimate the SO2 emission strengths for a number of point sources. Finally, we will outline our plan for further algorithm refinement and public data release.

SO2↗

Intrinsic Dimensionality as a Metric for the Impact of Mission Design Parameters

High-resolution space-based spectral imaging of the Earth's surface delivers critical information for monitoring changes in the Earth system as well as resource management and utilization. Orbiting spectrometers are built according to multiple design parameters, including ground sampling distance (GSD), spectral resolution, temporal resolution, and signal-to-noise ratio. Different applications drive divergent instrument designs, so optimization for wide-reaching missions is complex. The Surface Biology and Geology component of NASA's Earth System Observatory addresses science questions and meets applications needs across diverse fields, including terrestrial and aquatic ecosystems, natural disasters, and the cryosphere. The algorithms required to generate the geophysical variables from the observed spectral imagery each have their own inherent dependencies and sensitivities, and weighting these objectively is challenging. Here, we introduce intrinsic dimensionality (ID), a measure of information content, as an applications-agnostic, data-driven metric to quantify performance sensitivity to various design parameters. ID is computed through the analysis of the eigenvalues of the image covariance matrix, and can be thought of as the number of significant principal components. This metric is extremely powerful for quantifying the information content in high-dimensional data, such as spectrally resolved radiances and their changes over space and time. We find that the ID decreases for coarser GSD, decreased spectral resolution and range, less frequent acquisitions, and lower signal-to-noise levels. This decrease in information content has implications for all derived products. ID is simple to compute, providing a single quantitative standard to evaluate combinations of design parameters, irrespective of higher-level algorithms, products, applications, or disciplines.

Intrinsic dimensionality↗

A Disaggregation Algorithm for the High Resolution Soil Moisture Product from the Upcoming NISAR Mission

The NASA-ISRO Synthetic Aperture Radar (NISAR) is in the developmental stage and is planned to launch in Jan 2024 with two different microwave frequency bands L-band (~1.25 GHz) and S-band (~3.20 GHz), respectively, to provide fine-scale observations at resolutions of 5 to 10 meters. NISAR mission will provide a very high-resolution (200m) soil moisture product globally with a temporal resolution of 6 days, using L-band SAR observations. A data-driven approach is developed for disaggregating the coarse resolution (9 km) soil moisture data to a very high-resolution (200 m) soil moisture product using fine-scale (~ 10 m) NISAR L-band observations. In this study, we used ALOS PALSAR-2 L-band SAR observations in place of expected NISAR L-band observations. The developed disaggregation approach was tested on two different locations of India and USA and showed that the proposed approach has a great potential to estimate soil moisture at a very high resolution of 200m with very low uncertainties (0.02 m3/m3 – 0.04 m3/m3).

Vanama, Venkat↗

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

The Role in the Virtual Astronomical Observatory in the Era of Massive Data Sets

The Virtual Observatory (VO) is realizing global electronic integration of astronomy data. One of the long-term goals of the U.S. VO project, the Virtual Astronomical Observatory (VAO), is development of services and protocols that respond to the growing size and complexity of astronomy data sets. This paper describes how VAO staff are active in such development efforts, especially in innovative strategies and techniques that recognize the limited operating budgets likely available to astronomers even as demand increases. The project has a program of professional outreach whereby new services and protocols are evaluated.

data-driven science↗

Formulative Input into Future NASA Aeronautics Planning

This presentation covers industry input received for future work in NASA Aeronautics over the next 5 years. It is intended to present areas of significant imput and to stimulate further discussion.

future aeronautics planning↗