Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data-driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Natural Language Processing Analysis of Notices to Airmen for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized.

Natural Language Processing↗

Natural Language Processing (NLP) Analysis of NOTAMs for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized. Video is an mp4 download, with a play time of 9 min 35 secs.

Natural Language Processing↗

Wildfire Emergency Response Hazard Extraction and Analysis of Trends (HEAT) through Natural Language Processing and Time Series

A methodology for Hazard Extraction and Analysis of Trends (HEAT) is proposed and conducted on a data set of wildfire incident response forms, known as ICS-209-PLUS.The HEAT processes: (1) extract a set of hazards from a data set, (2) calculate hazard-relevant metrics in a primary analysis, (3) analyze trends over time in metrics using timeseries, and (4) examine potential explanations for metric trends using a secondary analysis. Hazards are extracted from narrative data in the ICS-209-PLUS based on a framework previously developed by the authors, using natural language processing. Metrics examined for each hazard include operational time to occurrence, rate of occurrence, frequency, and severity. Primary results include a taxonomy of hazards present in the data set with relevant quantitative metrics. The most frequent hazards identified are environmental and include hazardous terrain. Most hazards occur on average between 35-55% containment. Incidents with hazards tend to have a higher average severity score when compared to the average score for all incidents. Time series of the metrics and relevant predictors, including fire characteristics, fire intensity, and operations, are created to facilitate further analysis. Secondary results used to determine which factors best predict hazard frequency include a correlation matrix and regression analysis. These findings are relevant to safety for current, as well as emerging wildfire operations, and are an exploratory first step in developing historical data-driven risk assessment models.

Sequoia R. Andrade↗

Systems Health Management and Prognostics Approaches for Electric Aircrafts

As more and more electric vehicles emerge in our daily operation progressively, a very critical challenge lies in the prediction of remaining driving flying time/distance for the flying vehicles. This information is important, particularly in the case of auto vehicles, because such vehicles can become self-aware, autonomously compute its own capabilities, and identify how to best plan and successfully complete vehicular missions safely. In case of electric aircrafts, computing the remaining flying time is also safety-critical, since an aircraft that runs out of power (battery charge) while in the air will eventually lose control leading to catastrophe. To facilitate and solve the prediction problem, awareness of the current health state of the system is key, since it is necessary to perform condition-based predictions. To accurately predict the future state of any system, it is required to possess knowledge of its current health state and future operational conditions. Latest achievements of data-driven algorithms in regression of complex nonlinear functions and classification tasks have generated a growing interest in artificial intelligence for industrial applications. Complex multi-physics models as well as digital twins, once purely built on physics and corresponding simplified lumped parameter iterations, can now benefit from machine learning algorithms to mitigate the lack of understanding of some complex behavior. Given models of the current and future system behavior, a general approach of model-based prognostics can solve the prediction problem and further decision-making. A systematic prediction framework is implemented to identify all possible sources of uncertainty, quantify each of them individually, and mathematically estimate their combined effect on the system-level quantity of interest, in this case, the remaining flying time/distance of the unmanned aircraft. Note - This presentation contains all previously published information.

Systems Health Managent↗

Aircraft Engine Run-to-Failure Dataset Under Real Flight Conditions for Prognostics and Diagnostics

A key enabler of intelligent maintenance systems is the ability to predict the remaining useful lifetime (RUL) of its components, i.e., prognostics. The development of data-driven prognostics models requires datasets with run-to-failure trajectories. However, large representative run-to-failure datasets are often unavailable in real applications because failures are rare in many safety-critical systems. To foster the development of prognostics methods, we develop a new realistic dataset of run-to-failure trajectories for a fleet of aircraft engines under real flight conditions. The dataset was generated with the Commercial Modular Aero-Propulsion System Simulation (CMAPSS) model developed at NASA. The damage propagation modelling used in this dataset builds on the modelling strategy from previous work and incorporates two new levels of fidelity. First, it considers real flight conditions as recorded on board of a commercial jet. Second, it extends the degradation modelling by relating the degradation process to its operation history. This dataset also provides the health respectively fault class. Therefore, besides its applicability to prognostics problems, the dataset can be used for fault diagnostics.

CMAPPS↗

Usage-based Lifing of Lithium-Ion Battery with HybridPhysics-Informed Neural Networks

Lithium-ion batteries are commonly used to power unmanned aircraft vehicles (UAVs).The ability to model and forecast the remaining useful life of these batteries enables UAV reliability assurance. Building accurate models for battery state of charge and state of health based on first principles is challenging due to the complex electrochemistry that governs battery operations and computational complexity required to solve them. Therefore, reduced order models are often used due to their ability to capture the overall battery discharge. Un-fortunately, these simplifications lead to residual discrepancy between model predictions and observed data. In this paper, we present a hybrid modeling approach merging reduced-order models and neural networks. In this approach, while most of the input-output relationship is captured by Nernst and Butler-Volmer equations, data-driven kernels reduce the gap between predictions and observations. We validate our approach using data publicly available through the NASA Prognostics Center of Excellence repository. Results showed that our hybrid battery prognosis model can be successfully calibrated, even with a limited number of observations.

Lithium-ion Battery↗

Arm DevSummit Keynote: Environment for Data Engineering in Virtual Reality: Ethical Considerations

Currently, the US Government is going through a large-scale data Transformation effort, where the GSA playbook is guiding all agencies to make their processes data-driven with the help of emerging technologies. To address concerns about algorithm sharing, AI adoption, and vendor lock-in, the NASA Langley Research Center Digital Transformation Group has developed the Environment for Data Engineering in Virtual Reality (EnDEVR), a data science ecosystem that allows users to command and investigate customizable data analyses from a VR environment. The system has been evaluated in the Oculus Rift S and Quest environments, two popular VR systems powered by the ARM architecture. Current and future development will employ several AI capabilities to guide research and automation within the environment and to facilitate algorithm sharing. In this talk, we will discuss the results of an initial ethical investigation and recommended considerations for the use of AI within the system.

Artificial Intelligence↗

Model-Independent Time-Delay Interferometry Based on Principal Component Analysis

With a laser interferometric gravitational-wave detector in separate free flying spacecraft, the only way to achieve detection is to mitigate the dominant noise arising from the frequency fluctuations of the lasers via postprocessing. The noise can be effectively filtered out on the ground through a specific technique called time-delay interferometry (TDI), which relies on the measurements of time-delays between spacecraft and careful modeling of how laser noise enters the interferometric data. Recently, this technique has been recast into a matrix-based formalism by several authors, offering a different perspective on TDI, particularly by relating it to principal component analysis (PCA). In this work, we demonstrate that we can cancel laser frequency noise by directly applying PCA to a set of shifted data samples, without any prior knowledge of the relationship between single-link measurements and noise, nor time-delays. We show that this fully data-driven algorithm achieves a gravitational-wave sensitivity similar to classic TDI.

Quentin Baghi↗

Informing the Space Launch System Booster Separation Initial CFD Run Matrix with Observed Parametric Sensitivity

It currently requires significant computational cost to simulate the flow physics of the booster separation event on the Space Launch System. This comes from the large parametric space in which the event occurs, as pre-separation flight conditions and separated booster core-relative trajectories can vary. Functionally removing the risk of core-booster collision mandates careful assessment of the fluid dynamics in terms of several trajectory parameters. However, simulating the entire trajectory envelope is computationally intractable given the high parametric dimension. In order to reduce the uncertainty in the resulting low-parametric-resolution aerodynamic booster separation database, a data-driven approach was developed to select which breakpoints should be studied by simulations and experiments and which should be relegated to a regression-based interpolation procedure. This technique works by simulating cases where the flow physics are most sensitive to changes in the parameters and leaving the less parametrically sensitive regions for interpolation. The result is a booster separation run matrix whose computational cost is comparable to that of previous database generations but has lower interpolation errors.

SLS↗

Secondary Science Teachers’ Implementation of a Curricular Intervention When Teaching With Global Climate Models

In the past decade, emphasis on promoting “climate literacy” in K-16 science classrooms has increased. Teachers play a critical role in cultivating these opportunities, especially in secondary science classrooms. However, most prior climate education research has focused on students and student learning; little is known about how teachers implement climate-focused curricular interventions. Here, we report findings from a concurrent mixed methods, multiple-case study of four secondary science teachers’ implementation of a new, NGSS-aligned, model-centric climate curriculum module grounded in the use of a data-driven, computer-based climate modeling tool—Easy Global Climate Model (EzGCM). We employ multiple data sources, including video-recorded classroom observations, interviews, and instructional artifacts, and both qualitative and quantitative analyses, to investigate how teachers implemented the curriculum. Findings show that, overall, teachers implemented the curriculum in ways that were less model-centric than designed, placing greater emphasis on EzGCM itself rather than using the model to investigate Earth’s changing climate. Additionally, we present detailed single-case studies of each participant teacher that highlight differences in teachers’ implementation of the curriculum module and their reasoning for making observed instructional decisions. This research sheds light on the design of secondary science learning environments by illustrating the varied ways teachers implement a climate-focused curriculum to support students’ developing climate literacy. This has important implications for the design of climate-focused curriculum and supports for teachers.

Secondary science teaching↗

Upper Limits on the Isotropic Gravitational-Wave Background from Advanced LIGO and Advanced Virgo's Third Observing Run

We report results of a search for an isotropic gravitational-wave background (GWB) using data from Advanced LIGO's and Advanced Virgo's third observing run (O3) combined with upper limits from the earlier O1 and O2 runs. Unlike in previous observing runs in the advanced detector era, we include Virgo in the search for the GWB. The results of the search are consistent with uncorrelated noise, and therefore we place upper limits on the strength of the GWB. We find that the dimensionless energy density Ω(sub GW) ≤ 5.8 × 10(exp -9) at the 95% credible level for a at (frequency-independent) GWB, using a prior which is uniform in the log of the strength of the GWB, with 99% of the sensitivity coming from the band 20-76.6 Hz; Ω(sub GW)(f) ≤ 3.4 × 10(exp -9) at 25 Hz for a power-law GWB with a spectral index of 2/3 (consistent with expectations for compact binary coalescences), in the band 20-90.6 Hz; and Ω(sub GW)(f) ≤ 3.9 × 10(exp -10) at 25 Hz for a spectral index of 3, in the band 20-291.6 Hz. These upper limits improve over our previous results by a factor of 6.0 for a at GWB, 8.8 for a spectral index of 2/3, and 13.1 for a spectral index of 3. We also search for a GWB arising from scalar and vector modes, which are predicted by alternative theories of gravity; we do not find evidence of these, and place upper limits on the strength of GWBs with these polarizations. We demonstrate that there is no evidence of correlated noise of magnetic origin by performing a Bayesian analysis that allows for the presence of both a GWB and an effective magnetic background arising from geophysical Schumann resonances. We compare our upper limits to a fiducial model for the GWB from the merger of compact binaries, updating the model to use the most recent data-driven population inference from the systems detected during O3a. Finally, we combine our results with observations of individual mergers and show that, at design sensitivity, this joint approach may yield stronger constraints on the merger rate of binary black holes at z ≳ 2 than can be achieved with individually resolved mergers alone.

Ryan Abbott↗

Mapping Global Forest Age from Forest Inventories, Biomass and Climate Data

Forest age can determine the capacity of a forest to uptake carbon from the atmosphere. However, a lack of global diagnostics that reflect the forest stage and associated disturbance regimes hampers the quantification of age-related differences in forest carbon dynamics. This study provides a new global distribution of forest age circa 2010, estimated using a machine learning approach trained with more than 40 000 plots using forest inventory, biomass and climate data. First, an evaluation against the plot-level measurements of forest age reveals that the data-driven method has a relatively good predictive capacity of classifying old-growth vs. non-old-growth (precision = 0.81 and 0.99 for old-growth and non-old-growth, respectively) forests and estimating corresponding forest age estimates (NSE = 0.6 – Nash–Sutcliffe efficiency – and RMSE = 50 years – root-mean-square error). However, there are systematic biases of overestimation in young- and underestimation in old-forest stands, respectively. Globally, we find a large variability in forest age with the old-growth forests in the tropical regions of Amazon and Congo, young forests in China, and intermediate stands in Europe. Furthermore, we find that the regions with high rates of deforestation or forest degradation (e.g. the arc of deforestation in the Amazon) are composed mainly of younger stands. Assessment of forest age in the climate space shows that the old forests are either in cold and dry regions or warm and wet regions, while young–intermediate forests span a large climatic gradient. Finally, comparing the presented forest age estimates with a series of regional products reveals differences rooted in different approaches and different in situ observations and global-scale products. Despite showing robustness in cross-validation results, additional methodological insights on further developments should as much as possible harmonize data across the different approaches. The forest age dataset presented here provides additional insights into the global distribution of forest age to better understand the global dynamics in the forest water and carbon cycles. The forest age datasets are openly available at https://doi.org/10.17871/ForestAgeBGI.2021 (Besnard et al., 2021).

Simon Besnard↗

Advancing Open Science through Public-Private Partnerships

Rapid technology developments are changing the way data-driven research is performed within the science community. With the emergence of cloud computing, this has quickly become a viable approach for enabling “science at scale”. Researchers are no longer hindered by obstacles of data management and data wrangling, allowing them to quickly discover, access and perform analysis on extremely large datasets. Infrastructures that move data out of institutional silos and into a computational platform, will ensure that data and tools are accessible to all users. NASA’s Interagency Implementation and Advanced Concepts Team (IMPACT) seeks to address these rapid technology developments by establishing Space Act Agreements with selected partners from the public-private sector working in the area of cloud computing. These agreements aim to explore new opportunities with commercial cloud providers to accelerate open science and enable discovery, access and use of data sets on the cloud. In addition, they will also help establish training workshops for the science community to help researchers utilize the cloud for science. In this talk, we will present an overview of current and new partnerships we are developing to support open science and open data initiatives.

Elizabeth Fancher↗

Natural Language Processing Methods for Air Traffic Management Text and Speech Data

This presentation discusses two efforts of the NARI AI/ML Intern team during the Fall 2021 OSTEM Internship term. For Letters of Agreement (LoA), we have studied how LoAs are structured and explored the question ‘What is an LoA constraint?’ To do this, our approach is data-driven, iterative, and assisted by machine learning when available. In this presentation, we will walk through our tasks of manually scanning through documents, performing a preliminary entity labelling task, and our unsupervised analysis on LoA procedures sections. After this research phase, we define the smallest constraint unit in an LoA, and start to perform entity extraction. Looking towards constraint extraction, we are also exploring the use of a one-class support vector machine (OneClassSVM) model to identify patterns within the data. The second effort of our team this term is focused on Air Traffic Control System Command Center (ATCSCC) advisory meetings, and the subsequent advisory documents that get published from their content. These advisory documents are important to give readily accessible summaries of daily operations, so that data centers, airline officials, and other stakeholders can easily understand the context of these meetings in real time. In applying machine learning to this scenario, two natural language processing tasks are used. First is developing machine learning models to convert the meeting speech data into text. With this text, use of extractive and abstractive text summarization models are used to automatically generate preliminary versions of the advisory documents.

Natural Language Processing↗

Machine Vision based Sample-Tube Localization for Mars Sample Return

A potential Mars Sample Return (MSR) architecture is being jointly studied by NASA and ESA. As currently envisioned, the MSR campaign consists of a series of 3 missions: sample cache, fetch and return to Earth. In this paper, we focus on the fetch part of the MSR, and more specifically the problem of autonomously detecting and localizing sample tubes deposited on the Martian surface. Towards this end, we study two machine-vision based approaches: First, a geometrydriven approach based on template matching that uses hardcoded filters and a 3D shape model of the tube; and second, a data-driven approach based on convolutional neural networks (CNNs) and learned features. Furthermore, we present a large benchmark dataset of sample-tube images, collected in representative outdoor environments and annotated with ground truth segmentation masks and locations. The dataset was acquired systematically across different terrain, illumination conditions and dust-coverage; and benchmarking was performed to study the feasibility of each approach, their relative strengths and weaknesses, and robustness in the presence of adverse environmental conditions.

Detry, R.↗

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗