Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data-driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A materials-informatics based study of solid electrolytes and protective coatings for Li batteries

All-solid-state batteries with Li metal anode can address the safety issues surrounding traditional Li-ion batteries as well as the demand for higher energy densities. However, the development of solid electrolytes and protective coatings simultaneously possessing high ionic conductivity and wide electrochemical stability has proven to be a challenge. Here, we present a data-driven approach to explore the Li compound space for promising solid electrolytes and coatings. This is accomplished through the generation of a large database of battery-related materials properties of Li compounds by computing Li+ migration barriers using bond-valence-based pair potentials, and stability windows using density functional theory energies. Using this database, we implement machine learning models that can accurately predict migration barriers and electrochemical stability windows for any new Li compound. Through feature engineering, we ensure that our models are both accurate and interpretable. We perform feature importance analysis on our models to highlight materials properties that can be tuned for future design of coatings/electrolytes. Our database and informatics approach provide a valuable tool for the rapid discovery of new solid-state battery chemistries.

Solid state batteries↗

Proprioceptive Inference for Dual-Arm Grasping of Bulky Objects Using RoboSimian

This work demonstrates dual-arm lifting of bulky objects based on inferred object properties (center of mass (COM) location, weight, and shape) using proprioception (i.e. force torque measurements). Data-driven Bayesian models describe these quantities, which enables subsequent behaviors to depend on confidence of the learned models. Experiments were conducted using the NASA Jet Propulsion Laboratory’s (JPL) RoboSimian to lift a variety of cumbersome objects ranging in mass from 7kg to 25kg. The position of a supporting second manipulator was determined using a particle set and heuristics that were derived from inferred object properties. The supporting manipulator decreased the initial manipulator’s load and distributed the wrench load more equitably across each manipulator, for each bulky object. Knowledge of the objects came from pure proprioception (i.e. without reliance on vision or other exteroceptive sensors) throughout the experiments.

Backes, Paul↗

Analysis of Nonlinear Shrinkage for the Bound Metal Deposition Manufacturing using Multi-scale Approach

We consider problem of nonlinear shrinkage of the metal part during bound metal deposition manufacturing on the ground and in zero-G. To analyze this problem we developed multi-scale physics-based approach that spans atomistic dynamics at the scale of nanoseconds and the full part shrinkage at the time scale of hours. Using this approach we estimated the key parameters of the problem including grain boundary width, coefficient of surface diffusion, initial redistribution of particles during debinding stage, micro-structure evolution from round particles to densely packed grains and corresponding change of the total and chemical free energy, and sintering stress. The introduced method was used to predict shrinkage at the level of two particles, filament cross-section, sub-model, and the whole green, brown, and metal parts. To further improve accuracy and reliability of the shrinkage predictions we propose concept of intelligent additive manufacturing of metal powders in space that combines the strengths of both physics-based and data-driven methods of analysis of AM.

bound metal deposition↗

The Global Methane Budget 2000–2017

Understanding and quantifying the global methane (CH4) budget is important for assessing realistic pathways to mitigate climate change. Atmospheric emissions and concentrations of CH4 continue to increase, making CH4 the second most important human-influenced greenhouse gas in terms of climate forcing, after carbon dioxide (CO2). The relative importance of CH4 compared to CO2 depends on its shorter atmospheric lifetime, stronger warming potential, and variations in atmospheric growth rate over the past decade, the causes of which are still debated. Two major challenges in reducing uncertainties in the atmospheric growth rate arise from the variety of geographically overlapping CH4 sources and from the destruction of CH4 by short-lived hydroxyl radicals (OH). To address these challenges, we have established a consortium of multidisciplinary scientists under the umbrella of the Global Carbon Project to synthesize and stimulate new research aimed at improving and regularly updating the global methane budget. Following Saunois et al. (2016), we present here the second version of the living review paper dedicated to the decadal methane budget, integrating results of top-down studies (atmospheric observations within an atmospheric inverse-modelling framework) and bottom-up estimates (including process-based models for estimating land surface emissions and atmospheric chemistry, inventories of anthropogenic emissions, and data-driven extrapolations). For the 2008–2017 decade, global methane emissions are estimated by atmospheric inversions (a top-down approach) to be 576 Tg CH4/yr (range 550–594, corresponding to the minimum and maximum estimates of the model ensemble). Of this total, 359 Tg CH4/yr or ∼ 60 % is attributed to anthropogenic sources, that is emissions caused by direct human activity (i.e. anthropogenic emissions; range 336–376 Tg CH4/yr or 50 %–65 %). The mean annual total emission for the new decade (2008–2017) is 29 Tg CH4/yr larger than our estimate for the previous decade (2000–2009), and 24 Tg CH4/yr larger than the one reported in the previous budget for 2003–2012 (Saunois et al., 2016). Since 2012, global CH4 emissions have been tracking the warmest scenarios assessed by the Intergovernmental Panel on Climate Change. Bottom-up methods suggest almost 30 % larger global emissions (737 Tg CH4/yr, range 594–881) than top-down inversion methods. Indeed, bottom-up estimates for natural sources such as natural wetlands, other inland water systems, and geological sources are higher than top-down estimates. The atmospheric constraints on the top-down budget suggest that at least some of these bottom-up emissions are overestimated. The latitudinal distribution of atmospheric observation-based emissions indicates a predominance of tropical emissions (∼ 65 % of the global budget, < 30° N) compared to mid-latitudes (∼ 30 %, 30–60° N) and high northern latitudes (∼ 4 %, 60–90° N). The most important source of uncertainty in the methane budget is attributable to natural emissions, especially those from wetlands and other inland waters. Some of our global source estimates are smaller than those in previously published budgets (Saunois et al., 2016; Kirschke et al., 2013). In particular wetland emissions are about 35 Tg CH4/yr lower due to improved partition wetlands and other inland waters. Emissions from geological sources and wild animals are also found to be smaller by 7 Tg CH4/yr by 8 Tg CH4/yr, respectively. However, the overall discrepancy between bottom-up and top-down estimates has been reduced by only 5 % compared to Saunois et al. (2016), due to a higher estimate of emissions from inland waters, highlighting the need for more detailed research on emissions factors. Priorities for improving the methane budget include (i) a global, high-resolution map of water-saturated soils and inundated areas emitting methane based on a robust classification of different types of emitting habitats; (ii) further development of process-based models for inland-water emissions; (iii) intensification of methane observations at local scales (e.g., FLUXNET-CH4 measurements) and urban-scale monitoring to constrain bottom-up land surface models, and at regional scales (surface networks and satellites) to constrain atmospheric inversions; (iv) improvements of transport models and the representation of photochemical sinks in top-down inversions; and (v) development of a 3D variational inversion system using isotopic and/or co-emitted species such as ethane to improve source partitioning.

methane budget↗

Locating the Isolator Shock-Train Leading Edge with Limited Pressure Information

Real-time detection and control of the isolator shock-train leading edge (STLE) is important to the performance of high-speed air-breathing engines, such as dual-mode scramjets. Typically, the STLE location is determined using wall static-pressure measurements, but there are often restrictions on the placement and overall number of the pressure transducers, reducing the viability and accuracy of such approaches. To address these issues, we introduce the adaptive pressure profile (APP) method for estimating the STLE location. This method does not require extensive prior characterization of the isolator or engine model. Instead, it uses real-time pressure measurements from a small number of transducers to adaptively learn the isolator pressure profile and subsequently uses this deduced profile to estimate the STLE location in a data-driven manner. The APP method works well in situations with sparse transducer placement. It produces accurate estimates when the STLE location is 1) not bounded by two or more transducers or 2) between two transducers that are several isolator duct heights apart. We demonstrate the efficacy of the APP method using simulations and experimental data from direct-connect isolator models. This validation shows that the APP method is accurate and robust for different flow regimes, transducer configurations, and model geometries.

Gregory J. Hunt↗

Predicting Arrival and Departure Runway Assignments with Machine Learning

Runway assignments at major airports are made by air traffic controllers subject to various constraints, and to achieve various objectives. In this research, we describe our efforts training machine learning (ML) models to predict both departure and arrival runway assignments using an entirely data-driven approach. This approach is compared to existing rule-based approaches developed in previous research using input from Subject Matter Experts. The models have features derived from various FAA data feeds, and leverage multiple machine learning algorithms. Results for models trained for nine major U.S. airports are described and compared to one another across various important dimensions. Particular attention was paid to developing a repeatable framework for training these models so the approach could be scaled to other airports, and to developing models that are useful in a real-time environment. In addition, the models were designed to be functional in a real-time environment to support NASA’s ATD-2 project, as part of an ML-powered shadow system to compare against the performance of the fielded system.

machine learning↗

Natural Language Processing (NLP) Analysis of NOTAMs for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized. Video is an mp4 download, with a play time of 9 min 35 secs.

Natural Language Processing↗

Wildfire Emergency Response Hazard Extraction and Analysis of Trends (HEAT) through Natural Language Processing and Time Series

A methodology for Hazard Extraction and Analysis of Trends (HEAT) is proposed and conducted on a data set of wildfire incident response forms, known as ICS-209-PLUS.The HEAT processes: (1) extract a set of hazards from a data set, (2) calculate hazard-relevant metrics in a primary analysis, (3) analyze trends over time in metrics using timeseries, and (4) examine potential explanations for metric trends using a secondary analysis. Hazards are extracted from narrative data in the ICS-209-PLUS based on a framework previously developed by the authors, using natural language processing. Metrics examined for each hazard include operational time to occurrence, rate of occurrence, frequency, and severity. Primary results include a taxonomy of hazards present in the data set with relevant quantitative metrics. The most frequent hazards identified are environmental and include hazardous terrain. Most hazards occur on average between 35-55% containment. Incidents with hazards tend to have a higher average severity score when compared to the average score for all incidents. Time series of the metrics and relevant predictors, including fire characteristics, fire intensity, and operations, are created to facilitate further analysis. Secondary results used to determine which factors best predict hazard frequency include a correlation matrix and regression analysis. These findings are relevant to safety for current, as well as emerging wildfire operations, and are an exploratory first step in developing historical data-driven risk assessment models.

Sequoia R. Andrade↗

Aircraft Engine Run-to-Failure Dataset Under Real Flight Conditions for Prognostics and Diagnostics

A key enabler of intelligent maintenance systems is the ability to predict the remaining useful lifetime (RUL) of its components, i.e., prognostics. The development of data-driven prognostics models requires datasets with run-to-failure trajectories. However, large representative run-to-failure datasets are often unavailable in real applications because failures are rare in many safety-critical systems. To foster the development of prognostics methods, we develop a new realistic dataset of run-to-failure trajectories for a fleet of aircraft engines under real flight conditions. The dataset was generated with the Commercial Modular Aero-Propulsion System Simulation (CMAPSS) model developed at NASA. The damage propagation modelling used in this dataset builds on the modelling strategy from previous work and incorporates two new levels of fidelity. First, it considers real flight conditions as recorded on board of a commercial jet. Second, it extends the degradation modelling by relating the degradation process to its operation history. This dataset also provides the health respectively fault class. Therefore, besides its applicability to prognostics problems, the dataset can be used for fault diagnostics.

CMAPPS↗

Secondary Science Teachers’ Implementation of a Curricular Intervention When Teaching With Global Climate Models

In the past decade, emphasis on promoting “climate literacy” in K-16 science classrooms has increased. Teachers play a critical role in cultivating these opportunities, especially in secondary science classrooms. However, most prior climate education research has focused on students and student learning; little is known about how teachers implement climate-focused curricular interventions. Here, we report findings from a concurrent mixed methods, multiple-case study of four secondary science teachers’ implementation of a new, NGSS-aligned, model-centric climate curriculum module grounded in the use of a data-driven, computer-based climate modeling tool—Easy Global Climate Model (EzGCM). We employ multiple data sources, including video-recorded classroom observations, interviews, and instructional artifacts, and both qualitative and quantitative analyses, to investigate how teachers implemented the curriculum. Findings show that, overall, teachers implemented the curriculum in ways that were less model-centric than designed, placing greater emphasis on EzGCM itself rather than using the model to investigate Earth’s changing climate. Additionally, we present detailed single-case studies of each participant teacher that highlight differences in teachers’ implementation of the curriculum module and their reasoning for making observed instructional decisions. This research sheds light on the design of secondary science learning environments by illustrating the varied ways teachers implement a climate-focused curriculum to support students’ developing climate literacy. This has important implications for the design of climate-focused curriculum and supports for teachers.

Secondary science teaching↗

Upper Limits on the Isotropic Gravitational-Wave Background from Advanced LIGO and Advanced Virgo's Third Observing Run

We report results of a search for an isotropic gravitational-wave background (GWB) using data from Advanced LIGO's and Advanced Virgo's third observing run (O3) combined with upper limits from the earlier O1 and O2 runs. Unlike in previous observing runs in the advanced detector era, we include Virgo in the search for the GWB. The results of the search are consistent with uncorrelated noise, and therefore we place upper limits on the strength of the GWB. We find that the dimensionless energy density Ω(sub GW) ≤ 5.8 × 10(exp -9) at the 95% credible level for a at (frequency-independent) GWB, using a prior which is uniform in the log of the strength of the GWB, with 99% of the sensitivity coming from the band 20-76.6 Hz; Ω(sub GW)(f) ≤ 3.4 × 10(exp -9) at 25 Hz for a power-law GWB with a spectral index of 2/3 (consistent with expectations for compact binary coalescences), in the band 20-90.6 Hz; and Ω(sub GW)(f) ≤ 3.9 × 10(exp -10) at 25 Hz for a spectral index of 3, in the band 20-291.6 Hz. These upper limits improve over our previous results by a factor of 6.0 for a at GWB, 8.8 for a spectral index of 2/3, and 13.1 for a spectral index of 3. We also search for a GWB arising from scalar and vector modes, which are predicted by alternative theories of gravity; we do not find evidence of these, and place upper limits on the strength of GWBs with these polarizations. We demonstrate that there is no evidence of correlated noise of magnetic origin by performing a Bayesian analysis that allows for the presence of both a GWB and an effective magnetic background arising from geophysical Schumann resonances. We compare our upper limits to a fiducial model for the GWB from the merger of compact binaries, updating the model to use the most recent data-driven population inference from the systems detected during O3a. Finally, we combine our results with observations of individual mergers and show that, at design sensitivity, this joint approach may yield stronger constraints on the merger rate of binary black holes at z ≳ 2 than can be achieved with individually resolved mergers alone.

Ryan Abbott↗

Mapping Global Forest Age from Forest Inventories, Biomass and Climate Data

Forest age can determine the capacity of a forest to uptake carbon from the atmosphere. However, a lack of global diagnostics that reflect the forest stage and associated disturbance regimes hampers the quantification of age-related differences in forest carbon dynamics. This study provides a new global distribution of forest age circa 2010, estimated using a machine learning approach trained with more than 40 000 plots using forest inventory, biomass and climate data. First, an evaluation against the plot-level measurements of forest age reveals that the data-driven method has a relatively good predictive capacity of classifying old-growth vs. non-old-growth (precision = 0.81 and 0.99 for old-growth and non-old-growth, respectively) forests and estimating corresponding forest age estimates (NSE = 0.6 – Nash–Sutcliffe efficiency – and RMSE = 50 years – root-mean-square error). However, there are systematic biases of overestimation in young- and underestimation in old-forest stands, respectively. Globally, we find a large variability in forest age with the old-growth forests in the tropical regions of Amazon and Congo, young forests in China, and intermediate stands in Europe. Furthermore, we find that the regions with high rates of deforestation or forest degradation (e.g. the arc of deforestation in the Amazon) are composed mainly of younger stands. Assessment of forest age in the climate space shows that the old forests are either in cold and dry regions or warm and wet regions, while young–intermediate forests span a large climatic gradient. Finally, comparing the presented forest age estimates with a series of regional products reveals differences rooted in different approaches and different in situ observations and global-scale products. Despite showing robustness in cross-validation results, additional methodological insights on further developments should as much as possible harmonize data across the different approaches. The forest age dataset presented here provides additional insights into the global distribution of forest age to better understand the global dynamics in the forest water and carbon cycles. The forest age datasets are openly available at https://doi.org/10.17871/ForestAgeBGI.2021 (Besnard et al., 2021).

Simon Besnard↗

Machine Vision based Sample-Tube Localization for Mars Sample Return

A potential Mars Sample Return (MSR) architecture is being jointly studied by NASA and ESA. As currently envisioned, the MSR campaign consists of a series of 3 missions: sample cache, fetch and return to Earth. In this paper, we focus on the fetch part of the MSR, and more specifically the problem of autonomously detecting and localizing sample tubes deposited on the Martian surface. Towards this end, we study two machine-vision based approaches: First, a geometrydriven approach based on template matching that uses hardcoded filters and a 3D shape model of the tube; and second, a data-driven approach based on convolutional neural networks (CNNs) and learned features. Furthermore, we present a large benchmark dataset of sample-tube images, collected in representative outdoor environments and annotated with ground truth segmentation masks and locations. The dataset was acquired systematically across different terrain, illumination conditions and dust-coverage; and benchmarking was performed to study the feasibility of each approach, their relative strengths and weaknesses, and robustness in the presence of adverse environmental conditions.

Detry, R.↗

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

The Role in the Virtual Astronomical Observatory in the Era of Massive Data Sets

The Virtual Observatory (VO) is realizing global electronic integration of astronomy data. One of the long-term goals of the U.S. VO project, the Virtual Astronomical Observatory (VAO), is development of services and protocols that respond to the growing size and complexity of astronomy data sets. This paper describes how VAO staff are active in such development efforts, especially in innovative strategies and techniques that recognize the limited operating budgets likely available to astronomers even as demand increases. The project has a program of professional outreach whereby new services and protocols are evaluated.

data-driven science↗

Formulative Input into Future NASA Aeronautics Planning

This presentation covers industry input received for future work in NASA Aeronautics over the next 5 years. It is intended to present areas of significant imput and to stimulate further discussion.

future aeronautics planning↗