Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Predicting Large‐Scale Systematic Missing Pipe Attributes in Water Distribution Networks

Water distribution network (WDN) models are an essential tool used by water utilities for hydraulic analysis. Unfortunately, missing data and insufficient resources often make creating and maintaining these models unfeasible. Existing methods to address missing pipe properties, like sequential imputation for missing values and reconstruction using graph metrics, are designed to accommodate random patterns of missing information and require a significant percentage of the system's attributes to be known. However, these data completeness assumptions do not always align with real‐world scenarios where large sections of the WDN model have missing data. To address this challenge, this study proposes a data‐driven approach for estimating pipe diameter when considering different spatial patterns and degrees of data completeness (i.e., 0%–90%). Using data from 16 WDNs in Kentucky, this study compares the use of machine learning (ML) using topological and geospatial features against an existing deterministic approach. Results demonstrate that WDN models with pipe diameters predicted by the proposed ML method had comparable hydraulic performance to the ground truth models. Moreover, results showed that ML method performance varies between WDNs of differing topological classification. Insights from this study help advance the ability to leverage partial data to create and maintain WDN models amid uncertainty and inadequate resources.

Poff, Jason W. [Oregon State Univ., Corvallis, OR ↗

A universal language for finding mass spectrometry data patterns

Despite being information rich, the vast majority of untargeted mass spectrometry data are underutilized; most analytes are not used for downstream interpretation or reanalysis after publication. The inability to dive into these rich raw mass spectrometry datasets is due to the limited flexibility and scalability of existing software tools. Here, in this study, we introduce a new language, the Mass Spectrometry Query Language (MassQL), and an accompanying software ecosystem that addresses these issues by enabling the community to directly query mass spectrometry data with an expressive set of user-defined mass spectrometry patterns. Illustrated by real-world examples, MassQL provides a data-driven definition of chemical diversity by enabling the reanalysis of all public untargeted metabolomics data, empowering scientists across many disciplines to make new discoveries. MassQL has been widely implemented in multiple open-source and commercial mass spectrometry analysis tools, which enhances the ability, interoperability and reproducibility of mining of mass spectrometry data for the research community.

Damiani, Tito [Czech Academy of Sciences (CAS), Pr↗

Navigating Exascale Operational Data Analytics: From Inundation to Insight

In this paper, we address the challenges in achieving sustainable data-driven efficiency by providing a detailed exploration of the end-to-end operational data analytics (ODA) framework that evolved through two generations of supercomputer systems at the Oak Ridge Leadership Computing Facility (OLCF). This framework addresses large data streams ingested from heavily instrumented HPC environment that accumulates multi-terabytes per day. We outline the multifaceted data life cycle across HPC procurement, operations, and research & development, identifying key obstacles and design decisions that shape effective strategies in building and supporting data pipelines end-to-end. By sharing key insights and lessons learned from our experience, we offer recommendations for the HPC community on enabling sustainable operational data analytics and beyond. Our contributions aim to bridge the gap between potential and real benefits of operational data, guiding future efforts towards integrated and sustainable operational intelligence in high-performance computing environments.

Shin, Woong↗

Advancing Urban Water Resilience: Coproducing Knowledge through Civic–Academic Global Partnerships on Water and Climate

As extreme weather events become more pronounced, the vulnerabilities associated with the urban water supply and wastewater systems in megacities are intensified in multiple interconnected dimensions. These multifaceted water challenges can benefit from enhanced cross-sectoral collaboration and sharing of critical knowledge, which are essential for sustainable and adaptive water governance frameworks. In this context, the Megacity Alliance for Water and Climate (MAWAC)–Europe and North America Region (ENAR) Working Group convened a workshop in March 2023, followed by a subsequent workshop in London, United Kingdom, from 11 to 13 September 2024. These workshops aimed to investigate and devise solutions for the cascading hazards with water systems. The solutions examined various aspects focused on climate adaptation and mitigation, stormwater management, and the governance of water and wastewater systems. Additionally, discussions highlighted the importance of community engagement, economic considerations, equity, and effective communication in addressing these pressing challenges. Over the course of 3 days, experts from academia, government agencies, and industry engaged in meaningful discussions on digital modeling for integrated water management, climate-informed urban planning, and public–private–academic partnerships (Fig. 1). Case studies from cities such as New York, Los Angeles, London, Paris, and Chicago highlighted innovative governance strategies for managing water and wastewater systems, promoting water reuse, planning infrastructure, and fostering stakeholder-driven and stakeholder-informed adaptation. The workshop participants emphasized the need for data-driven decision-making, scalable governance models, and knowledge-sharing networks to enhance urban water governance for sustainability and resilience. This workshop report presents the key takeaways from the 3-day convening, providing a roadmap for integrating scientific research, policy frameworks, and emerging technologies to address water challenges faced by megacities.

Hydrologic models↗

Investigating Kinetic Mechanisms of Soot Formation in Plasma Pyrolysis of Methane via Active Learning (Final Technical Report)

Plasma pyrolysis of methane is an effective route for zero-carbon hydrogen production. Yet, soot generated from pyrolysis of hydrocarbons is detrimental to the climate and human health. There is ample experimental and theoretical evidence that suggests polycyclic aromatic hydrocarbons (PAHs) are the molecular precursors to soot particles. The reaction pathways of PAH formation are intricately dependent on a multitude of process parameters, whose kinetic mechanisms are not well-understood in plasma pyrolysis. This project aims to leverage advances in the kinetic modeling of soot formation in combustion, as well as in surrogate modeling and active learning, to systematically investigate the effects of process parameter on the kinetics of PAH formation in plasma pyrolysis of methane. To this end, we propose to use the PAH formation kinetics model developed by the PPPL/PU group based on the well-established ABF and HACA mechanisms, coupled with low-temperature plasma models. We will develop an active learning (AL) framework based on Bayesian optimization to systematically and data-efficiently explore the complex and multivariable parameter space of plasma pyrolysis in order to quantify the effects of plasma and feed parameters on the ABF and HACA kinetic pathways. AL is the branch of machine learning concerned with systematically querying samples from a system (experimental or computational) to train a data-driven model that maps design parameters to a performance criterion. We will use the data generated via AL to perform global sensitivity analysis, combined with uncertainty quantification, to elucidate the impact of different reaction pathways on minimizing formation of soot precursors. This study will result in an improved understanding of kinetics of PAH formation in plasma pyrolysis and can pave the way for more advanced mechanistic studies (e.g., soot nucleation mechanisms). Additionally, the findings will be useful for establishing practical strategies for increasing the pyrolysis efficiency and producing high-grade carbon for synthesis of nanomaterials.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Wind data for wind driven plant

Simple, averaged wind velocity data provide information on energy availability, facilitate generator site selection and enable appropriate operating ranges to be established for windpowered plants. They also provide a basis for the prediction of extreme wind speeds.

Stodhart, A. H.↗

Transportable Applications Environment (TAE) Plus: A NASA tool used to develop and manage graphical user interfaces

The Transportable Applications Environment (TAE) Plus was built to support the construction of graphical user interfaces (GUI's) for highly interactive applications, such as real-time processing systems and scientific analysis systems. It is a general purpose portable tool that includes a 'What You See Is What You Get' WorkBench that allows user interface designers to layout and manipulate windows and interaction objects. The WorkBench includes both user entry objects (e.g., radio buttons, menus) and data-driven objects (e.g., dials, gages, stripcharts), which dynamically change based on values of realtime data. Discussed here is what TAE Plus provides, how the implementation has utilized state-of-the-art technologies within graphic workstations, and how it has been used both within and without NASA.

Szczur, Martha R.↗

Machine Learning-Based Predictive Analytics for Aircraft Engine Conceptual Design

Big data and artificial intelligence/machine learning are transforming the global business environment. Data is now the most valuable asset for enterprises in every industry. Companies are using data-driven insights for competitive advantage. With that, the adoption of machine learning-based data analytics is rapidly taking hold across various industries, producing autonomous systems that support human decision-making. This work explored the application of machine learning to aircraft engine conceptual design. Supervised machine-learning algorithms for regression and classification were employed to study patterns in an existing, open-source database of production and research turbofan engines, and resulting in predictive analytics for use in predicting performance of new turbofan designs. Specifically, the author developed machine learning-based analytics to predict cruise thrust specific fuel consumption (TSFC) and core sizes of high-efficiency turbofan engines, using engine design parameters as the input. The predictive analytics were trained and deployed in Keras, an open-source neural networks application program interface (API) written in Python, with Google’s TensorFlow (an open source library for numerical computation) serving as the backend engine. The promising results of the predictive analytics show that machine-learning techniques merit further exploration for application in aircraft engine conceptual design.

deep-learning↗

Usage-based Lifing of Lithium-Ion Battery with HybridPhysics-Informed Neural Networks

Lithium-ion batteries are commonly used to power unmanned aircraft vehicles (UAVs).The ability to model and forecast the remaining useful life of these batteries enables UAV reliability assurance. Building accurate models for battery state of charge and state of health based on first principles is challenging due to the complex electrochemistry that governs battery operations and computational complexity required to solve them. Therefore, reduced order models are often used due to their ability to capture the overall battery discharge. Un-fortunately, these simplifications lead to residual discrepancy between model predictions and observed data. In this paper, we present a hybrid modeling approach merging reduced-order models and neural networks. In this approach, while most of the input-output relationship is captured by Nernst and Butler-Volmer equations, data-driven kernels reduce the gap between predictions and observations. We validate our approach using data publicly available through the NASA Prognostics Center of Excellence repository. Results showed that our hybrid battery prognosis model can be successfully calibrated, even with a limited number of observations.

Lithium-ion Battery↗

Re-Organizing Earth Observation Data Storage to Support Temporal Analysis of Big Data

The Earth Observing System Data and Information System archives many datasets that are critical to understanding long-term variations in Earth science properties. Thus, some of these are large, multi-decadal datasets. Yet the challenge in long time series analysis comes less from the sheer volume than the data organization, which is typically one (or a small number of) time steps per file. The overhead of opening and inventorying complex, API-driven data formats such as Hierarchical Data Format introduces a small latency at each time step, which nonetheless adds up for datasets with O(10^6) single-timestep files. Several approaches to reorganizing the data can mitigate this overhead by an order of magnitude: pre-aggregating data along the time axis (time-chunking); storing the data in a highly distributed file system; or storing data in distributed columnar databases. Storing a second copy of the data incurs extra costs, so some selection criteria must be employed, which would be driven by expected or actual usage by the end user community, balanced against the extra cost.

data storage↗

An Advanced Open-Source Platform for Air Quality Analysis, Visualization, and Prediction

Ambient air pollution is the largest environmental health risk factor, leading to several million premature deaths globally per year. The challenge of combating poor air quality is exacerbated by growing urban populations, changing emissions, and a warming climate. While there have been many advances monitoring and modeling of atmospheric composition, reflected in the dramatic increase in archived Earth Observations, there is no single measurement or method that alone can provide an accurate depiction of the entire atmosphere. The rapidly growing collections of observational and modeling data require us to be smarter about what data to include, and how such data is used. In recent years, NASA has invested significantly in advancing the concepts for Analytics Collaborative Framework (ACF) [5] and New Observing Strategies (NOS) [4] to tackle our software infrastructure need for harmonized data management and dynamic acquisition of diverse measurements for on-demand, interactive, multivariate analysis, and access [3]. It is not enough to have a big data, standalone analytics solution; it is critical that we start integrating data from remote sensing, modeling, and in-situ networks in a harmonized manner that enables timely and data-driven decision-making for air quality management. This work presents the design and development of an Air Quality Analytics Collaborative Framework (AQ ACF), as part of NASA’s Advanced Information Systems Technology (AIST) effort, to establish a data, machine-learning, and numerically driven platform for air quality analysis, visualization, and prediction.

Liu, Qian↗

Openet: Applications of Satellite-Based Evapotranspiration Data for Water Resources Management in the Western United States

Advancing water security in overallocated river basins globally requires consistent and reproducible information on consumptive use of water that can anchor the development of data-driven solutions to the challenge of balancing water supply and demand. OpenET is a fully automated system for field-scale (30 m), satellite-based mapping of evapotranspiration (ET) at daily, monthly and annual timesteps. OpenET currently provides spatially contiguous data throughout the 23 westernmost states in the continental US, and includes both current information as well as multi-year timeseries of ET. The OpenET consortium has implemented an ensemble of satellite-based ET models (ALEXI/DisALEXI, eeMETRIC, PT-JPL, geeSEBAL, SIMS and SSEBop) on Google Earth Engine, which provides a shared computing platform for collaboration on processing of data from Landsat and other satellites, land cover and meteorological inputs, leading to increased consistency and accuracy across the ensemble of models. Earth Engine also facilitates hosting and distribution of data via open data collections and an application programming interface. We provide updates on the OpenET framework, open data services and data access tools, approach to geographic expansion, recent accuracy assessments, and describe how a user-driven design approach has facilitated successful applications of OpenET data for a wide range of water resource management activities. Applications to date include: use of ET data to improve quantification of ET and consumptive use in Oregon, Utah and the Upper Colorado River Basin; streamlining of water use reporting requirements in the California Delta; support for calculation of water budgets for the implementation of the Sustainable Groundwater Management Act in California; and integration into decision support tools for irrigation management. The use cases demonstrate how satellite-derived ET data that are easily accessed and seen as broadly accepted can accelerate adoption of innovative water management practices at scale, and support advances in the sustainability of water supplies. Uptake and use of data by the OpenET science community has also led to advances in our understanding of the impacts of landcover change, irrigation intensification and wildfire events on hydrology and the water security.

Applications↗

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations↗

Code for the manuscript "Mori-Zwanzig Modal Decomposition"

We would like to create an open source repository in LANL's github on code written in Julia, in which we implement and extend the data-driven Mori-Zwanzig method for extracting large-scale spatio-temporal structures from data, which we call MZMD. This method is an extension of Dynamic Mode Decomposition (DMD) in which Mori-Zwanzig memory kernels are included into the associated companion matrix. In the code we would like to release, we apply MZMD to a flow over a cylinder with Reynolds number 100 rather than the much larger data set used in the associated manuscript. DMD is used extensively in the fluid dynamics community mainly for extracting large scale spatio-temporal structures (patters) from flow data. This is useful for understanding the key mechanisms that generate certain complex dynamical process relevant in engineering design. In MZMD, we improve upon DMD by adding the Mori-Zwanzig memory kernels, and show this improvement is especially important in strongly nonlinear regions of the flow.

Woodward, Michael↗

Secondary Science Teachers’ Implementation of a Curricular Intervention When Teaching With Global Climate Models

In the past decade, emphasis on promoting “climate literacy” in K-16 science classrooms has increased. Teachers play a critical role in cultivating these opportunities, especially in secondary science classrooms. However, most prior climate education research has focused on students and student learning; little is known about how teachers implement climate-focused curricular interventions. Here, we report findings from a concurrent mixed methods, multiple-case study of four secondary science teachers’ implementation of a new, NGSS-aligned, model-centric climate curriculum module grounded in the use of a data-driven, computer-based climate modeling tool—Easy Global Climate Model (EzGCM). We employ multiple data sources, including video-recorded classroom observations, interviews, and instructional artifacts, and both qualitative and quantitative analyses, to investigate how teachers implemented the curriculum. Findings show that, overall, teachers implemented the curriculum in ways that were less model-centric than designed, placing greater emphasis on EzGCM itself rather than using the model to investigate Earth’s changing climate. Additionally, we present detailed single-case studies of each participant teacher that highlight differences in teachers’ implementation of the curriculum module and their reasoning for making observed instructional decisions. This research sheds light on the design of secondary science learning environments by illustrating the varied ways teachers implement a climate-focused curriculum to support students’ developing climate literacy. This has important implications for the design of climate-focused curriculum and supports for teachers.

Secondary science teaching↗

Revisiting the Solar Research Cyberinfrastructure Needs: A White Paper of Findings and Recommendations

Solar and Heliosphere physics are areas of remarkable data-driven discoveries. Recent advances in high cadence, high-resolution multiwavelength observations, growing amounts of data from realistic modeling, and operational needs for uninterrupted science-quality data coverage generate the demand for a solar metadata standardization and overall healthy data infrastructure. This white paper is prepared as an effort of the working group “Uniform Semantics and Syntax of Solar Observations and Events” created within the “Towards Integration of Heliophysics Data, Modeling, and Analysis Tools” EarthCube Research Coordination Network (@HDMIEC RCN), with primary objectives to discuss current advances and identify future needs for the solar research cyberinfrastructure. The white paper summarizes presentations and discussions held during the special working group session at the EarthCube Annual Meeting on June 19th, 2020, as well as community contribution gathered during a series of preceding workshops and subsequent RCN working group sessions. The authors provide examples of the current standing of the solar research cyberinfrastructure, and describe the problems related to current data handling approaches. The list of the top-level recommendations agreed by the authors of the current white paper is presented at the beginning of the paper.

SMD↗

Leveraging Human Performance Data to Change the Narrative that People are the Safety Problem

The study of errors and failure has a long and productive history in the behavioral sciences. By studying how systems fail, we rule out various mechanisms for how those systems might work, thereby refining our theories of how they actually work. Human performance, however, includes more than errors; human performance comprises both failures and successes. A systematic bias to collect and analyze data only on error affects the decisions we make as a community by promoting the narrative that “people are the safety problem.” This narrative manifests in both obvious and subtle ways in the design of systems intended for human use. When the only safety data that are available are about human failure, then “data-driven” designs can only consider that humans fail. Changing this narrative will depend on new data and new ways to examine data – specifically, data on the processes by which human create and contribute to safety. An alternate narrative is that people represent a primary source of safety, through their capability to anticipate, monitor for, respond to, and learn from expected and unexpected change. This presentation will describe research efforts to expand the range of safety-relevant events to include not just rare safety failures but frequent safety successes. These efforts include use of data from both operations and simulations to develop methods and metrics for learning from structured observation, self-report, and system data.

Jon Holbrook↗

Extracting Material Property Measurements from Scientific Literature with Limited Annotations

Extracting material property data from scientific text is pivotal for advancing data-driven research in chemistry and materials science; however, the extensive annotation effort required to produce training data for named entity recognition (NER) models for this task often makes it a barrier to extracting specialized data sets. Here, in this work, we present a comparative study of the conventional, supervised NER methodology to alternative few-shot learning architectures and large language model (LLM)-based approaches that mitigate the need to label large training data sets. We find that the best-performing LLM (GPT-4o) not only excels in directly extracting relevant material properties based on limited examples but also enhances supervised learning through data augmentation. We supplement our findings with error and data quality assessments to provide a nuanced understanding of factors that impact property measurement extraction.

36 MATERIALS SCIENCE↗