Search NASASearch

SEARCH · Search NASA

Results for “data gap”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

GenAI-Based Digital Twins Aided Data Augmentation Increases Accuracy in Real-Time Cokurtosis-Based Anomaly Detection of Wearable Data

Early detection of potential infectious disease outbreaks is crucial for developing effective interventions. In this study, we introduce advanced anomaly detection methods tailored for health datasets collected from wearables, offering insights at both individual and population levels. Leveraging real-world physiological data from wearables, including heart rate and activity, we developed a framework for the early detection of infection in individuals. Despite the availability of data from recent pandemics, substantial gaps remain in data collection, hindering method development. To bridge this gap, we utilized Wasserstein Generative Adversarial Networks (WGANs) to generate realistic synthetic wearable data, augmenting our dataset for training. Subsequently, we use these augmented datasets to implement a cokurtosis-based technique for anomaly detection in multivariate time-series data. Our approach includes a comprehensive assessment of uncertainties in synthetic data compared to the actual data upon which it was modeled, as well as the uncertainty associated with fine-tuning anomaly detection thresholds in physiological measurements. Through our work, we present an enhanced method for early anomaly detection in multivariate datasets, with promising applications in healthcare and beyond. This framework could revolutionize early detection strategies and significantly impact public health response efforts in future pandemics.

Data-Driven Digital Twins

Shuttle Orbiter boundary-layer transition - A comparison of flight and wind tunnel data

Hypersonic boundary-layer transition data obtained on the windward centerline of the Shuttle Orbiter during entry for the first four flights are presented and analyzed. Because the Orbiter surface is composed of a large number of thermal protection tiles, the transition data include the effects of distributed roughness arising from tile misalignment and gaps. These data are used as a benchmark for assessing and improving the accuracy of boundary-layer transition predictions based on correlations of wind tunnel data taken on both aerodynamically rough and smooth Orbiter surfaces. By comparing these two data bases, the relative importance of tunnel free-stream noise and surface roughness on Orbiter boundary-layer transition correlation parameters can be assessed. This assessment indicates that accurate predictions of transition times can be made for the Orbiter at hypersonic flight conditions by using roughness dominated wind tunnel data. Specifically, times of transition onset and completion can be accurately predicted using a correlation based on critical and effective values of a roughness Reynolds number previously derived from wind tunnel data.

Goodrich, W. D.

Shuttle orbiter boundary layer transition at flight and wind tunnel conditions

Hypersonic boundary layer transition data obtained on the windward centerline of the Shuttle orbiter during entry for the first five flights are presented and analyzed. Because the orbiter surface is composed of a large number of thermal protection tiles, the transition data include the effects of distributed roughness arising from tile misalignment and gaps. These data are used as a benchmark for assessing and improving the accuracy of boundary layer transition predictions based on correlations of wind tunnel data taken on both aerodynamically rough and smooth orbiter surfaces. By comparing these two data bases, the relative importance of tunnel free stream noise and surface roughness on orbiter boundary layer transition correlation parameters can be assessed. This assessment indicates that accurate predications of transition times can be made for the orbiter at hypersonic flight conditions by using roughness dominated wind tunnel data. Specifically, times of transition onset and completion is accurately predicted using a correlation based on critical and effective values of a roughness Reynolds number previously derived from wind tunnel data.

Goodrich, W. D.

Continental Patterns of Bird Migration Linked to Climate Variability

For nearly 100 years, avian migration studies have divided North America into three or four primary flyways, at times based on subjective approaches or just for convenience. Those studies often fail to adequately reflect a critical characterization of migration —phenology. This shortcoming has been partly due to the lack of reliable continental-scale data, a gap filled by our current study. Here, we leveraged unique radar-based data quantifying migration phenology and used an objective regionalization approach to revisit the traditional spatial framework. Consequently, we identified two regions with distinct inter annual variability of spring migration across the contiguous U.S. This new data-driven framework has enabled us to explore the climatic cues affecting the inter annual variability of migration phenology, “specific to each region” across North America. For example, our “two-region” approach allowed us to identify an east-west dipole pattern in migratory behavior linked to atmospheric Ross by waves. Also, we revealed a low-frequency variability in migration movements over the western U.S. that is inversely related with temperature and the Pacific Decadal Oscillation (PDO). Our spatial platform would facilitate future work on better understanding the mechanisms responsible for broad-scale migration phenology and its potential future changes.

Atmosphere

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS

Remote Sensing and Fluxes Upscaling for Real-world Impact (Workshop Report)

The "Remote Sensing and Fluxes Upscaling for Real-world Impact" workshop, held on July 9-10, 2024, at Lawrence Berkeley National Lab, was a collaborative effort led by the AmeriFlux Management Project, NEON, and the Carbon Dew Community of Practice. The event brought together over 200 registrants and approximately 100 attendees each day, including leading experts, researchers, and practitioners. The primary focus was on bridging the gap between cutting-edge research and practical applications in environmental monitoring by integrating remote sensing and flux data. Key themes included the importance of site-level measurements for validating remote sensing products, providing nature-based climate solutions, and addressing challenges such as instrument costs and the need for standardized methods. At the regional scale, discussions centered on addressing spatial heterogeneity and using high-resolution remote sensing and machine learning methods to enhance data interpretation. Global scale challenges included data consistency, gap filling, and accurate emission source identification, with opportunities for international collaboration and standardized practices to improve global carbon budget assessments. The workshop emphasized the critical need for integrating data across local, regional, and global scales through explicit scale-matching and developed a workflow for scaling flux data using "straight shot" and "explicit nesting" approaches. The event highlighted the importance of connecting scientific research with real-world applications in carbon, energy, and water management, ensuring that advancements translate into tangible societal benefits. These insights will guide future research, technology transfer, and collaboration, maximizing the potential of environmental fluxes to address real-world challenges.

97 MATHEMATICS AND COMPUTING

An Agile-Like Approach to Hardware Development: The Ejectable Data Recorder (EDR) for Orion's Ascent Abort 2 (AA-2) Test Flight

On July 2, 2019, the Ascent Abort 2 (AA-2) Flight Test Vehicle was launched from Cape Canaveral, with the goal of demonstrating the performance of Orion’s Launch Abort System (LAS) and collecting data from hundreds of sensors throughout the vehicle. The data collected during this test flight is of paramount importance, as it will be used to certify the Orion vehicle for human spaceflight. Originally, the data was to be downlinked via a single string network of antennas on the LAS, with the associated risk of potential data dropouts, as well as loss of data once the LAS was jettisoned. Thus, additional antennas were added onto the crew module (CM) to support data downlink post-LAS jettison, a buffer rebroadcast capability was added to fill in any gaps in data downlink transmissions, and an ejectable data recorder (EDR) subsystem was added to the CM as a redundant measure to collect all the instrumentation data. The EDR subsystem was added to the project about one year after the project commenced, which significantly reduced the available development time when compared with the other subsystems of the AA-2 Test Flight. The project was further accelerated by six months, around the critical design review gate. Due to the schedule compression challenge and the fact that the EDR subsystem was a backup system and not flight critical, the EDR subsystem was further challenged to find a new and more efficient way to develop hardware. Thus, the EDR subsystem experimented with different management and systems engineering processes, team sizes, communication methods, and tools. Some examples are novel uses of SharePoint as a Data-centric Project Management & Systems Engineering environment, a continuous testing approach through the lifecycle, and a Skunkworks approach to managing the team. The EDR subsystem blended Commercial Off The Shelf (COTS) hardware with in-house developed hardware and software to create a novel data retrieval capability. The capability evolved rapidly through a hardware in the loop simulation environment that enabled incremental component updates for not only the EDR subsystem but across the entire Crew Module. This paper will present an overview of how the EDR subsystem was managed and compare it to an Agile approach to managing projects. The paper will further provide a recommended approach to future Agile-like hardware development that incorporates lessons learned from the EDR experience.

Agile

Basin-Scale Structural Features Database

The Basin-Scale Structural Features database provides spatial datasets of faults, fractures, folds, and earthquakes compiled from public, authoritative sources (e.g., U.S. Geological Survey and State Geological Surveys) and aggregated into derivative forms to support subsurface assessments. Recognizing that characterizing basin-scale structural features requires interpreting data that are often ambiguous or lack key information, the source data were evaluated using a knowledge-data framework and geospatial fuzzy logic method (Justman et al., 2020) to represent both measured (observed) and predicted (inferred or potential) structural features as derivative datasets. This workflow employs conceptual models for known structural features and predicted structural features, incorporating geospatial data to estimate potential, even with limited data. The aim is to aid and support an understanding of basin-scale features and identify potential gaps in data and knowledge. As of 4/30/2025, the database includes resources for nine sedimentary basins: Appalachian, Denver, U.S. Gulf Coast, Illinois, Michigan, Permian, Sacramento, San Joquin and Williston. The database is organized by basin and then data category: 1) Faults, fractures, folds, 2) Earthquakes, 3) Topographic, 4) Structural contours and isopachs, 5) Geophysical, and 6) Structural feature density assessment maps.

basin scale

Computational tools and data integration to accelerate vaccine development: challenges, opportunities, and future directions

The development of effective vaccines is crucial for combating current and emerging pathogens. Despite significant advances in the field of vaccine development there remain numerous challenges including the lack of standardized data reporting and curation practices, making it difficult to determine correlates of protection from experimental and clinical studies. Significant gaps in data and knowledge integration can hinder vaccine development which relies on a comprehensive understanding of the interplay between pathogens and the host immune system. In this review, we explore the current landscape of vaccine development, highlighting the computational challenges, limitations, and opportunities associated with integrating diverse data types for leveraging artificial intelligence (AI) and machine learning (ML) techniques in vaccine design. We discuss the role of natural language processing, semantic integration, and causal inference in extracting valuable insights from published literature and unstructured data sources, as well as the computational modeling of immune responses. Furthermore, we highlight specific challenges associated with uncertainty quantification in vaccine development and emphasize the importance of establishing standardized data formats and ontologies to facilitate the integration and analysis of heterogeneous data. Through data harmonization and integration, the development of safe and effective vaccines can be accelerated to improve public health outcomes. Looking to the future, we highlight the need for collaborative efforts among researchers, data scientists, and public health experts to realize the full potential of AI-assisted vaccine design and streamline the vaccine development process.

60 APPLIED LIFE SCIENCES

UAM Fleet Manager Gap Analysis

NASA's Urban Air Mobility (UAM) Sub-Project is engaged in research to support the introduction of air taxis into the National Airspace System. Such operations will require arrange of communication, navigation, and surveillance systems. Air vehicles for UAM are under development and will initially have human pilots. Separation from other aircraft, obstacles, and weather may be a pilot responsibility or provided by an operator's ground-based systems. Eventually, air taxis may be flown from the ground or fly autonomously. There will be a need for dispatch services for UAM. This report presents a gap analysis, data and capability requirements, and workstation design concepts for the UAM dispatcher or Fleet Manager (FM) position. This presentation reviews the gap analysis report and outlines user interface design.

Mogford, Richard

AWSD Reactive Burn Model for High Explosive LX‐14

ABSTRACT The results of an Arrhenius–Wescott–Stewart–Davis (AWSD) reactive flow calibration for the HMX‐based high explosive LX‐14 are presented. The parameters in the AWSD model are calibrated to experimental thermodynamic and gas gun data and to computational results from thermochemical calculations. There is no experimental rate stick data available for LX‐14; therefore, scaled experimental results from other PBX‐based high explosives are used in the calibration to fill this gap in data. Strong agreement is observed between the calibrated AWSD model and experimental data for LX‐14, including validation data that were not used in the calibration procedure. The developed model more accurately describes experimental shock‐to‐detonation results compared to several other reactive flow models for LX‐14 from the literature. The presented results illustrate that the AWSD model is capable of quantitatively describing the reactive burn of LX‐14.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF

Fungal Spore Seasons Advanced Across the US Over Two Decades of Climate Change

Abstract Phenological shifts due to climate change have been extensively studied in plants and animals. Yet, the responses of fungal spores—organisms important to ecosystems and major airborne allergens—remain understudied. This knowledge gap limits our understanding of their ecological and public health implications. To address this, we analyzed a long‐term (2003–2022), large‐scale (the continental US) data set of airborne fungal spores collected by the US National Allergy Bureau. We first pre‐processed the spore data by gap‐filling and smoothing. Afterward, we extracted 10 metrics describing the phenology (e.g., start and end of season) and intensity (e.g., peak concentration and integral) of fungal spore seasons. These metrics were derived using two complementary but not mutually exclusive approaches—ecological and public health approaches, defined as percentiles of total spore concentration and allergenic thresholds of spore concentration, respectively. Using linear mixed‐effects models, we quantified annual shifts in these metrics across the continental US. We revealed a significant advancement in the onset of the spore seasons defined in both ecological (11 days, 95% confidence interval: 0.4–23 days) and public health (22 days, 6–38 days) approaches over two decades. Meanwhile, total spore concentrations in an annual cycle and in a spore allergy season tended to decrease over time. The earlier start of the spore season was significantly correlated with climatic variables, such as warmer temperatures and altered precipitations. Overall, our findings suggest possible climate‐driven advanced fungal spore seasons, highlighting the importance of climate change mitigation and adaptation in public health decision‐making.

Environmental Sciences & Ecology

Unsteady Flowfield in a High-Pressure Turbine Modeled by TURBO

Forced response, or resonant vibrations, in turbomachinery components can cause blades to crack or fail because of the large vibratory blade stresses and subsequent high-cycle fatigue. Forced-response vibrations occur when turbomachinery blades are subjected to periodic excitation at a frequency close to their natural frequency. Rotor blades in a turbine are constantly subjected to periodic excitations when they pass through the spatially nonuniform flowfield created by upstream vanes. Accurate numerical prediction of the unsteady aerodynamics phenomena that cause forced-response vibrations can lead to an improved understanding of the problem and offer potential approaches to reduce or eliminate specific forced-response problems. The objective of the current work was to validate an unsteady aerodynamics code (named TURBO) for the modeling of the unsteady blade row interactions that can cause forced response vibrations. The three-dimensional, unsteady, multi-blade-row, Reynolds-averaged Navier-Stokes turbomachinery code named TURBO was used to model a high-pressure turbine stage for which benchmark data were recently acquired under a NASA contract by researchers at the Ohio State University. The test article was an initial design for a high-pressure turbine stage that experienced forced-response vibrations which were eliminated by increasing the axial gap. The data, acquired in a short duration or shock tunnel test facility, included unsteady blade surface pressures and vibratory strains.

Bakhle, Milind A.

Extending the Nuclide Inventory Validation Basis for High-Burnup Fuel with New Radiochemical Assay Data

Efforts are underway at Oak Ridge National Laboratory to improve the nuclide inventory validation basis for spent nuclear fuel at high burnups. Recently conducted radiochemical assay experiments provided new measurement data for nine samples of fuel irradiated in a pressurized water reactor, with estimated sample burnups in the 30 to 70 GWd/t range. This type of destructive assay data is essential for validating computational methods, tools, and nuclear data applied in nuclear safety analyses and for improving our understanding of the bias and uncertainty in code predictions. The measurement data include key actinides and fission products that span a gamut of needs and interests for nuclear science and engineering applications in criticality safety, reactor physics, nuclide inventory, decay heat, and radiation shielding. The SCALE 6.3 code system with ENDF/B-VII.1 cross-section libraries was used to simulate the irradiation histories of the measured fuel samples. The calculated nuclide concentrations are compared to corresponding measurement data. The significance of the comparisons is discussed, emphasizing how the addition of the new measurement data fills gaps in the validation basis at high burnups and contributes to the decrease in bias and uncertainty for predicted nuclide concentrations. The discussion addresses the effect of the sample burnup used in the simulation—which is based on reactor operator records or on calibration to measured data for burnup indicator fission products—on the validation results.

Nuclide inventory

Convective Weather Forecast Quality Metrics for Air Traffic Management Decision-Making

Since numerical weather prediction models are unable to accurately forecast the severity and the location of the storm cells several hours into the future when compared with observation data, there has been a growing interest in probabilistic description of convective weather. The classical approach for generating uncertainty bounds consists of integrating the state equations and covariance propagation equations forward in time. This step is readily recognized as the process update step of the Kalman Filter algorithm. The second well known method, known as the Monte Carlo method, consists of generating output samples by driving the forecast algorithm with input samples selected from distributions. The statistical properties of the distributions of the output samples are then used for defining the uncertainty bounds of the output variables. This method is computationally expensive for a complex model compared to the covariance propagation method. The main advantage of the Monte Carlo method is that a complex non-linear model can be easily handled. Recently, a few different methods for probabilistic forecasting have appeared in the literature. A method for computing probability of convection in a region using forecast data is described in Ref. 5. Probability at a grid location is computed as the fraction of grid points, within a box of specified dimensions around the grid location, with forecast convection precipitation exceeding a specified threshold. The main limitation of this method is that the results are dependent on the chosen dimensions of the box. The examples presented Ref. 5 show that this process is equivalent to low-pass filtering of the forecast data with a finite support spatial filter. References 6 and 7 describe the technique for computing percentage coverage within a 92 x 92 square-kilometer box and assigning the value to the center 4 x 4 square-kilometer box. This technique is same as that described in Ref. 5. Characterizing the forecast, following the process described in Refs. 5 through 7, in terms of percentage coverage or confidence level is notionally sound compared to characterizing in terms of probabilities because the probability of the forecast being correct can only be determined using actual observations. References 5 through 7 only use the forecast data and not the observations. The method for computing the probability of detection, false alarm ratio and several forecast quality metrics (Skill Scores) using both the forecast and observation data are given in Ref. 2. This paper extends the statistical verification method in Ref. 2 to determine co-occurrence probabilities. The method consists of computing the probability that a severe weather cell (grid location) is detected in the observation data in the neighborhood of the severe weather cell in the forecast data. Probabilities of occurrence at the grid location and in its neighborhood with higher severity, and with lower severity in the observation data compared to that in the forecast data are examined. The method proposed in Refs. 5 through 7 is used for computing the probability that a certain number of cells in the neighborhood of severe weather cells in the forecast data are seen as severe weather cells in the observation data. Finally, the probability of existence of gaps in the observation data in the neighborhood of severe weather cells in forecast data is computed. Gaps are defined as openings between severe weather cells through which an aircraft can safely fly to its intended destination. The rest of the paper is organized as follows. Section II summarizes the statistical verification method described in Ref. 2. The extension of this method for computing the co-occurrence probabilities in discussed in Section HI. Numerical examples using NCWF forecast data and NCWD observation data are presented in Section III to elucidate the characteristics of the co-occurrence probabilities. This section also discusses the procedure for computing throbabilities that the severity of convection in the observation data will be higher or lower in the neighborhood of grid locations compared to that indicated at the grid locations in the forecast data. The probability of coverage of neighborhood grid cells is also described via examples in this section. Section IV discusses the gap detection algorithm and presents a numerical example to illustrate the method. The locations of the detected gaps in the observation data are used along with the locations of convective weather cells in the forecast data to determine the probability of existence of gaps in the neighborhood of these cells. Finally, the paper is concluded in Section V.

Chatterji, Gano B.

Climate Informatics

The impacts of present and potential future climate change will be one of the most important scientific and societal challenges in the 21st century. Given observed changes in temperature, sea ice, and sea level, improving our understanding of the climate system is an international priority. This system is characterized by complex phenomena that are imperfectly observed and even more imperfectly simulated. But with an ever-growing supply of climate data from satellites and environmental sensors, the magnitude of data and climate model output is beginning to overwhelm the relatively simple tools currently used to analyze them. A computational approach will therefore be indispensable for these analysis challenges. This chapter introduces the fledgling research discipline climate informatics: collaborations between climate scientists and machine learning researchers in order to bridge this gap between data and understanding. We hope that the study of climate informatics will accelerate discovery in answering pressing questions in climate science.

Climate change