Search NASASearch

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

NASA GeneLab Space Omics Database: Expanding from Space to Ionizing Radiation Data on the Ground

NASA GeneLab is an open-access repository for omics datasets generated by biological experiments conducted in space or ground experiments relevant to spaceflight (e.g. simulated cosmic radiation, simulated microgravity, bed rest studies). The GeneLab Data Systems (GLDS) version 4.0 will be available on October 1st 2019, and will provide a state-of-the-art bioinformatics platform for the space biology and radiation communities to upload their data into an omics data commons, to process their data with vetted standard workflows and to compare with existing analyses. Started in 2015 as a repository designed to archive omics data from space experiments, GeneLab has expanded its scope to all ionizing radiation omics experiments conducted on the ground and has put considerable effort in providing carefully characterized radiation metadata on all datasets. GeneLab is also providing processed data derived from the raw data covering a large spectrum of omics (genome, epigenome, transcriptome, epitranscriptome, proteome, metabolome) to help users explore important questions: 1) Which genes or proteins are expressed differently in space for various living organisms? 2) What specific DNA mutations or epigenetic changes happen in space or after exposure to ionizing radiation? and 3) How does genetics affect these responses? Processed data available on GeneLab are derived by standard data analysis workflows vetted by hundreds of scientists who volunteered to join one of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). In this presentation, we will discuss how to bridge the gap between irradiation studies performed on earth and biological experiments conducted in space since the early 1990's. We will discuss how radiation dosimetry was estimated for datasets derived from samples collected during the Space Shuttle era on the International Space Station and on other orbiting platforms. Finally, we will address future strategies regarding dose monitoring in future missions into space, inter-agency efforts to unify data under one umbrella, and knowledge dissemination across the radiation research community and the space biology community.

open-science

Hub stability in the calcium calmodulin-dependent protein kinase II

The calcium calmodulin protein kinase II (CaMKII) is a multi-subunit ring assembly with a central hub formed by the association domains. There is evidence for hub polymorphism between and within CaMKII isoforms, but the link between polymorphism and subunit exchange has not been resolved. Here, we present near-atomic resolution cryogenic electron microscopy (cryo-EM) structures revealing that hubs from the α and β isoforms, either standalone or within an β holoenzyme, coexist as 12 and 14 subunit assemblies. Single-molecule fluorescence microscopy of Venus-tagged holoenzymes detects intermediate assemblies and progressive dimer loss due to intrinsic holoenzyme lability, and holoenzyme disassembly into dimers upon mutagenesis of a conserved inter-domain contact. Molecular dynamics (MD) simulations show the flexibility of 4-subunit precursors, extracted in-silico from the β hub polymorphs, encompassing the curvature of both polymorphs. The MD explains how an open hub structure also obtained from the β holoenzyme sample could be created by dimer loss and analysis of its cryo-EM dataset reveals how the gap could open further. An assembly model, considering dimer concentration dependence and strain differences between polymorphs, proposes a mechanism for intrinsic hub lability to fine-tune the stoichiometry of αβ heterooligomers for their dynamic localization within synapses in neurons.

59 BASIC BIOLOGICAL SCIENCES

ORBIT-2 Dataset for Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling

This dataset release corresponds to the work conducted in ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling, where large-scale AI methods were applied to improve climate and weather resolution. The collection integrates four widely used, publicly available datasets: ERA5, PRISM, DAYMET, and IMERG. To prepare the data for ORBIT-2 model training and evaluation, we applied a preprocessing pipeline that generates paired low-resolution and high-resolution samples, enabling supervised downscaling experiments. The transformation from coarse to fine scales was performed using bilinear regridding, consistent with the procedures described in WeatherBench2, a community benchmark for weather and climate AI models. This dataset supports the development and evaluation of foundation models designed for weather and climate downscaling at exascale. Additional details on methodology and applications can be found in Wang et al., ORBIT-2 (arXiv:2505.04802, 2025).

54 ENVIRONMENTAL SCIENCES

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN

High-n Rydberg transition spectroscopy for heavy impurity transport studies in W7-X (invited)

Here, we present a novel spectroscopy approach to investigate impurity transport by analyzing line-radiation following high-n Rydberg transitions. While high-n Rydberg states of impurity ions are unlikely to be populated via impact excitation, they can be accessed by charge exchange (CX) reactions along the neutral beams in high-temperature plasmas. Hence, localized radiation of highly ionized impurities, free of passive contributions, can be observed at multiple wavelengths in the visible range. For the analysis and modeling of the observed Rydberg transitions, a technique for calculating effective emission coefficients is presented that can well reproduce the energy dependence seen in datasets available on the OPEN-ADAS database. By using the rate coefficients and comparing modeling results with the new high-n Rydberg CX measurements, impurity transport coefficients are determined with well-documented 2σ confidence intervals for the first time. This demonstrates that high-n Rydberg spectroscopy provides important constraints on the determination of impurity transport coefficients. By additionally considering Bolometer measurements, which provide constraints on the overall impurity emissivity and, therefore, impurity densities, error bars can be reduced even further.

Instruments & Instrumentation

Synthesizing land use and demographic change in Southeast Asia’s smaller urbanized areas from 2000–2015

The majority of the human population now reside in urban areas today. The United Nations estimates that nearly half of all urban dwellers currently live in cities smaller than 500 000 persons and the majority of future urban growth will take place in Asia and Africa, likely in these smaller urban areas, not mega cities. Thus, understanding the factors that influence urban demographic trajectories in small urban areas is critical to address sustainable and equitable policy initiatives related to food security, changing climate hazard exposure, and economic opportunities. Here we focus on Southeast Asia—a region historically characterized by lower urban population proportions, yet with a rapidly shifting dynamic demographic—to examine correlates of demographic change among smaller cities. We combine two open-source satellite-informed datasets: GHS urban center database (2015) and age-sex gridded data from WorldPop to calculate socio-demographic characteristics to model drivers of change in annualized urban population growth from 2000–2015 for 505 urbanized places. We find a general pattern of decreasing dependency ratios as city-size increases for most urban areas in Southeast Asia. Higher rates of growth and more variation is observed for smaller cities—those with fewer than 300 000 persons, the lowest population limit for UN data on urbanization. When examining covariates of urban population growth, we find significant statistical associations of population change in smaller urbanized areas with climatic, economic, and land cover/land use variables, but with country-specific variations. Characterizing a continuum of urban population development in the context of changing environmental, economic and climate conditions has been an important sustainable development and equity issue for decades, but newer analysis of city-level drivers allows for systematic inquiry thus moving beyond total population counts for policy-relevant insight.

Southeast Asia synthesis

SO(3)-invariant PCA with application to molecular data

Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.

Fraiman, Michael [Tel Aviv Univ., Tel Aviv (Israel

New model for the ion collection by cylindrical probes over a wide range of collisionality

Langmuir probes remain one of the most important diagnostic tools for plasma processing applications. Modern probe analysis usually relies on the electron current part of the Langmuir probe characteristic using the Druyvesteyn method. However, for electronegative plasmas or for discharges containing dust the analysis of the ion current attracted by the probe can be desirable to determine the ion density. But, even at low pressures of a few Pa, the ion current is affected by collisions due to the large cross section for charge exchange. Available theories for collisional or collision-enhanced ion currents onto probes are complex and not well validated. Thus, in this contribution, we compare available collisional probe theories for the ion current to results of particle-in-cell (PIC) simulations. To this end, the probe surrounded by a semi-infinite plasma is simulated using a modified version of the open-source code EDIPIC. A dataset of currents for different neutral gas pressures is obtained and compared to the different theories from the literature. Based on these results, we propose a simpler and more intuitive model for the ion current collected by the probe, based on the model of Gatti and Kortshagen (Phys. Rev. E 78, 046402, 2008), developed for the charging of dust particles.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Biomolecular Analysis Capability for Cellular and Omics Research on the International Space Station

International Space Station (ISS) assembly complete ushered a new era focused on utilization of this state-of-the-art orbiting laboratory to advance science and technology research in a wide array of disciplines, with benefits to Earth and space exploration. ISS enabling capability for research in cellular and molecular biology includes equipment for in situ, on-orbit analysis of biomolecules. Applications of this growing capability range from biomedicine and biotechnology to the emerging field of Omics. For example, Biomolecule Sequencer is a space-based miniature DNA sequencer that provides nucleotide sequence data for entire samples, which may be used for purposes such as microorganism identification and astrobiology. It complements the use of WetLab-2 SmartCycler"TradeMark", which extracts RNA and provides real-time quantitative gene expression data analysis from biospecimens sampled or cultured onboard the ISS, for downlink to ground investigators, with applications ranging from clinical tissue evaluation to multigenerational assessment of organismal alterations. And the Genes in Space-1 investigation, aimed at examining epigenetic changes, employs polymerase chain reaction to detect immune system alterations. In addition, an increasing assortment of tools to visualize the subcellular distribution of tagged macromolecules is becoming available onboard the ISS. For instance, the NASA LMM (Light Microscopy Module) is a flexible light microscopy imaging facility that enables imaging of physical and biological microscopic phenomena in microgravity. Another light microscopy system modified for use in space to image life sciences payloads is initially used by the Heart Cells investigation ("Effects of Microgravity on Stem Cell-Derived Cardiomyocytes for Human Cardiovascular Disease Modeling and Drug Discovery"). Also, the JAXA Microscope system can perform remotely controllable light, phase-contrast, and fluorescent observations. And upcoming confocal microscopy capability will allow for optical sectioning of biological tissues to determine microanatomical localization of biomarkers. Furthermore, NASA's geneLAB effort addresses integration of genomic, epigenomic, transcriptomic, proteomic and metabolomic datasets, by applying an innovative open source science platform for multi-investigator high throughput utilization of the ISS. In sum, the expanding ISS capability for analysis of biomolecules is enabling innovative research in a broad spectrum of areas such as cellular and molecular biology, biotechnology, tissue engineering, biomedicine, and Omics, providing manifold benefits for humanity.

Guinart-Ramirez, Y.

Climates of Warm Earth-Like Planets. I. 3D Model Simulations

We present a large ensemble of simulations of an Earth-like world with increasing insolation and rotation rate. Unlike previous work utilizing idealized aquaplanet congurations we focus our simulations on modern Earth-like topography. The orbital period is the same as modern Earth, but with zero obliquity and eccentricity. The atmosphere is 1 bar N2-dominated with CO2=400 ppmv and CH4=1 ppmv. The simulations include two types of oceans; one without ocean heat transport (OHT) between grid cells as has been commonly used in the exoplanet literature, while the other is a fully coupled dynamic bathtub type ocean. The dynamical regime transitions that occur as day length increases induce climate feedbacks producing cooler temperatures, rst via the reduction of water vapor with increasing rotation period despite decreasing shortwave cooling by clouds, and then via decreasing water vapor and increasing shortwave cloud cooling, except at the highest insolations. Simulations without OHT are more sensitive to insolation changes for fast rotations while slower rotations are relatively insensitive to ocean choice. OHT runs with faster rotations tend to be similar with gyres transporting heat poleward making them warmer than those without OHT. For slower rotations OHT is directed equator-ward and no high latitude gyres are apparent. Uncertainties in cloud parameterization preclude a precise determination of habitability but do not a affect robust aspects of exoplanet climate sensitivity. This is the first paper in a series that will investigate aspects of habitability in the simulations presented herein. The datasets from this study are open source and publicly available.

Planetary systems

DNA Damage Response to Low and High-LET in a Large Cohort of Mice and Humans and Latest Advancement in NASA Space Omics

This presentation will first focus on a thorough evaluation of the DNA damage response to both low and high-LET in a cohort of 76 mice primary skin fibroblast derived from 15 different strains or in human blood mononuclear cells derived from 550 healthy donors. In both the human and mice work, we have hypothesized that DNA repair capacity can be used as a marker to evaluate and differentiate individual radiation sensitivity. More specifically, this work is based on the concept that the combined time-dose dependence of radiation-induced foci (RIF) of p53-binding protein 1 (53BP1) following low-LET exposure contains sufficient information to infer sensitivity to any other LET. This work is one of the most extensive studies on the kinetics and possible genetic underpinnings of radiation-induced DNA damage and repair. Results on humans are still preliminary as we are still in the process of collecting and isolating primary blood mononuclear cells from 500 to 800 healthy subjects of European descent, 18-75 years of age, 50/50 male/female distribution. We have analyzed 53BP1+ RIF formation as well as oxidative stress and cell death in primary cells from 192 subjects in response to the same HZE particles as used in mice: 600 MeV/n Fe, 350 MeV/n Ar and 350 MeV/n Si, 1.1 and 3 particles/100m2, 4 and 24 hours after irradiation. The second part of the talk will focus on describing GeneLab: The NASA Systems Biology Platform for Space Omics Repository, Analysis and Visualization. NASA GeneLab is an open-access repository for omics datasets generated by biological experiments conducted in space or experiments relevant to spaceflight (e.g. simulated cosmic radiation, simulated microgravity, bed rest studies). Started as a repository designed to archive precious omics from space experiments, GeneLab has expanded its scope to maximize the intelligibility of the raw data (e.g. RNAseq, microarray, WGBS, metagenome), particularly for users with limited bioinformatics knowledge. As such GeneLab is now providing processed data derived from the raw data covering a large spectrum of omics (genome, epigenome, transcriptome, epitranscriptome, proteome, metabolome), to help users explore important questions: Which genes or proteins are expressed differently in space for various living organisms? What are the consequences arising from these changes? What specifics DNA mutations or epigenetic changes happen in space? What species or genetic features lead to better adaption to such a unique environment? In this presentation, we will report on the current and future objectives for GeneLab, and review recent published studies relating molecular changes observed in various animal models and tissue with microgravity, radiation, circadian rhythm, hydration and carbon dioxide conditions.

DNA repair kinetics

A Quantitative Analysis on the Use of Supervised Machine Learning in Earth Science

Recent review papers (Ball et al., 2017; Reichstein et al., 2019) have investigated the opportunities and challenges in applying supervised machine learning (ML) techniques to Earth science problems. A common challenge is the lack of training (or labeled) data. Supervised ML, and especially deep learning (DL), require large training datasets. While there are large, open access Earth science archives, the data typically require preprocessing in preparation for supervised ML, frequently including manual labeling. Our objective is to understand the landscape of supervised ML in the Earth sciences, including which research communities have most rapidly adopted supervised ML, which algorithms are applied, and what data are used to train these algorithms. We conducted a literature survey of Earth science papers published during the last 10 years in journals from the American Geophysical Union (AGU), American Meteorological Society (AMS), the Institute of Electrical and Electronics Engineers(IEEE), and the Society of Photo-Optical Instrumentation Engineers (SPIE). We identified papers containing the terms ML, DL, or the names of individual supervised ML algorithms. "Earth science" is an additional required search term for IEEE and SPIE. We investigate trends in supervised ML usage during the 10-year study period, and manually analyzed AGU papers from 2018-2019 to enable deep-dive statistics.

Katrina S Virts

Hydrologic Model Data for the East Fork Poplar Creek Watershed Simulated with the Advanced Terrestrial Simulator (ATS): Streamflow and Network Expansion–Contraction Dynamics

This dataset supports hydrologic modeling and stream network expansion–contraction analysis for the East Fork Poplar Creek (EFPC) Watershed in Tennessee. It includes a Jupyter notebook for model setup, model configuration files, simulation outputs, and derived products used to evaluate model performance and investigate stream dynamics under varying hydrologic conditions. The dataset was generated using the Watershed Workflow Python package and the Advanced Terrestrial Simulator (ATS), enabling integrated surface–subsurface hydrologic simulations using a stream-aligned mesh. Outputs include high-resolution time series of streamflow, active network length, water table depth, and related hydrologic variables. Also included are spatially explicit stream persistency indices and classifications of reaches as perennial or non-perennial. These data facilitate reproducibility and support further research on stream intermittency and variability in network extent.The model data archive is organized in following directories:1) model_setup_inputsContains the Watershed Workflow Jupyter notebooks (accessed through any open source code editor), selected input datasets, and resulting ATS input files, including XML files (access through any open source code editor), computational mesh (.exo files can be viewed using Paraview), and meteorological forcing files (.h5 files can be accessed through h5py python package and HDFView open source software). 2) model_outputsIncludes ATS simulation outputs relevant to this study. Time series of spatially integrated or averaged variables (e.g., streamflow, water table depth) are provided as CSV files. Select spatial fields (e.g., ponded depth and water table depth) are saved as pickled Python objects to reduce file size, and can be accessed through pickle package in Python. Key geometry objects from Watershed Workflow—such as the surface mesh and river tree—are also included to support analysis of streamflow persistency and expansion–contraction dynamics. These files can also be accessed through Watershed Workflow Python package.3) model_evaluationProvides observed streamflow time series and field survey-based flow regime classifications used to evaluate model performance. Jupyter notebooks for processing ATS outputs and comparing model predictions with observations to build confidence in the model prior to scientific analysis are also included.4) Q_L_relationshipsContains workflows for generating time series of discharge, active network length, and related hydrologic variables used in the stream network expansion–contraction analysis. Includes routines for delineating baseflow-dominated periods. For each catchment, notebooks and processed data (as pickled DataFrames accessed through Pandas Python package) are provided. 5) figure_scriptsProvides the Jupyter notebooks used to generate the figures presented in the paper.

54 ENVIRONMENTAL SCIENCES

Videos and front speeds of frontal ring-opening metathesis polymerization (FROMP) of DCPD/ENB with norbornene-functionalized PDMS comonomers

This dataset contains videos and front speed measurements for 16 frontal ring-opening metastasis polymerization experiments of dicyclopentadiene (DCPD)/5-ethylidene-2-norbornene (ENB) resins and norbornene-functionalized polydimethylsiloxane (nor-PDMS) comonomers. Each run was carried out in a 10 mm diameter glass test tube and recorded to quantify front propagation behavior. Reported front speeds were extracted by video tracking and reported maximum front temperatures were measured with a thermocouple.

Clarke, Brandon R.

Wind Turbine Sound Setbacks and Supply Curves: Ordinances and Extrapolated Trends, 110 Hub Height, 130 Rotor Diameter

This dataset provides a comprehensive set of wind turbine sound setbacks from every residential structure in the contiguous United States (CONUS). A sound setback is defined as the minimum required distance between a residential structure and a hypothetical turbine installation site to ensure that modeled sound levels received at the residence do not exceed local sound ordinances, which are commonly expressed in A-weighted decibels (dBA). Therefore, sound setbacks are a local spatial assessment combining multiple factors, including the sound pressure curve as a function of the observer location (distance and direction) relative to the turbine, local sound regulations, and the geographical distribution of residential structures. The dataset is organized into multiple scenario-based products, detailed as follows: 1. Existing and extrapolated sound setbacks. An existing scenario characterizes sound setbacks only in states or counties that have implemented sound regulations as of 2022. The extrapolated scenarios extend a constant sound threshold to counties that lack explicit sound regulations, with thresholds ranging from 35 to 60 dBA, in 5-dBA increments reflecting the variation observed in current sound ordinances. 2. Sound setbacks in directional and worst scenarios. The directional scenario accounts for the distance and orientation of residential structures relative to a hypothetical turbine location, utilizing the turbine's sound emissions in that specific direction. In contrast, the worst scenario takes loudest sound level at each distance step from the turbine, irrespective of directional considerations, which aligns with current industry practice. 3. Supply curves for Open and Reference Access scenarios. This dataset includes supply curves generated by the reV model, which integrates each of the above sound setbacks into both Open and Reference siting scenarios. In addition, two Open and Reference baselines scenarios were included which do not consider sound setbacks for comparative analysis. All sound setback data are stored in TIF files, with partial maps of the data provided in PNG format. The values in the sound setback raster range from 0 to 1, representing the fraction of developable land within a 90 meter by 90 meter pixel due to sound ordinances. A value of 0 indicates areas where wind energy development is prohibited, while a value of 1 signifies areas fully permissible. The wind turbine parameters used in the sound modeling are based on the land-based turbine from International Energy Agency (IEA), featuring a rated electrical power of 3.4 MW, a rotor diameter of 130 meters, and a hub height of 110 meters. The atmospheric conditions, including wind speed/direction, turbulence, air temperature, relative humidity, and air pressure, that drive the sound generation are obtained from the WIND Toolkit dataset.

Array

Advancing Open Science in Atmospheric Research: Integrating Data Usability and Machine Learning

In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences." "In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences.

Jennifer Wei

An Exploratory Data Mining Investigation for Constructing a Publicly Sourced Dataset of Foreign Hypersonic Tests

This document details a data mining exercise that resulted in an exploratory dataset of publicly reported foreign (non-US) hypersonic vehicle test events. Using a combination of targeted English language searches and country-specific queries, the study aggregates information from digital news media, official press releases, and social media posts. The resulting list of events captures the publicly available accounts of foreign hypersonic tests, although it does not represent an exhaustive record. Limitations such as inconsistent reporting, translation challenges, and the inherently provisional nature of open-source data are acknowledged. This dataset serves as an initial reference point for further inquiries into high-speed atmospheric phenomena and may facilitate future efforts to correlate these events with geophysical measurements.

33 ADVANCED PROPULSION SYSTEMS

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS