Search NASA⌕ Search

SEARCH · Search NASA

Results for “Scientific Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Global Corn Heat Stress: Mean and SD of Degree Days Above 29°C based on NEX-GDDP-CMIP6 Climate Projections

Description This global dataset provides the estimated mean and standard deviation (SD) of corn heat stress (degree days above 29°C) for a set of climate models in NEX-GDDP-CMIP6 at 0.25-degree resolution. The NEX-GDDP-CMIP6 dataset is comprised of global downscaled climate scenarios derived from the General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 6 (CMIP6). The current dataset includes: Long-Term Average Degree Days Above 29°C- Historical Long-Term Average Degree Days Above 29°C- SSP245 Long-Term Standard Deviation of Degree Days Above 29°C- Historical Long-Term Standard Deviation of Degree Days Above 29°C- SSP245 The mean and SD are calculated over 1985-2014 for the historical period and over 2035-2064 for future projections. A full description of methods, including growing season, daily temperature distribution, and statistical coefficients, can be found in Haqiqi (2024). The source climate data are obtained from https://ds.nccs.nasa.gov/thredds2/catalog/catalog.html and are described in Thrasher et al (2022). The codes used to create this dataset are available at https://github.com/ihaqiqi/dd29c_nex_cmip6. Acknowledgments This work was supported by the US Department of Energy, Office of Science, Biological and Environmental Research Program, Earth and Environmental Systems Modeling, MultiSector Dynamics under Cooperative Agreement DE-SC0022141. The data processing, computation, and storage were completed on Purdue Anvil supercomputer and cyberinfrastructure supported by the National Science Foundation HDR award # 2118329: "NSF Institute for Geospatial Understanding through an Integrative Discovery Environment (I-GUIDE)". References Haqiqi. I. (2024). Trade can buffer climate-induced risks and volatilities in crop supply. Environmental Research: Food Systems. https://doi.org/10.1088/2976-601X/ad7d12 Thrasher, B., Wang, W., Michaelis, A., Melton, F., Lee, T., & Nemani, R. (2022). NASA global daily downscaled projections, CMIP6. Scientific Data, 9(1), 262. https://doi.org/10.1038/s41597-022-01393-4

Climate Change↗

Large-Scale Visualization of 3D Unstructured Groundwater Model Using Cave Automated Virtual Environment

The immersive three-dimensional (3D) virtual reality (VR) visualization of groundwater models allows us to deepen our understanding of aquifer systems and provide better solutions to present groundwater-related problems, such as groundwater recharge, water quality, and sustainability. Visualization assists in accurately developing groundwater models and revealing important subsurface features, including faulting, folding, and unconformity. However, assessing model accuracy poses challenges due to the complexity of geology and groundwater systems. This research demonstrates a workflow to visualize and analyze raw 3D unstructured groundwater model data using an immersive Cave Automated Virtual Environment (CAVE). To visualize the unstructured groundwater model data, the raw dataset is converted into interactive CAVE-compatible formats utilizing a set of tools: ParaView, Blender, and Unity. This enables researchers to immerse themselves in the data, identifying influential patterns and relationships. e resulting insights can inform the development of sophisticated machine-learning models for groundwater level prediction. The CAVE’s immersive capabilities allow intuitive exploration from various perspectives, providing a more holistic understanding of the factors affecting groundwater levels. These insights are crucial to improve predictive models. The CAVE results also facilitate collaborative analysis and have potential applications in training and education. is research demonstrates the value of immersive VR tools such as the CAVE for unraveling intricacies within high-dimensional scientific data to drive real-world forecasting and modeling applications.

54 ENVIRONMENTAL SCIENCES↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

Interdisciplinary Approaches Improve Understanding of Cryptogenic Species: A Historical Case Study of Crayfish in Montana, USA

ABSTRACT Cryptogenic species are those that are not yet “demonstrably native or introduced” in a given area, such as crayfish in western Montana, USA. Delving into evidence from diverse fields can help clarify the status of cryptogenic species. Primary historical sources and indigenous knowledge have informed various ecological questions but are seldom used to clarify the status of cryptogenic species, especially aquatic taxa. We summarize types of evidence used to illuminate species status and offer a case study applying historical sources to cryptogenic crayfish in Montana. We searched for crayfish mentions in dictionaries of indigenous languages and primary historical documents (e.g., travel journals, newspaper accounts, early biological surveys). Early explorer accounts we examined did not mention crayfish in Montana, though they noted them in the nearby Snake River drainage of Idaho and Wyoming. The first mention of crayfish in Montana that we found was from an 1868 newspaper account, probably referencing the upper Missouri River. The species was not noted, but an 1887 article suggested signal crayfish Pacifastacus leniusculus introductions near Bozeman, where non‐native signal crayfish still persist. The earliest crayfish mention we found west of the Great Divide in Montana was from 1934 in Ninepipes and probably referred to non‐native virile crayfish ( Faxonius virilis ). Signal crayfish were not mentioned west of the Divide until a 1944 newspaper account noted their 1937 introduction to the Bitterroot River drainage. The case study demonstrates the value of historical sources in identifying the absence of species and documenting their presence or introduction during periods predating scientific data.

Kennedy, Hampton L. [USDA Forest Service, Southern↗

Non-stationary precipitation design standards for stormwater infrastructure modernization at USAF installations

The resilience of defense infrastructure systems to a changing climate is critical for national security. Climate induced recurrent flooding is already impacting over 20 U.S. Air Force installations, underscoring the urgency of revisiting precipitation standards and stormwater infrastructure design. Despite growing scientific knowledge and an expanding set of tools for updating outdated precipitation standards based on the assumption of climate stationarity, the adoption of climate informed analyses remain limited in practice. This study utilizes an existing framework to update Intensity (or Depth)-Duration-Frequency (DDF) curves using an ensemble of future climate projections. Change factors in precipitation estimates are derived and applied to six USAF installations across the U.S. The analysis is further extended to evaluate the implications of climate-informed DDFs on stormwater infrastructure performance and flood analysis at Tyndall AFB. Results indicate that the current design precipitation estimates are likely to become obsolete in all six USAF bases by the end of the century. The wide range of change factors across 32 GCM ensembles highlights the need to integrate uncertainty and evolving scientific data into infrastructure planning. The study also finds that the impacts of a changing climate vary spatially and temporally, emphasizing the value of localized analysis for infrastructure decision-making. The work advances ongoing DoD and societal efforts to implement adaptation strategies aimed at enhancing infrastructure resilience.

Intensity-duration-frequency curves↗

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models

Abstract We propose masked particle modeling (MPM) as a self-supervised method for learning generic, transferable, and reusable representations on unordered sets of inputs for use in high energy physics (HEP) scientific data. This work provides a novel scheme to perform masked modeling based pre-training to learn permutation invariant functions on sets. More generally, this work provides a step towards building large foundation models for HEP that can be generically pre-trained with self-supervised learning and later fine-tuned for a variety of down-stream tasks. In MPM, particles in a set are masked and the training objective is to recover their identity, as defined by a discretized token representation of a pre-trained vector quantized variational autoencoder. We study the efficacy of the method in samples of high energy jets at collider physics experiments, including studies on the impact of discretization, permutation invariance, and ordering. We also study the fine-tuning capability of the model, showing that it can be adapted to tasks such as supervised and weakly supervised jet classification, and that the model can transfer efficiently with small fine-tuning data sets to new classes and new data domains.

Heinrich, Lukas (ORCID:0000000240487584)↗

eCounter: Inline Per-IP Network Monitoring at Millisecond Resolution via eBPF

Scientific data acquisition (SciDAQ) systems are shifting from archive-based workflows to streaming paradigms, where real-time, fine-grained network monitoring becomes essential. While P4-enabled devices offer per-packet in-band observability, they require specialized switches and routers. Host-side tools like Prometheus exporters lack sufficient temporal granularity. To bridge this gap, we present eCounter, a lightweight, hardware-agnostic, inline telemetry agent built on extended Berkeley Packet Filter (eBPF). eCounter captures per-interface ingress and egress traffic, categorized by IP address and protocol, at millisecond to sub-millisecond resolution. In a 100 Gbps environment, it continuously exports up to 3,257 time-series bins per second with only 4% CPU utilization at a 35¿KiB/s data rate. We evaluate eCounter across diverse NIC MTU settings, hook types, CPU architectures and operating systems, and observed negligible impact on concurrent high-throughput streaming applications. Complexity analysis confirms that it can be readily scaled to distributed SciDAQ deployments.

Mei, Xinxin [Computational Sciences and Technology↗

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat↗

Critical Literature Review of Low Global Warming Potential (GWP) Refrigerants and their Environmental Impact

Refrigeration and air conditioning currently account for ~20% of the total electricity consumption in buildings around the world. Over the next three decades as global temperatures are projected to increase, urbanization and economic growth will lead to an increased demand for refrigeration and cooling. Most commonly used refrigerants belong to the five following classes: (i) chlorofluorocarbons, (ii) hydrochlorofluorocarbons, (iii) hydrofluorocarbons (HFCs), (iv) hydrofluoroolefins (HFOs), and (v) natural refrigerants. Over the past century, there have been shifts in which compounds were used for refrigeration to improve safety and durability, allow for ozone protection, and, most recently, to reduce global warming potential (GWP). Although technological advances have led to increased cooling capacity and safer refrigerants, emissions from refrigeration systems can affect the environment by contributing to greenhouse gas emissions or by depleting the ozone layer, depending on the gas emitted. The focus is increasingly on adopting compounds that are both efficient at cooling and effective for reducing emissions and other adverse environmental impacts. Because of policy and regulatory changes to avert ozone depletion and global climate change, much discussion has centered on the environmental impacts of next-generation refrigerants. Of particular interest are the fluorinated refrigerants, HFCs and HFOs, most of which are defined as per- and polyfluoroalkyl substances (PFAS) and their breakdown products (especially trifluoroacetic acid). The US Environmental Protection Agency in 2021 drafted a Strategic Roadmap for PFAS, which has already resulted in an increase in investment in research on these compounds and has restricted the release of PFAS into the environment through the implementation of monitoring and reporting requirements. A critical evaluation of fluorinated refrigerants and their breakdown products with respect to persistence, biodegradation and toxicity, and global warming potential is needed to guide environmental regulations. This document aims to perform a critical review of the relevant scientific data on the most common refrigerants currently used, their degradation products, and their alternatives. Where available, estimates of precursor production quantities and existing environmental regulatory information are reviewed. Key data of interest for the evaluation include physicochemical properties, environmental fate parameters, ecological or human health toxicity/risk information, and GWP for compounds of interest.

54 ENVIRONMENTAL SCIENCES↗

Critical Literature Review of Low Global Warming Potential (GWP) Refrigerants and their Environmental Impact

Refrigeration and air conditioning currently account for ~20% of the total electricity consumption in buildings around the world. Over the next three decades as global temperatures are projected to increase, urbanization and economic growth will lead to an increased demand for refrigeration and cooling. Most commonly used refrigerants belong to the five following classes: (i) chlorofluorocarbons, (ii) hydrochlorofluorocarbons, (iii) hydrofluorocarbons (HFCs), (iv) hydrofluoroolefins (HFOs), and (v) natural refrigerants. Over the past century, there have been shifts in which compounds were used for refrigeration to improve safety and durability, allow for ozone protection, and, most recently, to reduce global warming potential (GWP). Although technological advances have led to increased cooling capacity and safer refrigerants, emissions from refrigeration systems can affect the environment by contributing to greenhouse gas emissions or by depleting the ozone layer, depending on the gas emitted. The focus is increasingly on adopting compounds that are both efficient at cooling and effective for reducing emissions and other adverse environmental impacts. Because of policy and regulatory changes to avert ozone depletion and global climate change, much discussion has centered on the environmental impacts of next-generation refrigerants. Of particular interest are the fluorinated refrigerants, HFCs and HFOs, most of which are defined as per- and polyfluoroalkyl substances (PFAS) and their breakdown products (especially trifluoroacetic acid). The US Environmental Protection Agency in 2021 drafted a Strategic Roadmap for PFAS, which has already resulted in an increase in investment in research on these compounds and has restricted the release of PFAS into the environment through the implementation of monitoring and reporting requirements. A critical evaluation of fluorinated refrigerants and their breakdown products with respect to persistence, biodegradation and toxicity, and global warming potential is needed to guide environmental regulations. This document aims to perform a critical review of the relevant scientific data on the most common refrigerants currently used, their degradation products, and their alternatives. Where available, estimates of precursor production quantities and existing environmental regulatory information are reviewed. Key data of interest for the evaluation include physicochemical properties, environmental fate parameters, ecological or human health toxicity/risk information, and GWP for compounds of interest.

54 ENVIRONMENTAL SCIENCES↗

Quantifying and Optimizing the Energy Benefits of Mass Timber Construction

The International Mass Timber Alliance (IMTA) is a global organization of industry leaders, engineers, scientists, and associations dedicated to advancing mass timber construction. Its mission is to generate and disseminate scientific data supporting the development of standardized construction and energy efficient practices that promote the adoption of mass timber worldwide. IMTA collaborated with Oak Ridge National Laboratory (ORNL) to leverage ORNL’s expertise in building envelope modeling and testing to evaluate how mass timber construction can reduce peak heating and cooling demand, lower overall energy use, and improve resilience during power outages. A previous study of 80 mass timber buildings in Finland found measured energy use up to 50% lower than predicted by simulation. This project aimed to validate and extend those findings for U.S. buildings through analytical modeling, laboratory testing, and full-scale building evaluations. The research focused on the thermal performance of low-embodied-energy wall assemblies, such as cross-laminated timber (CLT) panels and log walls, with particular attention to the effects of thermal inertia on indoor comfort and energy performance. While mass timber’s structural and fire-resistance properties are well documented, its whole-building thermal behavior has received limited attention. Field data, simulation results, and resilience testing from this study will inform future modeling practices, design guidelines, and construction practices by quantifying the unique thermal and demand-flexibility benefits of mass timber construction.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Model-Agnostic Signal Discovery with Machine Learning: Bridging the Gap Between Theory and Practice

Searches for new phenomena in complex scientific data are predominantly model-dependent, optimized for specific hypotheses, and therefore limited in their coverage of the space of possible signals. Recently, new AI-based model-agnostic search strategies, many of which have been pioneered in high-energy physics, have been proposed which provide a complementary paradigm, prioritizing broad exploration over tailored analyses. These techniques offer an opportunity to enhance the overall discovery potential of modern experiments, especially in regimes where theoretical guidance is scarce. In this document, we review the conceptual framework behind the main classes of AI-based model-agnostic strategies. We discuss the potential pitfalls of these methods, and strategies for their validation and interpretation. We aim for this document to serve as a useful reference both for practitioners and for researchers interested in learning more about these model-agnostic search strategies.

Amram, Oz [Fermilab] (ORCID:0000000237653123)↗

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING↗

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john↗

From Raw to Curated Data: A Lakehouse Approach for Scientific Workflows

This report provides a technical overview of how to go from raw to curated data in three stages using a lakehouse approach. We focus on the application of open source tools in scientific use cases (while noting parallels to enterprise and commercial alternatives). Our goal is to provide scientific data managers and infrastructure providers with a common frame of reference for understanding and applying modern lakehouse technologies and approaches.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Data Placement Optimization for ATLAS in a Multi-Tiered Storage System within a Data Center

Scientific experiments and computations, especially in High Energy Physics, are generating and accumulating data at an unprecedented rate. Effectively managing this vast volume of data while ensuring efficient data analysis poses a significant challenge for data centers, which must integrate various storage technologies. This paper proposes addressing this challenge by designing and developing a precise data popularity prediction model utilizing state-of-theart AI/ML techniques. This model is crafted from the analysis of ATLAS data and access patterns. It enables us to migrate infrequently accessed data to more economical storage media, such as tape drives, while storing frequently accessed data on faster yet costlier storage media like HDD or SSD. This strategic approach ensures data is placed optimally into the appropriate storage classes, thereby maximizing storage capacity while minimizing data access latency for end-users. Furthermore, the paper includes a performance evaluation of the prediction model using various key metrics such as F1 score, accuracy, precision and recall. Finally, we present a prototype use case, leveraging real-world file access data to assess the model’s impact on performance.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗