Search NASA⌕ Search

SEARCH · Search NASA

Results for “science data system”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Development of a strip-shaped X-ray mapping system for 9-cell superconducting cavities

Electrons emitted via field emission during superconducting (SC) radio-frequency (RF) cavity tests at vertical test stands often collide with the iris region inside the cavity, generating X-rays at these locations. In 1.3 GHz 9-cell SC RF cavities designed for the International Linear Collider (ILC), stiffener rings located outside the iris region between cells can interfere with X-ray detection, complicating the precise identification of field emission sites. Hence, in this study, we developed a high-density strip X-ray mapping systems (sX-map) that can be inserted into the iris region of ILC-type 9-cell SC RF cavities. This sX-map facilitates efficient and accurate detection of X-rays generated near the irises, unaffected by the presence of stiffener rings. The sX-map consisted of 32 sensors per strip, with sensors spaced approximately 10 mm apart. It was deployed in every iris of the 9-cell cavity, using a total of 320 sensors. A multiplexer was employed to facilitate the readout of a large number of detectors using a minimal number of signal lines, connecting the strips inside within the vertical test cryostat. In a vertical test conducted at Jefferson Lab (JLab), we demonstrated the capability of sX-map to detect X-rays despite the presence of a stiffener ring. This paper presents the detailed design of the sX-map and the results from the vertical test at JLab.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Basic Research Needs for Inverse Methods for Complex Systems under Uncertainty [Brochure]

The four priority research directions outlined in this brochure represent a cohesive vision for advancing the science of inverse problems for complex systems under uncertainty. Together, they address the critical challenges of: discovering, exploiting, and preserving physical and problem structure; overcoming model limitations; integrating disparate, multimodal, and/or dynamic data; and tailoring the solution of inverse problems to downstream tasks. While each PRD focuses on a distinct aspect of inverse-problem research, their interconnected nature highlights the importance of a holistic approach that leverages progress across all areas to achieve transformative solutions. This agenda calls for research across mathematics, statistics, and computer science disciplines, which are guided and complemented by rapid advances in artificial intelligence, high-performance computing, and experimental facilities, to unlock new capabilities, maximize scientific impact, and meet the growing demands of inverse problems that arise across applications that are critical to DOE's mission.

97 MATHEMATICS AND COMPUTING↗

SOHIP Abel Transform and Onion Peeling Model Module

This software provides tools for analyzing and modeling physical systems using mathematical transforms and layered models. It includes (1) functions for performing the Abel transform, which is used to relate measurements of bending angles to properties such as refractive index and radius in a medium. The code can compute bending angles from input profiles and also reconstruct these profiles from observed data; (2) the functions for modeling systems with multiple layers using an onion-peeling approach, allowing users to simulate and analyze the behavior of layered materials or structures. These capabilities are useful for researchers and engineers working in fields such as optics, atmospheric science, and materials analysis, enabling them to interpret and model data from experiments or simulations relates to refraction in spherical symmetric medium.

Xu, Shuang [Lawrence Livermore National Laboratory↗

Uncertainty-Driven Rapid Thermodynamic Assessment of Nb-Ta-Zr System and Effects of C impurities (L25GF9298S): Annual Progress Report

An integrated computational materials engineering (ICME) method is in development for rapid thermodynamic experimental investigation and high-fidelity computational modeling of refractory multi-principal element alloys (RMPEAs). These ultra-high-temperature (UHT) alloys are of interest for structural applications in extreme environments, but deficiency of reliable data, especially melting temperatures, impedes the prediction of alloys with favorable properties. The method leverages UHT capabilities and computational expertise of LLNL’s Materials Science Division and the McCormack Lab’s UHT conical nozzle levitation (CNL) system to iteratively map the Nb-Ta-Zr phase space, with focus on the liquidus surface, through targeted experiments selected by quantifying uncertainty in the thermodynamic model fitting parameters. This method will reduce the time to map uncharted RMPEA phase space and thereby accelerate discovery and development of advanced materials for applications in extreme environments.

36 MATERIALS SCIENCE↗

CMIP7 Data Request: atmosphere priorities and opportunities

This paper presents a comprehensive overview of the Coupled Model Intercomparison Project Phase 7 (CMIP7) request for data unlocking key research avenues in atmospheric science and provides justification for the resources needed to produce this data. Topics within the CMIP7 Atmosphere Theme centre around processes and feedbacks in atmospheric science such as clouds, aerosols and atmospheric chemistry, atmospheric circulation, temperature variability and extremes, radiative forcings, and Earth system model evaluation. These topics are summarised in this paper as scientific “opportunities” which will be realised through CMIP7 experiments and Earth system model outputs. These opportunities were submitted by a thematic group of atmospheric science community representatives combined with an extended consultation process. The production of these variables will close key gaps and uncertainties identified during previous rounds of CMIP, and will be broadly used by scientific, policy, governmental, industry, and other communities that rely on climate model projections for research and decision making, including supporting the 7th Intergovernmental Panel on Climate Change Assessment Report (AR7). As an author group, we also reflect on the process used to collate this data request and make recommendations to future CMIP governance on implementing a consultation on this scale in the future.

58 GEOSCIENCES↗

Curating Carbon Storage Data for Reuse: Enabling Research and Modeling from Earth’s Surface to Subsurface

The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.

Morkner, Paige↗

Towards Next-Generation Urban Decision Support Systems through AI-Powered Construction of Scientific Ontology Using Large Language Models—A Case in Optimizing Intermodal Freight Transportation

The incorporation of Artificial Intelligence (AI) models into various optimization systems is on the rise. However, addressing complex urban and environmental management challenges often demands deep expertise in domain science and informatics. This expertise is essential for deriving data and simulation-driven insights that support informed decision-making. In this context, we investigate the potential of leveraging the pre-trained Large Language Models (LLMs) to create knowledge representations for supporting operations research. By adopting ChatGPT-4 API as the reasoning core, we outline an applied workflow that encompasses natural language processing, Methontology-based prompt tuning, and Generative Pre-trained Transformer (GPT), to automate the construction of scenario-based ontologies using existing research articles and technical manuals of urban datasets and simulations. From these ontologies, knowledge graphs can be derived using widely adopted formats and protocols, guiding various tasks towards data-informed decision support. The performance of our methodology is evaluated through a comparative analysis that contrasts our AI-generated ontology with the widely recognized pizza ontology, commonly used in tutorials for popular ontology software. We conclude with a real-world case study on optimizing the complex system of multi-modal freight transportation. Our approach advances urban decision support systems by enhancing data and metadata modeling, improving data integration and simulation coupling, and guiding the development of decision support strategies and essential software components.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

Repository of HydroSMADE: Hydropower Site-level Monthly Availability Data Ensemble for 1950-2100 at Existing and Potential Global Sites

This repository presents HydroSMADE—Hydropower Site-level Monthly Availability Data Ensemble, a new open dataset that provides monthly hydropower availability for 1,593 existing and 124,333 potential sites worldwide over the period 1950–2100. The dataset is generated by using a global hydrologic model (Xanthos) with explicit representation of hydropower operation. Specifically, HydroSMADE distinguishes between storage and diversion sites, applies optimized operating rules, and incorporates site-specific characteristics such as generation capacity, maximum turbine flow, and reservoir storage. Driven by bias-corrected meteorological inputs, the data is provided for 30 alternative future scenarios. The scenarios consist of the full factorial combination of three standard CMIP6 atmospheric forcing pathways (SSP1-2.6, SSP3-7.0, and SSP5-8.5) and ten CMIP6 General Circulation Models (GCMs): GFDL-ESM4, IPSL-CM6A-LR, MPI-ESM1-2-HR, MRI-ESM2-0, EC-Earth3, CanESM5, MIROC6, CNRM-ESM2-1, UKESM1-0-LL, and CNRM-CM6-1. The repository contains a total of 122 files: a text file (readme.txt) containing a brief description of the included data, a CSV file containing site attributes, and the remaining 120 files (in CSV) containing site-level monthly hydropower availability. Example Jupyter Notebooks to explore the HydroSMADE dataset are available on GitHub at https://github.com/kamal0013/HydroSMADE More details on the methods and technical validation of HydroSMADE are available in the following paper by the same authors: Chowdhury, A. K., Abeshu, G. W., Zhao, M., Wild, T. B., Hassan, N., Ying, Z., Kim, G. J., Matthew, B., Jonathan, L., & Li, H.-Y. (Submitted). Hydropower Site-level Monthly Availability Data Ensemble for 1950-2100 at Existing and Potential Global Sites.

Existing and Potential Sites↗

Employing artificial intelligence to steer exascale workflows with colmena

Computational workflows are a common class of application on supercomputers, yet the loosely coupled and heterogeneous nature of workflows often fails to take full advantage of their capabilities. We created Colmena to leverage the massive parallelism of a supercomputer by using Artificial Intelligence (AI) to learn from and adapt a workflow as it executes. Colmena allows scientists to define how their application should respond to events (e.g., task completion) as a series of cooperative agents. In this paper, we describe the design of Colmena, the challenges we overcame while deploying applications on exascale systems, and the science workflows we have enhanced through interweaving AI. The scaling challenges we discuss include developing steering strategies that maximize node utilization, introducing data fabrics that reduce communication overhead of data-intensive tasks, and implementing workflow tasks that cache costly operations between invocations. These innovations coupled with a variety of application patterns accessible through our agent-based steering model have enabled science advances in chemistry, biophysics, and materials science using different types of AI. In conclusion, our vision is that Colmena will spur creative solutions that harness AI across many domains of scientific computing.

Workflows↗

TEAMER - Field Demonstration of MarineSitu’s Marine Energy Monitoring Tools - CRADA 664 (Abstract)

In order to effectively monitor for marine life around marine energy devices and thus minimize the risk of collision, multiple sensors working in coordination and augmented with around-the-clock automated monitoring algorithms need to be installed in challenging high-energy tidal and wave environments. Such systems are often too expensive for widespread adoption, or lack sufficient sensors or smarts to enable around-the-clock, real-time monitoring without human involvement. MarineSitu has been working to tackle this problem by developing a low-cost, combined sonar and stereo camera sensor array with connected real-time AI-based algorithms for automatically detecting marine life in these marine energy suitable environments. In this TEAMER project with Pacific Northwest National Lab (PNNL), MarineSitu will be testing this novel sensor system for the first time in the high-energy tidal channel environment at PNNL’s Marine and Coastal Research Lab. Throughout this deployment, MarineSitu will be monitoring their system and running analytics on the sensor’s data in real-time. Meanwhile, PNNL Data Scientists and Ocean Engineers, will be evaluating the system’s effectiveness and ease of use both as a tool for plug-and-play environmental monitoring and novel environmental monitoring research. In doing so, the team will improve MarineSitu’s system and software, produce insightful data products, and develop novel visualizations and AI algorithms for combining and analyzing the data produced by systems like MarineSitu’s.

16 TIDAL AND WAVE POWER↗

The Vera C. Rubin Observatory Data Preview 1

We present Rubin Data Preview 1 (DP1), the first data from the National Science Foundation–Department of Energy Vera C. Rubin Observatory, comprising raw and calibrated single-epoch images, coadds, difference images, detection catalogs, and ancillary data products. DP1 is based on 1792 optical–near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera (LSSTComCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile in late 2024. DP1 covers ∼15 deg 2 distributed across seven roughly equal-sized noncontiguous fields, each independently observed in six broad photometric bands, ugrizy. The median FWHM of the point-spread function across all bands is approximately 1"14, with the sharpest images reaching about 0." 58. The 5σ point-source depths for coadded images in the deepest field, the Extended Chandra Deep Field South, are u = 24.55, g = 26.18, r = 25.96, i = 25.71, z = 25.07, and y = 23.1. Other fields are no more than 2.2 mag shallower in any band, where they have nonzero coverage. DP1 contains approximately 2.3 million distinct astrophysical objects, of which 1.6 million are extended in at least one band in coadds, and 431 solar system objects, of which 93 are new discoveries. DP1 is approximately 3.5 TB in size and is available to Vera C. Rubin Observatory data rights holders via the Rubin Science Platform, a cloud-based environment for the analysis of petascale astronomical data. While small compared to future LSST releases, its high quality and diversity of data support a broad range of early science investigations ahead of full operations in 2026.

Ground-based astronomy↗

An information-matching approach to optimal experimental design and active learning

The efficacy of mathematical models heavily depends on the quality of the training data, yet collecting sufficient data is often expensive and challenging. Many modeling applications require inferring parameters only as a means to predict other quantities of interest (QoI). Because models often contain many unidentifiable (sloppy) parameters, QoIs often depend on a relatively small number of parameter combinations. Therefore, we introduce an information-matching criterion based on the Fisher information matrix to select the most informative training data from a candidate pool. This method ensures that the selected data contain sufficient information to learn only those parameters that are needed to constrain downstream QoIs. It is formulated as a convex optimization problem, making it scalable to large models and datasets. Here, we demonstrate the effectiveness of this approach across various modeling problems in diverse scientific fields, including power systems and underwater acoustics. Finally, we use information-matching as a query function within an active learning (AL) loop for materials science applications. In all these applications, we find that a relatively small set of optimal training data can provide the necessary information for achieving precise predictions. These results are encouraging for diverse future applications, particularly AL in large machine-learning models.

Materials science↗

Baseline Climate Variables for Earth System Modelling

The Baseline Climate Variables for Earth System Modelling (ESM-BCVs) are defined as a list of 135 variables which have high utility for the evaluation and exploitation of climate simulations. The list reflects the most frequently used elements of the Coupled Model Intercomparison Project Phase 6 (CMIP6) archive. Successive phases of CMIP have supported strong results in science and substantially influence international climate policy formulation. This paper responds to both interest in exploiting CMIP data standards in a broader range of climate modelling activities and a need to achieve greater clarity about the significance and intention of variables in the CMIP Data Request. As Earth system modelling archives grow in scale and complexity, there are emerging problems associated with weak standardisation at the variable collection level. That is, there are good standards covering how specific variables should be archived, but this paper fills a gap in the standardisation of which variables should be archived. The ESM-BCV list is intended as a resource for ESM intercomparison projects (MIPs) developing requests to enable greater consistency among MIPs and as a reference for modelling centres to enhance consistency within MIPs. Provisional planning for the CMIP7 Data Request exploits the ESM-BCVs as a core element. The baseline variable list includes 98 variables which have modest or minor data volume footprints and could be generated systematically when simulations are produced and archived for exploitation by the World Climate Research Programme (WCRP) community. A further 35 variables are classed as “high volume” and are only suitable for production when the resource implications are justified.

Juckes, Martin [University of Oxford (United Kingd↗

Indicators of Global Climate Change 2023: annual update of key indicators of the state of the climate system and human influence

Intergovernmental Panel on Climate Change (IPCC) assessments are the trusted source of scientific evidence for climate negotiations taking place under the United Nations Framework Convention on Climate Change (UNFCCC). Evidence-based decision-making needs to be informed by up-to-date and timely information on key indicators of the state of the climate system and of the human influence on the global climate system. However, successive IPCC reports are published at intervals of 5–10 years, creating potential for an information gap between report cycles. We follow methods as close as possible to those used in the IPCC Sixth Assessment Report (AR6) Working Group One (WGI) report. We compile monitoring datasets to produce estimates for key climate indicators related to forcing of the climate system: emissions of greenhouse gases and short-lived climate forcers, greenhouse gas concentrations, radiative forcing, the Earth's energy imbalance, surface temperature changes, warming attributed to human activities, the remaining carbon budget, and estimates of global temperature extremes. The purpose of this effort, grounded in an open-data, open-science approach, is to make annually updated reliable global climate indicators available in the public domain. As they are traceable to IPCC report methods, they can be trusted by all parties involved in UNFCCC negotiations and help convey wider understanding of the latest knowledge of the climate system and its direction of travel. The indicators show that, for the 2014–2023 decade average, observed warming was 1.19 [1.06 to 1.30] °C, of which 1.19 [1.0 to 1.4] °C was human-induced. For the single-year average, human-induced warming reached 1.31 [1.1 to 1.7] °C in 2023 relative to 1850–1900. The best estimate is below the 2023-observed warming record of 1.43 [1.32 to 1.53] °C, indicating a substantial contribution of internal variability in the 2023 record. Human-induced warming has been increasing at a rate that is unprecedented in the instrumental record, reaching 0.26 [0.2–0.4] °C per decade over 2014–2023. This high rate of warming is caused by a combination of net greenhouse gas emissions being at a persistent high of 53±5.4 Gt CO 2 e yr -1 over the last decade, as well as reductions in the strength of aerosol cooling. Despite this, there is evidence that the rate of increase in CO 2 emissions over the last decade has slowed compared to the 2000s, and depending on societal choices, a continued series of these annual updates over the critical 2020s decade could track a change of direction for some of the indicators presented here.

54 ENVIRONMENTAL SCIENCES↗

Supporting Special Values in ZFP

This white paper outlines potential approaches to supporting special values in the ZFP numerical compressor without breaking backwards compatibility. Other than infinities and NaNs, special values are often used to indicate the absence of data, where no value is defined, for example by designating finite but extreme “fill values” as special. Such fill values are commonly used in earth system science, among other applications, but if left as is during compression lead to artifacts and loss of precision in nearby true values. Multiple candidate solutions that would allow ZFP to recognize special values are here proposed. Until such support is available, we also sketch available workarounds.

97 MATHEMATICS AND COMPUTING↗