Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Reasoning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Sizing Accuracy of Low-Cost Optical Particle Sensors Under Controlled Laboratory Conditions

Low-cost particulate matter sensors have seen increased use for monitoring at personal and local levels due to their affordability, ease of operation, and high time resolution. However, the quality of data reported by these sensors can be questionable, and a thorough evaluation of their performance is necessary. This study evaluated the particle sizing accuracy of several commonly used optical sensors, including the Alphasense optical particle counter (OPC), TSI DustTrak DRX aerosol monitor, Plantower PMS5003 sensor, and Sensirion SPS30 sensor, using laboratory-generated monodisperse particles. The OPC and DRX agreed partially with reference instruments and showed promise in detecting coarse-size particles. However, the PMS5003 and SPS30 did not correctly size fine and coarse particles. Furthermore, their reported mass distributions do not directly correspond to their number distribution. Despite these limitations, field measurements involving a dust storm period showed that the SPS30 correlated reasonably well with reference instruments for both PM2.5 and PM10, though the regression slopes differed significantly. These findings underscore the need for caution when interpreting data from low-cost optical sensors, particularly for coarse particles. Recommendations for improving the performance of these sensors are also provided.

Gautam, Prakash↗

Assessment of Extinction‐, Satellite‐, and Model‐Based Vertical Cloud Condensation Nuclei (CCN) Retrieval Methods Using Airborne CCN Measurements Over the Southern Great Plains

Abstract Accurate estimates of the vertical profile of cloud condensation nuclei (CCN) concentration are crucial to better quantify aerosol‐cloud interactions. We assessed the correlation between the vertical CCN concentrations obtained from extinction‐, satellite‐, and model‐based retrieval methods and airborne CCN concentrations collected at 0.24% supersaturation within the 3, 9, 27, and 81 km regions centered over the U.S. Department of Energy's Atmospheric Radiation Measurement User Facility Southern Great Plains (SGP) site during the spring and summer of 2016. The extinction profiles at a wavelength 355 nm were provided by the ground‐based Raman lidar. Our analysis showed moderate correlation between dry‐corrected extinction and airborne CCN data. We found the retrieved number concentration of CCN (RNCCN) method showed regression best‐fit slopes close to unity and consistent prediction errors for the majority of the data. The Lenhardt et al. (2023, https://doi.org/10.5194/amt‐16‐2037‐2023 ) method showed similar conclusions but only during spring, whereas the Mamouri and Ansmann (2016, https://doi.org/10.5194/acp‐16‐5905‐2016 ) method showed poor correlation. The Shinozuka et al. (2015, https://doi.org/10.5194/acp‐15‐7585‐2015 ) satellite‐based method exhibited reasonable agreement during summer but poor correlation during periods where both high (∼1,400 #/cm 3 ) and low (∼50 #/cm 3 ) airborne CCN concentrations were observed. The Copernicus Atmosphere Monitoring Service reanalysis modeled 3‐D CCN data set showed a moderate to weak positive correlation but performed poorly at high airborne CCN concentrations. Our analysis suggests the extinction‐based RNCCN method performed better than other methods across most observation periods under the diverse meteorological conditions observed at the SGP site.

54 ENVIRONMENTAL SCIENCES↗

Should We Conserve Entropy or Energy when Computing CAPE with Mixed-Phase Precipitation Physics?

Abstract The rapidly increasing resolution of global atmospheric reanalysis and climate model datasets necessitates finding methods for computing convective available potential energy (CAPE) both efficiently and accurately. To this end, this article compares two common methods for computing CAPE which conserve either energy or entropy. Inaccuracies in these computations arise from both physical and numerical errors. For instance, computing CAPE with entropy conserved results in physical errors from nonequilibrium phase transitions but minimizes numerical errors because solutions are analytic at each height. In contrast, computing CAPE with energy conserved avoids these physical errors, but accumulates numerical errors that are grid-resolution-dependent because the numerical integration of a differential equation is required. Analysis of CAPE computed with large databases of soundings from the tropical Amazon and midlatitude storm environments shows that physical errors from the entropy method are typically 1%–3% as large as CAPE, which is comparable to the numerical errors from conserving energy with grid spacing of 25 and 250 m using explicit first-order and second-order integration schemes, respectively. Errors in entropy-based CAPE calculations are also insensitive to vertical grid spacing, in contrast to energy-based calculations whose error strongly scales with the grid spacing. It is shown that entropy-based methods are advantageous when intercomparing datasets with differing vertical resolution because they produce accurate and reasonably fast results that are insensitive to grid resolution, whereas a second-order energy-based method is advantageous when analyzing data with a consistent vertical resolution because of its superior computational efficiency. Significance Statement Convective available potential energy (CAPE) is a measure of instability in the atmosphere that helps forecasters and researchers understand when and where thunderstorms will form. The purpose of this article is to identify the most efficient and accurate methods for computing CAPE. Two methods are considered here, one that relates to the entropy (a measure of thermodynamic disorder) of an air parcel and one that relates to the energy of an air parcel. Results indicate that the entropy method is most accurate and insensitive to the resolution of the data used for the calculation (which can vary considerably), whereas the energy method uses the least computation time.

Peters, John M.↗

WRF-ELM v1.0: a regional climate model to study land–atmosphere interactions over heterogeneous land use regions

Abstract. The Energy Exascale Earth System Model (E3SM) Land Model (ELM) is a state-of-the-art land surface model that simulates the intricate interactions between the terrestrial land surface and other components of the Earth system. Originating from the Community Land Model (CLM) version 4.5, ELM has been under active development, with added new features and functionality, including plant hydraulics, radiation–topography interaction, subsurface multiphase flow, and more explicit land use and management practices. This study integrates ELM v2.1 with the Weather Research and Forecasting (WRF; WRF-ELM) model through a modified Lightweight Infrastructure for Land Atmosphere Coupling (LILAC) framework, enabling affordable high-resolution regional modeling by leveraging ELM's innovative features alongside WRF's diverse atmospheric parameterization options. This framework includes a top-level driver for variable communication between WRF and ELM and Earth System Modeling Framework (ESMF) caps for the WRF atmospheric component and ELM workflow control, encompassing initialization, execution, and finalization. Importantly, this LILAC–ESMF framework demonstrates a more modular approach compared to previous coupling efforts between WRF and land surface models. It maintains the integrity of ELM's source code structure and facilitates the transfer of future developments in ELM to WRF-ELM. To test the ability of the coupled model to capture land–atmosphere interactions over regions with a variety of land uses and land covers, we conducted high-resolution (4 km) WRF-ELM ensemble simulations over the Great Lakes region (GLR) in the summer of 2018 and systematically compared the results against observations, reanalysis data, and WRF-CTSM (WRF coupled with the Community Terrestrial Systems Model). In general, the coupled WRF-ELM model has reasonably captured the spatial distribution of surface state variables and fluxes across the GLR, particularly over the natural vegetation areas. The evaluation results provide a baseline reference for further improvements in ELM in the regional application of high-resolution weather and climate predictions. Our work serves as an example to the model development community for expanding an advanced land surface model's capability to represent fully-coupled land–atmosphere interactions at fine spatial scales. The development and release of WRF-ELM marks a significant advancement for the ELM user community, providing opportunities for fine-scale regional representation, parameter calibration in coupled mode, and examination of new schemes with atmospheric feedback.

54 ENVIRONMENTAL SCIENCES↗

AI-Ready Semantic Infrastructure for CEBAF: From CED to PALS Knowledge Graphs

JLab and PNNL are jointly developing an AI-ready data ecosystem that exposes the Continuous Electron Beam Acceleration Facility’s (CEBAF’s) operational configuration, lattice description, and control-system channels to agentic optimization frameworks through a standards-based semantic layer. The effort integrates the existing facility-specific CEBAF Element Database (CED) with extensions of the emerging facility-agnostic Particle Accelerator Lattice Standard (PALS) to produce a knowledge graph (KG) containing coherent, machine-interpretable views of devices, signals, and regions. With this KG, CEBAF’s setpoints, readbacks, and device hierarchies become queryable using a uniform declarative graph query language (e.g., Neo4j Cypher), providing intents and inspectable semantics suitable for agentic control. The resulting graph-backed interfaces will allow autonomous agents to retrieve authoritative machine configurations, reason over device- and signal-level relationships, and execute tuning and diagnostic workflows without bespoke CEBAF-specific logic, thereby delivering a scalable pathway from operational data to trustworthy agentic accelerator tuning frameworks.

Zhang, He [Thomas Jefferson National Accelerator F↗

Data Interfaces for Automated Vehicle Services - A Municipality Perspective

As Automated Vehicle (AV) services proliferate, data sharing between AV operators and municipal agents is assuming greater importance. Information on the dynamic nature of the road system such as incidents to avoid, weather hazards (such as flooding), construction and detours, as well as active safety concerns (e.g. - riots) is important for AV operators. Such information cannot be directly sensed from a vehicle's sensor array, but instead must be communicated in a timely and trustworthy channel. Municipalities are interested in pushing this information to AV operators to support emergency response efforts, reduce traffic in construction zones, and generally improve operation of the system. Similarly, information on vehicle safety such as disengagements, as well as critical information on the use of roadway system (trips, origin and destination patterns) are important performance factors for municipalities to understand utilization and plan for appropriate infrastructure. As mobility shifts to on-demand options, the need for safe and coordinated pick-up and drop-off zones will increase (potentially reducing parking needs). For all of these reasons, communication flows between AV operators and municipalities are becoming increasingly important. This paper investigates the functions, emerging practices and protocols for sharing of such critical data, and identifies gaps in and challenges in existing practices. Additionally, case studies are used to highlight the impacts of data sharing between AV operators and municipalities.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Theorems in Service of Sound Composition, Rapid Modeling and Scalable Analysis

This project extends the state of the art in formal verification modeling with modules and automatically checkable data-sharing patterns such that component modules can retain their assurance case when composed within a larger system. For users, smaller models make reasoning easier and help to ensure they accurately reflect text specifications. For automated methods, smaller models give exponential benefits for verification algorithm execution time.

97 MATHEMATICS AND COMPUTING↗

Integrating Immersive Visualization in Molten-Salt Reactor Waste Management for Experimental Design and Planning

Molten-salt reactors (MSRs) represent a promising solution for next-generation nuclear energy, offering advantages in safety, fuel efficiency, and waste minimization. However, their liquid-fueled design presents unique challenges for spent fuel management, making post-shutdown waste characterization essential for developing effective strategies. Despite this need, there is a notable absence of visualization platforms specifically tailored to the unique characteristics and analytical requirements of MSR waste management. Existing tools in the nuclear industry are primarily designed for reactor operations or generic data exploration and lack both integration with MSR-specific multiphysics frameworks and the ability to simultaneously visualize time-dependent thermal fields, chemical composition evolution, and radiation distribution patterns. To address these limitations, this paper presents an immersive virtual reality (VR) visualization platform that processes and displays high-fidelity multiphysics simulation output from the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework in real-time, using Unity. The platform visualizes MSR waste characteristics such as nuclide decay, salt cooling, and corrosion by using Exodus II output data and running on a VR headset. It includes a user-friendly interface with features such as visibility toggling, cross-sectional slicing, and time-series animation for exploring simulation data. These capabilities support experimental design, stakeholder engagement, and public communication by making complex reactor behavior more accessible and understandable. By enhancing spatial reasoning and reducing cognitive load, this immersive environment fosters more effective communication and decision-making in MSR waste management.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING↗

CMB lensing and Ly⁢ α forest cross bispectrum from DESI’s first-year quasar sample

The squeezed cross-bispectrum B κ,Ly α between the gravitational lensing in the cosmic microwave background and the 1D Ly α forest power spectrum can constrain bias parameters and break degeneracies between σ8 and other cosmological parameters. We detect B κ,Ly ⁢α with 4.8⁢σ significance at an effective redshift z eff =2.4 using Planck PR3 lensing map and over 280,000 quasar spectra from the Dark Energy Spectroscopic Instrument’s first-year data. We test our measurement against metal contamination and foregrounds such as Galactic extinction and clusters of galaxies by deprojecting the thermal Sunyaev-Zeldovich effect. Finally, we compare our results to a tree-level perturbation theory calculation and find reasonable agreement between the model and measurement.

79 ASTRONOMY AND ASTROPHYSICS↗

An infrared, Raman, and X-ray database of battery interphase components

Further improvements to lithium-ion and emerging battery technologies can be enabled by an improved understanding of the chemistry and working mechanisms of interphases that form at electrochemically active battery interfaces. However, it is difficult to collect and interpret spectra of interphases for several reasons, including the presence of a variety of compounds. To address this challenge, we herein present a vibrational spectroscopy and X-ray diffraction data library of ten compounds that have been identified as interphase constituents in lithium-ion or emerging battery chemistries. The data library includes attenuated total reflectance Fourier transform infrared spectroscopy, Raman spectroscopy, and X-ray diffraction data, collected in inert atmospheres provided by custom sample chambers. The data library presented in this work (and online repository) simplifies access to reference data that is otherwise either diffusely spread throughout the literature or non-existent, and provides energy storage researchers streamlined access to vital interphase-relevant data that can accelerate battery research efforts.

25 ENERGY STORAGE↗

A procedure for rule extraction from a Self-Organising plasma disruption predictor for JET

In a previous paper, a Self-Organizing Map had proven to be able to identify the regions of the plasma operative space characterizing the pre-disruptive phase at JET without relying on any a priori information. One of the strengths of this disruption predictor lies in its inherent self-organization capability. The Self-Organizing Map discovers non-trivial relationships and captures the complicated interplay of device diagnostics on the internal plasma states directly from the experimental data. Moreover, the provided model allows the visualization of high-dimensional plasma parameters and facilitates easy interrogation of the model to understand the reasons behind its correlations. In this paper, an additional step is taken towards the interpretability of models for predicting disruptions by training a Decision Tree to classify the plasma states according to the interpretation provided by the Self-Organizing Map (stable or at high risk of disruptions). The Decision tree provides a set of rules which describe the transition of the plasma towards the pre-disruptive phase as visualized in the Self-Organizing Map. The obtained rules for the database explored in the study identify four regions in the map, two of which are at risk of disruption. These regions correspond to partitions of a 3D space based on the peaking factors of the core and divertor radiation, as well as the Locked Mode. The agreement between the Self-Organizing Map answers and the rules supplied by the Decision Tree is confirmed by the comparison of the performance exhibited by the two models in the prediction of disruptions.

Setzu, Samuele [Univ. of Cagliari, Monserrato, Cag↗

Announcing the Biomedical Data Translator: Initial Public Release

ABSTRACT The growing availability of biomedical data offers vast potential to improve human health, but the complexity and lack of integration of these datasets often limit their utility. To address this, the Biomedical Data Translator Consortium has developed an open‐source knowledge graph–based system—Translator—designed to integrate, harmonize, and make inferences over diverse biomedical data sources. We announce here Translator's initial public release and provide an overview of its architecture, standards, user interface, and core features. Translator employs a scalable, federated, knowledge graph framework for the integration of clinical, genomic, pharmacological, and other biomedical knowledge sources, enabling query retrieval, inference, and hypothesis generation. Translator's user interface is designed to support the exploration of knowledge relationships and the generation of insights, without requiring deep technical expertise and gradually revealing more detailed evidence, provenance, and confidence information, as needed by a given user. To demonstrate Translator's application and impact, we highlight features of the user interface in the context of three real‐world use cases: suggesting potential therapeutics for patients with rare disease; explaining the mechanism of action of a pipeline drug; and screening and validating drug candidates in a model organism. We discuss strengths and limitations of reasoning within a largely federated system and the need for rich concept modeling and deep provenance tracking. Finally, we outline future directions for enhancing Translator's functionality and expanding its data sources. Translator represents a significant step forward in making complex biomedical knowledge more accessible and actionable, aiming to accelerate translational research and improve patient care.

Research & Experimental Medicine↗

The Short Life of Upvalley Wind in a High‐Altitude Valley in the Colorado Rocky Mountains

Thermally driven upvalley (UV) wind in the upper East River Valley in the Colorado Rocky Mountains often unexpectedly stops in midmorning and reverses back to downvalley (DV) wind. We use a comprehensive observational data set for a nearly two‐year long period to analyze the wind system and boundary layer evolution in this high‐altitude valley and determine the reason for this early wind reversal. Days with short UV wind predominantly occur during the warm season when the valley floor is free of snow and the convective boundary layer (CBL) grows well above the height of the surrounding ridges. UV wind persists throughout the day only on a few days during the warm season. We link differences in valley wind evolution to wind direction at upper levels at and above ridge height and propose forced channeling mechanisms to describe coupling between valley and upper‐level wind when the CBL grows above ridge height. The frequency distribution of upper‐level wind direction is such that channeling in the DV direction is favored, which explains the predominance of days with short UV wind. The deep CBL is supported by the presence of a deep weakly stably stratified residual layer with high aerosol content, which is regularly present over the mountain range during the warm season. On days when the CBL does not grow above ridge height, for example, when the valley floor is covered by snow, thermally driven UV wind is able to persist throughout the day independent of upper‐level wind direction.

54 ENVIRONMENTAL SCIENCES↗

Feature-agnostic metabolomics for determining effective subcytotoxic doses of common pesticides in human cells

Although classical molecular biology assays can provide a measure of cellular response to chemical challenges, they rely on a single biological phenomenon to infer a broader measure of cellular metabolic response. These methods do not always afford the necessary sensitivity to answer questions of subcytotoxic effects, nor do they work for all cell types. Likewise, boutique assays such as cardiomyocyte beat rate may indirectly measure cellular metabolic response, but they too, are limited to measuring a specific biological phenomenon and are often limited to a single cell type. For these reasons, toxicological researchers need new approaches to determine metabolic changes across various doses in differing cell types, especially within the low-dose regime. Here, the data collected herein demonstrate that LC-MS/MS-based untargeted metabolomics with a feature-agnostic view of the data, combined with a suite of statistical methods including an adapted environmental threshold analysis, provides a versatile, robust, and holistic approach to directly monitoring the overall cellular metabolomic response to pesticides. When employing this method in investigating two different cell types, human cardiomyocytes and neurons, this approach revealed separate subcytotoxic metabolomic responses at doses of 0.1 and 1 µM of chlorpyrifos and carbaryl. These findings suggest that this agnostic approach to untargeted metabolomics can provide a new tool for determining effective dose by metabolomics of chemical challenges, such as pesticides, in a direct measurement of metabolomic response that is not cell type-specific or observable using traditional assays.

59 BASIC BIOLOGICAL SCIENCES↗

Participation in and Assessment of the Second DNCSH Public Workshop

The DOE/NRC Criticality Safety for Commercial-Scale HALEU Fuel Cycle and Transportation (DNCSH) project was established through the Inflation Reduction Act of 2022 (H.R. 5376) to support the US Nuclear Regulatory Commission (NRC) and industry in addressing critical experiment validation gaps that impede the licensing basis and regulatory approval of high-assay low-enriched uranium (HALEU) operations. An initial public workshop was held in February 2024 to address HALEU transportation validation gaps. The resulting call for proposals was released in April and resulted in funding for the execution and/or evaluation of 16 critical experiments. A second public workshop was held in August 2025 to address facility and operational validation gaps, precluding a second call for proposals. A list of attendees is provided in APPENDIX A, Table A-1. A total of 319 participants joined the meeting, which was hosted online via Microsoft Teams as well as in person. The slides from the meeting were uploaded online to the NRC’s Agencywide Documents Access and Management System (ADAMS). The meeting agenda is provided in Table 1-1. In preparation for the meeting, a study was performed to examine expected fissile forms for the fuel cycles of various fuel types at different stages of production and the apparent validation gaps. The resulting report, titled “Benchmark Gap Assessment for the Manufacturing of High-Assay Low-Enriched Uranium Fuels,” provided the foundation for the discussions that took place during the workshop. The discussions and the validation gaps in the report were used to develop the second call for proposals. The present report presents the feedback received before, during, and after the second workshop. All the data presented are based on voluntarily self-reported identification, opinions from workshop participants, and survey responses and are assumed to be as accurate as practically reasonable. The discussions during the workshop and the subsequent survey responses were intended to direct attention to industry-specific areas of interest and to collect feedback on the work performed to date by the DNCSH project.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Towards Next-Generation Urban Decision Support Systems through AI-Powered Construction of Scientific Ontology Using Large Language Models—A Case in Optimizing Intermodal Freight Transportation

The incorporation of Artificial Intelligence (AI) models into various optimization systems is on the rise. However, addressing complex urban and environmental management challenges often demands deep expertise in domain science and informatics. This expertise is essential for deriving data and simulation-driven insights that support informed decision-making. In this context, we investigate the potential of leveraging the pre-trained Large Language Models (LLMs) to create knowledge representations for supporting operations research. By adopting ChatGPT-4 API as the reasoning core, we outline an applied workflow that encompasses natural language processing, Methontology-based prompt tuning, and Generative Pre-trained Transformer (GPT), to automate the construction of scenario-based ontologies using existing research articles and technical manuals of urban datasets and simulations. From these ontologies, knowledge graphs can be derived using widely adopted formats and protocols, guiding various tasks towards data-informed decision support. The performance of our methodology is evaluated through a comparative analysis that contrasts our AI-generated ontology with the widely recognized pizza ontology, commonly used in tutorials for popular ontology software. We conclude with a real-world case study on optimizing the complex system of multi-modal freight transportation. Our approach advances urban decision support systems by enhancing data and metadata modeling, improving data integration and simulation coupling, and guiding the development of decision support strategies and essential software components.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

DaYu: Optimizing Distributed Scientific Workflows by Decoding Dataflow Semantics and Dynamics

The combination of ever-growing scientific datasets and distributed workflow complexity creates I/O performance bottlenecks due to data volume, velocity, and variety. Although the increasing use of descriptive data formats (e.g., HDF5, netCDF) helps organize these datasets, it also creates obscure bottlenecks due to the need to translate high level operations into file addresses and then into low-level I/O operations. To address this challenge, we introduce DaYu, a method and toolset for analyzing (a) semantic relationships between logical datasets and file addresses, (b) how dataset operations translate into I/O, and (c) the combination across entire workflows. DaYu's analysis and visualization enables identification of critical bottlenecks and reasoning about remediation. We describe our methodology and propose optimization guidelines. Evaluation on scientific workflows demonstrates up to 3.7x performance improvements in I/O time for obscure bottlenecks. The time and storage overhead for DaYu's time-ordered data is typically under 0.2% of runtime and 0.25% of data volume, respectively.

Tang, Meng↗