Search NASA⌕ Search

SEARCH · Search NASA

Results for “software ecosystem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

AmeriFlux FLUXNET-1F US-NC5 NC Butner Farm

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-NC5 NC Butner Farm. This is the FLUXNET version of the carbon flux data for the site US-NC5 NC Butner Farm produced by applying the standard ONEFlux (1F) software. Site Description - The US-NC5 flux tower is located within an 80-year-old mixed pine-hardwood forest at the Umstead Research Farm in Butner, North Carolina. The northern section of this 20-hectare Fall Lake Watershed of the Neuse River Basin in the Piedmont of North Carolina, USA. The Northern portion is currently a managed cattle farm, which is slated for expansion—necessitating forest clearing in the flux site. To establish a reference baseline, a year-long, all-season eddy covariance flux monitoring campaign will be conducted from April 2025 to March 2026. This effort aims to capture the carbon flux dynamics of the mature forest ecosystem prior to a planned land-use conversion. The site will be transitioned into a silvopasture, maintained through prescribed burning and cattle grazing to promote an open-canopy watershed structure. Flux measurements will continue after the conversion.

Sun, Ge [USDA Forest Service]↗

Data for Roebuck et al. (2025), "Differences in dissolved organic matter composition between rivers and estuaries is conserved across freshwater and saltwater coastal regions"

Dissolved organic matter (DOM) in coastal surface waters influences local water quality and is an important component of biogeochemical cycling in coastal systems, but the processes that alter DOM composition along lower reaches of rivers and estuarine waters are poorly understood. Roebuck et al. (2025) leveraged a spatially distributed community sampling effort in coastal ecosystems across two regions to identify broad spatial drivers of surface water DOM composition and identify transferable trends between saltwater and freshwater coastal systems. Samples were collected by community members from 47 locations within the mid-Atlantic and Great Lakes coastal regions.This dataset includes:* A selection of commonly reported absorbance and fluorescence peaks normalized to dissolved organic carbon concentrations* Parallel factor output from the EC1 fluorescence datasets* A selection of commonly reported absorbance and fluorescence peaks * Spectral indices output from matlab script for absorbance and fluorescence datasets* CO2sys calculations of pH changes under varying temperatures and a constant salinity, DIC, and alkalinity concentrationAll data files are plain-text CSV (comma separated value) and no special software is required to read them.

54 ENVIRONMENTAL SCIENCES↗

Coupling of high-resolution mass spectrometer and photosynthesis system for comprehensive leaf volatile metabolite profiling

Background Leaf-level biogenic volatile organic compounds (BVOCs) emissions represent a major source of organic gases in the atmosphere, influencing both climate and air quality. These emissions are strongly driven by environmental perturbations, which affect individual plant- to ecosystem-level processes. Uncovering all the BVOCs and understanding how their emissions respond to altered environmental conditions provide critical insights into vegetation-driven changes in atmospheric chemistry. We developed a tandem instrumentation setup that integrates a proton transfer reaction time-of-flight mass spectrometer (PTR-ToF-MS) with parts-per-trillion detection limits and a photosynthetic infrared gas exchange system for the untargeted survey of all the BVOCs. This novel system enables simultaneous, real-time monitoring of BVOC emissions and photosynthetic parameters at the leaf level, offering new opportunities to disentangle the physiological and environmental drivers of VOC release. Furthermore, we established the VOC Analysis and Processing Optimization Resource (VAPOR), an open-access software tool designed for rapid data post-processing and the analysis of the variability of hundreds of BVOCs. We assessed the performance of the tandem system under varying background conditions, using standard gas mixtures and a range of environmental factors. Results Blank emissions were substantially lower for major BVOCs (e.g., isoprene) compared to those observed in plant emissions. Despite this, the observation of background-level VOCs highlights the importance of routinely acquiring and accounting for blank measurements in analyses using the coupled instrumentation. Introduction of known VOC concentrations to the system demonstrated a linear response across different compounds with varying molecular compositions, indicating minimal gas loss regardless of chemical moieties within the coupled instrumentation. We applied the optimized system to investigate the physiological mechanisms driving BVOC emissions across different genotypes of poplar and pennycress. The high mass resolution capabilities of the PTR-ToF-MS, coupled with comprehensive VAPOR-driven data analysis, enabled the identification of several important BVOCs, including methanol and methanethiol; these BVOCs displayed substantial variation across pennycress genotypes and showed concentrations ~ 100–350% higher than the blank. Moreover, isoprene emissions varied significantly among poplar genotypes grown in different potting media. Conclusions Tandem instrumentation offers a powerful tool for profiling volatile molecular markers and elucidating their genetic and environmental underpinnings. This approach enhances our ability to predict BVOC emissions in response to genotype by environmental interactions and contributes to a deeper understanding of vegetation responses to environmental changes.

Biogenic volatile organic compounds↗

TEMPEST3 surface runoff water chemistry and organic matter composition

Coastal flooding, driven by storm surges and sea level rise, can mobilize organic matter (OM) via runoff, while introducing compositionally distinct OM (e.g., estuarine OM) into the system. To understand event-scale OM dynamics, we monitored source waters and surface runoff during an ecosystem-scale field manipulation experiment, TEMPEST (Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments), in June 2024. The TEMPEST experiment is part of the COMPASS-FME (Coastal Observations, Mechanisms, and Predictions Across Systems and Scales – Field, Measurements, and Experiments) project and designed to investigate biogeochemical and ecological impacts of freshwater and seawater flooding on coastal terrestrial-aquatic interface ecosystems by simulating freshwater and seawater storm events in two 2000m2 coastal upland forest plots (freshwater and brackish seawater plots). The temporal coverage of this dataset is during the TEMPESTⅢ event (June 11-13, 2024). This dataset contains: - Surface runoff discharge measured by flumes - Sensor data (specific conductivity, salinity, dissolved oxygen, and temperature) - Particle size distribution - Total suspended sediment concentrations (TSS), particulate and dissolved organic carbon (POC, DOC) concentrations, total nitrogen and total dissolved nitrogen (TN, TDN) concentrations - Bulk particulate and dissolved OM compositions (stable C and N isotopes of particulates and optical measurements of chromophoric dissolved OM) - High resolution mass spectrometry analysis data - Water isotope data All data files are plain-text CSV (comma-separated value), and no special software is required to read them.

COMPASS-FME↗

Data for Machado-Silva et al. (2024), "Short-Term Groundwater Level Fluctuations Drive Subsurface Redox Variability"

This dataset contains the analytical data reported in Machado-Silva et al. (2024) as part of the COMPASS-FME project, which seeks to advance a scalable, predictive understanding of the fundamental biogeochemical processes, ecological structure, and ecosystem dynamics that distinguish coastal terrestrial-aquatic interfaces from the purely terrestrial or aquatic systems to which they are coupled. The dataset consists of water quality parameters as well as redox potential, water content, and electrical conductivity. These data were collected in 2022 in Crane Creek (CRC), Portage River (PTR), and Old Woman Creek (OWC). Each of these sites included uplands (UP), transitions (TR), wetland-transition edge (WTE), and wetland (W) zones. The sites represent replicates of the Lake Erie terrestrial-aquatic interface under fluctuating water levels and are located in well-preserved areas with natural or restored marsh and forest cover.This dataset consists of a single data file (Machado_Silva_et_al_2024_EST_data.csv) that is in comma-separated value (CSV) format. No special software is required to read it.This dataset uses the ESS-DIVE Hydrologic Monitoring Reporting Format 1.0.

54 ENVIRONMENTAL SCIENCES↗

AmeriFlux FLUXNET-1F MX-Aog Alamos Old-Growth tropical dry forest

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site MX-Aog Alamos Old-Growth tropical dry forest. This is the FLUXNET version of the carbon flux data for the site MX-Aog Alamos Old-Growth tropical dry forest produced by applying the standard ONEFlux (1F) software. Site Description - This tower is located at a patch of remnant old growth tropical dry forest where the dominant vegetation are leguminous trees and a notorious presence of the genus Brusera. The site is at about 18 km from the municipality of Alamos Sonora Mexico within a private ranch managed by the NGO Nature Culture International. The area is also within a federally protected land named “Area de Proteccion de Flora y Fauna Sierra de Alamos Rio-Cuchijaqui" in the catalog of the “Comision Nacional de Areas Naturales Protegidas (CONANP-Mexico)”. The preserved ranch is surrounded by a mosaic of forest patches with different successional stages of tropical dry forest, and some places that are used for local agriculture and livestock. This ecosystem lies in a highly seasonal region under the influence of the North American Monsoon that brings about 70% of the rains from July to September.

Yepez, Enrico A. [Instituto Tecnologico de Sonora]↗

Variation in forest root image annotation by experts, novices, and AI

Abstract Background The manual study of root dynamics using images requires huge investments of time and resources and is prone to previously poorly quantified annotator bias. Artificial intelligence (AI) image-processing tools have been successful in overcoming limitations of manual annotation in homogeneous soils, but their efficiency and accuracy is yet to be widely tested on less homogenous, non-agricultural soil profiles, e.g., that of forests, from which data on root dynamics are key to understanding the carbon cycle. Here, we quantify variance in root length measured by human annotators with varying experience levels. We evaluate the application of a convolutional neural network (CNN) model, trained on a software accessible to researchers without a machine learning background, on a heterogeneous minirhizotron image dataset taken in a multispecies, mature, deciduous temperate forest. Results Less experienced annotators consistently identified more root length than experienced annotators. Root length annotation also varied between experienced annotators. The CNN root length results were neither precise nor accurate, taking ~ 10% of the time but significantly overestimating root length compared to expert manual annotation ( p = 0.01). The CNN net root length change results were closer to manual ( p = 0.08) but there remained substantial variation. Conclusions Manual root length annotation is contingent on the individual annotator. The only accessible CNN model cannot yet produce root data of sufficient accuracy and precision for ecological applications when applied to a complex, heterogeneous forest image dataset. A continuing evaluation and development of accessible CNNs for natural ecosystems is required.

Handy, Grace↗

Data for Myers-Pigg et al. (2026), "Short-term coastal forest responses to a hurricane-scale freshwater and saltwater flooding experiment"

Coastal upland forests are exposed to intensifying precipitation regimes and sea level rise, increasing tree mortality and transforming these coastal forests into wetland ecosystems. Despite these well-known risks, the differing degrees to which hydrological, biogeochemical, and biological components of upland forests respond to novel salinity exposure is relatively unknown. The Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experiment decouples two distinct disturbances associated with hydrological extremes: (1) flooding from heavy precipitation and (2) exposure to saline conditions from storm surge. This dataset includes data reported in Myers-Pigg et al. (2025), which analyzed data from the first TEMPEST flooding treatment in 2022. This includes: - Colored dissolved organic matter in porewaters - Soil temperature and oxygen - Groundwater temperature and chemistry - Dissolved organic carbon concentrations in porewaters - Soil-to-atmosphere CH4 and CO2 fluxes - Soil temperature, water content, and electrical conductivity - Root-influenced CH4 and CO2 flux - Tree sap flow velocity - The R analytical code and documentation about the computational environmental in which it was run (the "sessionInfo.txt" file) All data files are plain-text comma separated value (CSV) and no special software is required to read them.

54 ENVIRONMENTAL SCIENCES↗

AmeriFlux FLUXNET-1F US-MEF Manitou Experimental Forest

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-MEF Manitou Experimental Forest. This is the FLUXNET version of the carbon flux data for the site US-MEF Manitou Experimental Forest produced by applying the standard ONEFlux (1F) software. Site Description - This site was established on an existing 30-m tower structure at the Forest Service's Manitou Experimental Forest, Colorado. Manitou Experimental Forest is located approximately 38 km northwest of Colorado Springs, within the Pike National Forest. The tower was constructed and used between 2008 and 2013 by NCAR for the BEACHON project and has line power and internet connections. The tower is located in approximately 12 km^2 of ponderosa pine forest and savannah with a land use history of thinning and prescribed burning, though these disturbances have not occurred for many years. The tower is situated in a broad, flat valley that drains to the north. Early research at the site was focused mainly on grazing, while more recently it has turned broadly to the ecology and management of ponderosa pine ecosystems. In 2002, the Hayman Fire burned approximately 56,000 ha to the west and northwest of the tower location.

Frank, John [US Forest Service, Rocky Mountain Res↗

Data for Stetten et al. (2025), "Biogeochemical controls on iron speciation and cycling across upland to shoreline gradients in freshwater and estuarine coastal soils (Lake Erie and Chesapeake Bay, United States)"

Coastal environments are dynamic interfaces that mediate carbon and nutrient exchanges between terrestrial landscapes and open waters, but it is unclear how biogeochemical reactions, in particular iron (Fe) redox transformations, affect the understanding and prediction of coastal ecosystem functions. This dataset includes measurements from two freshwater sites in the Western and Central basins of Lake Erie (Ohio, United States) and two estuarine sites in the Chesapeake Bay (Maryland, United States); the analytical results were reported by Stetten et al. (2025) in Science of the Total Environment. It was produced as part of the COMPASS-FME project, which seeks to advance a scalable, predictive understanding of the fundamental biogeochemical processes, ecological structure, and ecosystem dynamics that distinguish coastal terrestrial-aquatic interfaces from the purely terrestrial or aquatic systems to which they are coupled. The sites were sampled in November 2022 (CRC), December 2022 (MSM), February 2023 (GCW), and March 2023 (OWC); site codes follow those used by Pennington et al. (2025).The dataset consists of the following soil data:- Solid data (Fe concentration, etc.)- Porewater data (sulfate, sulfide, etc.)- Linear combination fitting results of X-ray absorption near edge structure (XANES) spectra; i.e., quantitative results of the oxidation state of Fe, indicated as a proportion of pure Fe(III) and Fe(II) model compounds- Linear combination fitting results of EXAFS (extended X-ray absorption fine structure) spectra, indicated as proportion of of Fe-model compounds (illite, smectite, etc.)Each data type has a single file in comma-separated value (CSV) format. No special software is required to read it.

54 ENVIRONMENTAL SCIENCES↗

Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces

These data are from Bandopadhyay et al., "Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces". This study aims to understand the soil microbial ecology along terrestrial-aquatic interfaces of a freshwater and estuarine region and how it relates to organic matter. We analyzed soil microbial (16S rRNA gene) and organic matter (Fourier-transform ion cyclotron resonance mass spectrometry, FTICR-MS) composition from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie (freshwater) and Chesapeake Bay (estuarine) regions. This dataset includes 16S rRNA gene amplicon data (only processed file types included here) and organic matter composition from FTICR-MS data (raw and processed files included here) from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie and Chesapeake Bay regions. These sites are part of the COMPASS-FME project (https://compass.pnnl.gov/FME/COMPASSFME). File formats and software needed to access files: 16S rRNA gene amplicon data: These files follow the format reported here https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format#updates-in-v1.0.1. As per this format, there are four file types reported: 1. Taxon tables (also called sequence-by-sample or OTU (operational taxonomic unit)/ESV (exact sequence variant) tables) : available in a .txt file format and accessible using TextEdit or MS Excel. 2. Representative sequences (also called consensus sequences) : available in a .fasta format and accessible using TextEdit. 3. Sequencing metadata : available in a MS Excel workbook file format and CSV file format 4. Bioinformatic metadata : available in a MS Excel workbook file format and CSV file format FTICR-MS data: 1. Raw data converted to a processed file with intensities of the peaks in the given samples : available in a MS Excel CSV file format 2. Processed file used in analyses and visualizations (appended as icr_long_) : available in a MS Excel CSV file format 3. Metadata file for ICR features (appended as icr_meta) : available in a MS Excel CSV file format

54 ENVIRONMENTAL SCIENCES↗

Blockchain-Enabled Secure Device-to-Device Communication in Software-Defined Networking

The Internet of Things (IoT) continues to increase the demand for seamless communication among IoT devices. The rapid growth of IoT devices has led to an exponential increase in device-to-device (D2D) communication within the Software-Defined Networking (SDN), though it enables a flexible archi-tecture for managing network resources. However, traditional security models face challenges (e.g., Security, privacy, and trust) in addressing the dynamic and decentralized nature of these communications. Despite of these challenges, this paper proposes a novel approach that leverages blockchain technology to enhance the security, privacy, and trustworthiness of D2D communication within an SDN environment. The proposed approach integrates blockchain nodes in sDN components to establish a decentralized ledger for transparent and verifiable records. Smart contracts enforce authentication rules to ensure that only authenticated devices can access the network and engage in transactions securely. It also automates the security policies to ensure temper resistance execution using the cryptographic mechanism for data integrity and authentic communication. The Implementation of the proposed algorithms validates the resilience of the proposed approach against cyberattacks. Overall, the proposed approach enables efficient and secure D2D communication for resilient SDN infrastructure in IoT ecosystems.

Das, Debashis↗

Integrating ORNL’s HPC and Neutron Facilities with a Performance-Portable CPU/GPU Ecosystem

We explore the development of a performance-portable CPU/GPU ecosystem to integrate two of the US Department of Energy’s (DOE’s) largest scientific instruments, the Oak Ridge Leadership Computing facility and the Spallation Neutron Source (SNS), both of which are housed at Oak Ridge National Laboratory. We select a relevant data reduction workflow use-case to obtain the differential scattering cross-section from data collected by SNS’s CORELLI and TOPAZ instruments. We compare the current CPU-only production implementation using the Garnet Python multiprocess package based on the Mantid C++ framework against our proposed CPU/GPU implementation that uses the LLVM-based, just-in-time Julia scientific language and the JACC.jl performance-portable package. Two proxy apps were developed: (i) an app for extracting relevant Mantid kernels (MDNorm) in C++ and (ii) the Julia MiniVATES.jl miniapp. We present performance results for NVIDIA A100 and AMD MI100 GPUs and AMD EPYC 7513 and 7662 CPUs. The results provide insights for future generations of data reduction software that can embrace performance portability for an integrated research infrastructure across DOE’s experimental and computational facilities.

Hahn, Steven↗

Fungal mat growth and leaf colonization at the TRACE warming experiment, Mar - Aug 2024, Luquillo, Puerto Rico

This data package contains processed measurements on the growth of litter mat-forming fungi and the time to leaf colonization at the Tropical Responses to Altered Climate Experiment (TRACE). Located near the Sabana Field Research Station in Luquillo, Puerto Rico, the TRACE site is located in a mature, closed-canopy tropical rainforest within the Luquillo Experimental Forest (LEF). These data quantify fungal mat growth and the time to leaf colonization of fungi species Gymnopus johnstonii and Marasmius crinis-equi. The experiment was conducted in ambient (control) and experimentally warmed plots (4°C above ambient) during spring and summer periods to assess how litter mat-forming fungi respond to a range of environmental conditions of tropical wet forests. The data files include tables of relative fungal mat growth rates, time to leaf colonization, averages of soil temperature (°C), and number of dry days before leaf attachment. The data are stored in comma-separated values (CSV) format and viewable with any text editor, spreadsheet, or statistical software (e.g., R, Python, Excel). Associated metadata describe plot identifiers, measurement descriptions, and processing steps.

Agaric fungi↗

Five Years of Dissolved Oxygen, Temperature, Salinity, Depth, Weather Data from a Transitioning Wetland at Beaver Creek, Washington, USA

Groundwater dissolved oxygen (DO) variability in coastal system remains poorly understood despite its importance for biogeochemical cycling and ecosystem modeling. Here we investigate the temporal variability in groundwater DO and its hydro-climatic drivers across hourly to seasonal timescales in a transitioning wetland at Beaver Creek, Washington, USA. The site is transitioning from a freshwater forest to a brackish tidal wetland following removal of a barrier in 2014 that prevented tides from accessing the freshwater creek. By utilizing novel optical dissolved oxygen instrumentation (Opti O2, LLC) we obtained continuous, high-frequency (5-minute), in-situ measurements of DO from the flood-plain from June 26th, 2019 through September 30th, 2024. This 63 month dataset is comprised of groundwater dissolved oxygen, temperature, water level and salinity timeseries from the floodplain. This dataset also includes rainfall, air pressure, air temperature, and solar radiation data collected with a co-located Campbell ClimaVUE50 weather sensor. All data is contained within a single csv (2019-06-26 to 2024-09-30 Beaver Creek DO, saln, BGS, temp, weather.csv) that can easily be viewed either using software such as Excel or using any text editor.

54 ENVIRONMENTAL SCIENCES↗

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Throughput Estimation of Data Transport Networks From Digital Twin Measurements

Digital twins of networked infrastructures, known as Virtual Infrastructure Twins (VITs), are increasingly used for software development, pre-deployment testing, and design space exploration. While VITs avoid the costs and potential disruptions associated with experiments on operational networks, their throughput measurements are typically not sufficiently accurate for performance profiling of wide-area networks that they emulate. Here, machine learning (ML) methods are developed to transform these inaccurate VIT network throughput measurements to closely match in peak and overall profile of those from a physical testbed or production network. First, a micro kernel network reflecting a physical network is utilized to collect one-time measurements on a host to support this ML transformation. Then, a generic multi-modal ML method is developed to learn a map that transforms measurements from subsequent VITs on the same host to match past, current and follow-on testbed and cloud networks. ML generalization equations are derived to establish its correctness and probabilistically guarantee its generalization accuracy. Experimental results are presented for a variety of VIT hosts with target testbed and cloud networks; they include a case study of a four-site science ecosystem wherein inaccurate convex VIT measurement profiles are transformed into accurate concave profiles of target networks.

97 MATHEMATICS AND COMPUTING↗