Search NASA⌕ Search

SEARCH · Search NASA

Results for “Multivariate outputs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

High‐Resolution National‐Scale Water Modeling Is Enhanced by Multiscale Differentiable Physics‐Informed Machine Learning

Abstract The National Water Model (NWM) is a key tool for flood forecasting, planning, and water management. Key challenges facing the NWM include calibration and parameter regionalization when confronted with big data. We present two novel versions of high‐resolution (∼37 km 2 ) differentiable models (a type of hybrid model): one with implicit, unit‐hydrograph‐style routing and another with explicit Muskingum‐Cunge routing in the river network. The former predicts streamflow at basin outlets whereas the latter presents a discretized product that seamlessly covers rivers in the conterminous United States (CONUS). Both versions use neural networks to provide a multiscale parameterization and process‐based equations to provide a structural backbone, which were trained simultaneously (“end‐to‐end”) on 2,807 basins across the CONUS and evaluated on 4,997 basins. Both versions show great potential to elevate future NWM performance for extensively calibrated as well as ungauged sites: the median daily Nash‐Sutcliffe efficiency of all 4,997 basins is improved to around 0.68 from 0.48 of NWM3.0. As they resolve spatial heterogeneity, both versions greatly improved simulations in the western CONUS and also in the Prairie Pothole Region, a long‐standing modeling challenge. The Muskingum‐Cunge version further improved performance for basins >10,000 km 2 . Overall, our results show how neural‐network‐based parameterizations can improve NWM performance for providing operational flood predictions while maintaining interpretability and multivariate outputs. The modeling system supports the Basic Model Interface (BMI), which allows seamless integration with the next‐generation NWM. We also provide a CONUS‐scale hydrologic data set for further evaluation and use.

Song, Yalan [Civil and Environmental Engineering T↗

STITCHES: a Python package to amalgamate existing Earth system model output into new scenario realizations

Understanding the interaction between humans and the Earth system is a computationally daunting task, with many possible approaches depending on resources available and questions of interest. For example, state-of-the-art impact models require decade-long time series of relatively high frequency, spatially resolved and often multiple variables representing climatic impact-drivers (Ruane et al., 2022). Most commonly these are derived from the outputs of detailed, computationally expensive Earth System Models (ESMs) run according to a standard, limited set of future scenarios, the latest being the SSP-RCPs run under CMIP6/ScenarioMIP (Eyring et al., 2016; O’Neill et al., 2016). At the time of writing, O’Neill et al. (2016) has been cited more than 1750 times and Eyring et al. (2016) more than 5000 times, highlighting the broad, general applications of this data. Often, however, impact modeling seeks to explore new scenarios that were not part of the ScenarioMIP protocol, and/or needs a larger set of initial condition ensemble members than are typically available to quantify the effects of ESM internal variability. In addition, the recognition that the human and Earth systems are fundamentally intertwined, and may feature potentially significant feedback loops, is making integrated, simultaneous modeling of the coupled human-Earth system increasingly necessary, if computationally challenging with most existing tools (Thornton et al., 2017). For impact modelers, climate model emulators can be the answer to meet both the needs of: 1) creating realizations for novel scenarios and 2) achieving a simplified, computationally tractable representation of ESM behavior in a coupled human-Earth system modeling framework. We proposed a new, comprehensive approach to such emulation of gridded, multivariate ESM outputs for novel scenarios without the computational cost of a full ESM, STITCHES (Tebaldi et al., 2022). The approach outlined in Tebaldi et al. (2022) should be extensible to future CMIP eras, although the STITCHES software at present is strictly focused on CMIP6/ScenarioMIP data hosted on Pangeo (https://gallery.pangeo.io/repos/pangeo-gallery/cmip6/). The corresponding STITCHES Python package uses existing archives of ESMs’ scenario experiments from CMIP6/ScenarioMIP to construct gridded, multivariate realizations of new scenarios provided by reduced complexity climate models (Hartin et al., 2015; Meinshausen et al., 2011; Smith et al., 2018), or to enrich existing initial condition ensembles. Its output provides the same characteristics as the emulated ESM output: multivariate (spanning potentially all variables that the ESM has saved), spatially resolved (down to the native grid of the ESM), and preserving the same high frequency as the original data. A new realization of multiple variables can be generated on the order of minutes with STITCHES, rather than the hours or sometimes days that ESMs require.

97 MATHEMATICS AND COMPUTING↗

Kernel-based global sensitivity analysis obtained from a single data set

Results from global sensitivity analysis (GSA) often guide the understanding of complicated input–output systems. Kernel-based GSA methods have recently been proposed for their capability of treating a broad scope of complex systems. In this paper, we develop a new set of kernel GSA tools when only a single set of input–output data is available. Three key advances are made: (1) A new numerical estimator is proposed that demonstrates an empirical improvement over previous procedures. (2) A computational method for generating inner statistical functions from a single data set is presented. (3) A theoretical extension is made to define conditional sensitivity indices, which reveal the degree that the inputs carry shared information about the output when inherent input–input correlations are present. Utilizing these conditional sensitivity indices, a decomposition is derived for the output uncertainty based on what is called the optimal learning sequence of the input variables, which remains consistent when correlations exist between the input variables. Further, while these advances cover a range of GSA subjects, a common single data set numerical solution is provided by a technique known as the conditional mean embedding of distributions. The new methodology is implemented on benchmark systems to demonstrate the provided insights.

42 ENGINEERING↗

National serosurvey and risk mapping reveal widespread distribution of Coxiella burnetii in Kenya

Coxiella burnetii, the causative agent of Q fever, is an emerging pathogen that has the potential to cause severe chronic infections in animals and humans worldwide. The detrimental impact on public health is projected to be higher in the low- and middle-income countries given their lower capacity to sustain effective surveillance and response measures. We implemented a national serosurvey of cattle in Kenya to map the spatial distribution of the pathogen. The study used serum samples that were collected from randomly selected cattle in different ago-ecological zones across the country. These samples were screened for the pathogen using PrioCHECK Ruminant Q Fever AB Plate ELISA kit. The laboratory findings were analyzed using INLA package to identify risk factors for C. burnetii exposure from herd- and animal-level factors, area, and bioclimatic datasets accessed from online databases. A total of 6,593 cattle were recruited for the study; of these, 7.9% (95% CI; 7.2–8.5) were seropositive. Outputs from the multivariable analysis revealed that the animal age and some of the geographical variables including wind speed, area under shrubs and “petric calcisols” type of soil were significantly associated with C. burnetii seropositivity. Being a calf, weaner or subadult was associated with lower odds of exposure compared to being an adult by 0.24 (credibility interval: 2.5% and 97.5%), 0.41 (0.30–0.55) and 0.51 (0.38–0.69), respectively. In addition, a unit increase in the wind speed increased the odds of C. burnetii seropositivity by 1.27 (1.05–1.52) while an increase on the land area under shrubs was associated with lower odds of exposure (0.67 [0.47–0.69]). The effect of petric calcisols was non-linear; an increase of the land area with this soil type was associated with an exponential increase in C. burnetii seropositivity. This study provides new data on C. burnetii seroprevalence, information of its risk factors and a prevalence map that can be used for C. burnetii risk surveillance and control. The identification of environmental risk factors for C. burnetii exposure, and the increasing awareness of the zoonotic potential of the pathogen, calls for the need to enhance the existing collaborations for the surveillance and control of C. burnetii in line with the One Health framework. The evidence generated on the potential role of environmental factors can also be used to design nature-based interventions, such as replacement of vegetation in denuded areas, to reduce potential for the aerosolization of the pathogen. Livestock vaccination in the hotspots would also reduce animal infections and hence the contamination of the environment.

60 APPLIED LIFE SCIENCES↗

Probabilistic projections of the Amery Ice Shelf catchment, Antarctica, under conditions of high ice-shelf basal melt

Abstract. Antarctica's Lambert Glacier drains about one-sixth of the ice from the East Antarctic Ice Sheet and is considered stable due to the strong buttressing provided by the Amery Ice Shelf. While previous projections of the sea-level contribution from this sector of the ice sheet have predicted significant mass loss only with near-complete removal of the ice shelf, the ocean warming necessary for this was deemed unlikely. Recent climate projections through 2300 indicate that sufficient ocean warming is a distinct possibility after 2100. This work explores the impact of parametric uncertainty on projections of the response of the Lambert–Amery system (hereafter “the Amery sector”) to abrupt ocean warming through Bayesian calibration of a perturbed-parameter ice-sheet model ensemble. We address the computational cost of uncertainty quantification for ice-sheet model projections via statistical emulation, which employs surrogate models for fast and inexpensive parameter space exploration while retaining critical features of the high-fidelity simulations. To this end, we build Gaussian process (GP) emulators from simulations of the Amery sector at a medium resolution (4–20 km mesh) using the Model for Prediction Across Scales (MPAS)-Albany Land Ice (MALI) model. We consider six input parameters that control basal friction, ice stiffness, calving, and ice-shelf basal melting. From these, we generate 200 perturbed input parameter initializations using space filling Sobol sampling. For our end-to-end probabilistic modeling workflow, we first train emulators on the simulation ensemble and then calibrate the input parameters using observations of the mass balance, grounding line movement, and calving front movement with priors assigned via expert knowledge. Next, we use MALI to project a subset of simulations to 2300 using ocean and atmosphere forcings from a climate model for both low- and high-greenhouse-gas-emission scenarios. From these simulation outputs, we build multivariate emulators by combining GP regression with principal component dimension reduction to emulate multivariate sea-level contribution time series data from the MALI simulations. We then use these emulators to propagate uncertainty from model input parameters to predictions of glacier mass loss through 2300, demonstrating that the calibrated posterior distributions have both greater mass loss and reduced variance compared to the uncalibrated prior distributions. Parametric uncertainty is large enough through about 2130 that the two projections under different emission scenarios are indistinguishable from one another. However, after rapid ocean warming in the first half of the 22nd century, the projections become statistically distinct within decades. Overall, this study demonstrates an efficient Bayesian calibration and uncertainty propagation workflow for ice-sheet model projections and identifies the potential for large sea-level rise contributions from the Amery sector of the Antarctic Ice Sheet after 2100 under high-greenhouse-gas-emission scenarios.

54 ENVIRONMENTAL SCIENCES↗

ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation

Modern climate projections lack adequate spatial and temporal resolution due to computational constraints. A consequence is inaccurate and imprecise predictions of critical processes such as storms. Hybrid methods that combine physics with machine learning (ML) have introduced a new generation of higher fidelity climate simulators that can sidestep Moore’s Law by outsourcing compute-hungry, short, high-resolution simulations to ML emulators. However, this hybrid ML-physics simulation approach requires domain-specific treatment and has been inaccessible to MLexperts because of lack of training data and relevant, easy-to-use workflows. Wepresent ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state. The dataset is global in coverage, spans multiple years at high sampling frequency, and is designed such that resulting emulators are compatible with downstream coupling into operational climate simulators. We implement a range of deterministic and stochastic regression baselines to highlight the ML challenges and their scoring. The data (https://huggingface.co/datasets/LEAP/ClimSim_high-res2) and code(https://leap-stc.github.io/ClimSim)arereleasedopenlytosupport the development of hybrid ML-physics and high-fidelity climate simulations for the benefit of science and society.

artificial intelligence, machine learning↗

Application of Partial Least Squares Approaches to Pyroprocessing ER Data

Multivariate approaches show promise for application to process monitoring for safeguards of pyroprocessing. Past MPACT work explored the application of Principal Component Analysis (PCA) to detect off-normal conditions in pyroprocessing electrorefiner (ER) data from in the Hot Fuel Examination Facility (HFEF) at Idaho National Laboratory (INL) known as the Scalable Pyrochemical Recycling testbed (SPyRe) ER. PCA, however, does not consider the output variables. In FY24, multivariate analysis was extended from PCA to Partial Least Squares (PLS) analysis. PLS maximizes the variance between both the input signals and output variables. In the case of this work, PLS was applied in two different manners: Predictive PLS and Discriminant PLS. Predictive PLS maximizes the covariance between the process variables of the ER and the measured U concentration from in-situ voltammetry. Discriminant PLS maximizes the covariance between the process variables and a set of training process “states” such as known off-normal conditions. By projecting into the latent variable space in PLS, the process variables can be regressed onto the outputs and predictions can be made for new data sets. In this work, by applying predictive PLS, a penalized non-linear PLS approach was able to make predictions of concentration based on test and training data and detect when operations were off-normal. However, the predictive PLS does not classify the signals to which off-normal operations are attributable. Discriminant PLS can be used to classify off-normal operations but is inadequate to properly classify specific off-normal classes like power supply faults when the Discriminant PLS model is only specifically trained to detect that off-normal class. When all faults are trained against the observation data, all three operational classes are accurately classified and distinguished. Thus, future application of latent variable techniques should not select any given method, but should use a mixture of PCA, Predictive PLS, and Discriminant PLS.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Learning likelihood ratios with neural network classifiers

The likelihood ratio is a crucial quantity for statistical inference in science that enables hypothesis testing, construction of confidence intervals, reweighting of distributions, and more. Many modern scientific applications, however, make use of data- or simulation-driven models for which computing the likelihood ratio can be very difficult or even impossible. By applying the so-called “likelihood ratio trick,” approximations of the likelihood ratio may be computed using clever parametrizations of neural network-based classifiers. A number of different neural network setups can be defined to satisfy this procedure, each with varying performance in approximating the likelihood ratio when using finite training data. We present a series of empirical studies detailing the performance of several common loss functionals and parametrizations of the classifier output in approximating the likelihood ratio of two univariate and multivariate Gaussian distributions as well as simulated high-energy particle physics datasets.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

STSR-INR: Spatiotemporal super-resolution for multivariate time-varying volumetric data via implicit neural representation

Implicit neural representation (INR) has surfaced as a promising direction for solving different scientific visualization tasks due to its continuous representation and flexible input and output settings. We present STSR-INR, an INR solution for generating simultaneous spatiotemporal super-resolution for multivariate time-varying volumetric data. Inheriting the benefits of the INR-based approach, STSR-INR supports unsupervised learning and permits data upscaling with arbitrary spatial and temporal scale factors. Unlike existing GAN- or INR-based super-resolution methods, STSR-INR focuses on tackling variables or ensembles and enabling joint training across datasets of various spatiotemporal resolutions. Here we achieve this capability via a variable embedding scheme that learns latent vectors for different variables. In conjunction with a modulated structure in the network design, we employ a variational auto-decoder to optimize the learnable latent vectors to enable latent-space interpolation. To combat the slow training of INR, we leverage a multi-head strategy to improve training and inference speed with significant speedup. We demonstrate the effectiveness of STSR-INR with multiple scalar field datasets and compare it with conventional tricubic+linear interpolation and state-of-the-art deep-learning-based solutions (STNet and CoordNet).

97 MATHEMATICS AND COMPUTING↗

Downscaled CMIP5 projections of physical fire risk understate historical trends

Reliable projections of wildfire risk are important for multi-sector impacts analysis. Statistically downscaled and bias-corrected Earth system model ensemble products are routinely used to analyze regional physical wildfire risk, but evaluations of historical observed trends and variability are lacking. Here, we evaluate physical fire risk over the western United States using the Canadian Forest Fire Weather Index (FWI) by comparing model outputs from the Coupled Model Intercomparison Project Phase 5 (CMIP5), statistically downscaled via the Multivariate Adaptive Constructed Analogs (MACA) approach, against the observational target dataset gridMET, a gridded high-resolution surface meteorological product. We analyze multidecadal trends and interannual variability in seasonal average FWI for the historical period and future projections under two emissions scenarios, and we compare MACA-CMIP5 ensemble results with a simple time series model that generates historical and future projections of seasonal FWI based on bootstrapping observed historical trends and variability. Our findings indicate that MACA-CMIP5 accurately captures the magnitude and spatial patterns of seasonally averaged FWI but tends to underestimate historical decadal trends. We show that future increases in fire risk may be underestimated relative to the simple time series model that projects historical variability into the future. We also highlight that model biases in relative humidity contribute significantly to model-data differences. Our results underscore the importance of historical hindcasting exercises for informing broader multi-sector applications.

FWI↗

Explainable AI for Multivariate Time Series Pattern Exploration: Latent Space Visual Analytics With Temporal Fusion Transformer and Variational Autoencoders in Power Grid Event Diagnosis

Detecting and analyzing complex patterns in multivariate time-series data is crucial for decision-making in urban and environmental system operations. However, challenges arise from the high dimensionality, intricate complexity, and interconnected nature of complex patterns, which hinder the understanding of their underlying physical processes. Existing AI methods often face limitations in interpretability, computational efficiency, and scalability, reducing their applicability in real-world scenarios. This paper proposes a novel visual analytics framework that integrates two generative AI models, Temporal Fusion Transformer (TFT) and Variational Autoencoders (VAEs), to reduce complex patterns into lower-dimensional latent spaces and visualize them in 2D using dimensionality reduction techniques such as PCA, t-SNE, and UMAP with DBSCAN. These visualizations, presented through coordinated and interactive views and tailored glyphs, enable intuitive exploration of complex multivariate temporal patterns, identifying patterns’ similarities and uncover their potential correlations for a better interpretability of the AI outputs. The framework is demonstrated through a case study on power grid signal data, where it identifies multi-label grid event signatures, including faults and anomalies with diverse root causes. Additionally, novel metrics and visualizations are introduced to validate the models and assess the performance, efficiency, and consistency of latent maps generated by VAE, which have been utilized in prior studies for latent space cartography and used as a benchmark in this study, and the emerging TFT architecture under various configurations. These analyses provide actionable insights for model parameter tuning and reliability improvements. Comparative results highlight that TFT achieves shorter run times and superior scalability to diverse time-series data shapes compared to VAE. This work advances fault diagnosis in multivariate time series, fostering explainable AI to support critical system operations.

Explainable AI↗

A New Simple-to-Configure Self-Perturbing Multivariable Extremum-Seeking Controller

This paper presents a new stochastic relay-based extremum-seeking controller (ESC) for multi-input-single-output (MISO) systems. The algorithm was developed with the goal of simplifying configuration to enable easier deployment to real-world problems. A solution is developed first for a static map and then adapted for a general class of dynamic systems. The number of configurable parameters is one per input channel for the static case and only one additional parameter is needed for the dynamic version. The problem of gradient identifiability is solved via the use of stochastic relay gains and a simple stability proof for the static case is presented. Simulation tests demonstrate the performance of the strategy for optimizing both static and dynamic systems.

Salsbury, Timothy [BATTELLE (PACIFIC NW LAB)]↗

Machine learning-based ethylene and carbon monoxide estimation, real-time optimization, and multivariable feedback control of an experimental electrochemical reactor

Electrochemical reduction of CO 2 gas is a novel CO 2 utilization technique that has the potential to mitigate the global climate crisis caused by anthropogenic CO 2 emissions, and enable the large-scale storage of energy generated from renewable sources in the form of carbon-based chemicals and fuels. However, due to the complexity of the electrochemical reactions, the explicit first-principles models for CO2 reduction are not available yet, and there has been a limited effort to develop process modeling, optimization and control of CO 2 electrochemical reactors. To this end, a rotating cylinder electrode (RCE) reactor has been constructed at UCLA to understand the mass transfer and reaction kinetics effects separately on the productivity. In the RCE reactor, the applied potential strongly influences the reaction energetics and the electrode rotation speed affects the hydrodynamic boundary layer and modifies the film mass transfer coefficient, which involves convective and diffusive transport. Further, the present work aims to develop a multi-input multi-output (MIMO) control scheme for the RCE reactor that integrates techniques from artificial and recurrent neural network modeling, nonlinear optimization, and process controller design. Specifically, production rates of two products from the experimental reactor, ethylene and carbon monoxide, are controlled by manipulating two inputs, applied potential and catalyst rotation speed. Process dynamics and controllability are analyzed, a feedback control strategy is designed and the controllers are tuned accordingly. The experimental electrochemical cell is employed to gather data for process modeling and implement the multivariable control system. Finally, the experimental results are presented which demonstrate excellent closed-loop performance by the control system and regulation of the outputs at three different set-points including an economically-optimal set-point.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parallel hybrid quantum-classical machine learning for kernelized time-series classification

Supervised time-series classification garners widespread interest because of its applicability throughout a broad application domain including finance, astronomy, biosensors, and many others. Here, in this work, we tackle this problem with hybrid quantum-classical machine learning, deducing pairwise temporal relationships between time-series instances using a timeseries Hamiltonian kernel (TSHK). A TSHK is constructed with a sum of inner products generated by quantum states evolved using a parameterized time evolution operator. This sum is then optimally weighted using techniques derived from multiple kernel learning. Because we treat the kernel weighting step as a differentiable convex optimization problem, our method can be regarded as an end-to-end learnable hybrid quantum-classical-convex neural network, or QCC-net, whose output is a data set-generalized kernel function suitable for use in any kernelized machine learning technique such as the support vector machine (SVM). Using our TSHK as input to a SVM, we classify univariate and multivariate time-series using quantum circuit simulators and demonstrate the efficient parallel deployment of the algorithm to 127-qubit superconducting quantum processors using quantum multi-programming.

97 MATHEMATICS AND COMPUTING↗

Mass Spectral Imaging to Map Plant–Microbe Interactions

Plant–microbe interactions are of rising interest in plant sustainability, biomass production, plant biology, and systems biology. These interactions have been a challenge to detect until recent advancements in mass spectrometry imaging. Plants and microbes interact in four main regions within the plant, the rhizosphere, endosphere, phyllosphere, and spermosphere. This mini review covers the challenges within investigations of plant and microbe interactions. We highlight the importance of sample preparation and comparisons among time-of-flight secondary ion mass spectroscopy (ToF-SIMS), matrix-assisted laser desorption/ionization (MALDI), laser desorption ionization (LDI/LDPI), and desorption electrospray ionization (DESI) techniques used for the analysis of these interactions. Using mass spectral imaging (MSI) to study plants and microbes offers advantages in understanding microbe and host interactions at the molecular level with single-cell and community communication information. More research utilizing MSI has emerged in the past several years. We first introduce the principles of major MSI techniques that have been employed in the research of microorganisms. An overview of proper sample preparation methods is offered as a prerequisite for successful MSI analysis. Traditionally, dried or cryogenically prepared, frozen samples have been used; however, they do not provide a true representation of the bacterial biofilms compared to living cell analysis and chemical imaging. New developments such as microfluidic devices that can be used under a vacuum are highly desirable for the application of MSI techniques, such as ToF-SIMS, because they have a subcellular spatial resolution to map and image plant and microbe interactions, including the potential to elucidate metabolic pathways and cell-to-cell interactions. Promising results due to recent MSI advancements in the past five years are selected and highlighted. The latest developments utilizing machine learning are captured as an important outlook for maximal output using MSI to study microorganisms.

59 BASIC BIOLOGICAL SCIENCES↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Visual Analytics of Multivariate Networks With Representation Learning and Composite Variable Construction

Multivariate networks are commonly found in real-world data-driven applications. Uncovering and understanding the relations of interest in multivariate networks is not a trivial task. This article presents a visual analytics workflow for studying multivariate networks to extract associations between different structural and semantic characteristics of the networks (e.g., what are the combinations of attributes largely relating to the density of a social network?). The workflow consists of a neural-network-based learning phase to classify the data based on the chosen input and output attributes, a dimensionality reduction and optimization phase to produce a simplified set of results for examination, and finally an interpreting phase conducted by the user through an interactive visualization interface. A key part of our design is a composite variable construction step that remodels nonlinear features obtained by neural networks into linear features that are intuitive to interpret. We demonstrate the capabilities of this workflow with multiple case studies on networks derived from social media usage and also evaluate the workflow with qualitative feedback from experts.

97 MATHEMATICS AND COMPUTING↗

Feedforward-feedback ammonia control at a water resource recovery facility based on a digital twin with hybrid model

Ammonia-based aeration control (ABAC) at full-scale Water Resource Recovery Facilities (WRRFs) can be challenged by diurnal loading and transport delays. This work addressed these challenges using a hybrid feedforward–feedback controller built on Activated Sludge Model 1 (ASM1), marking the first full-scale deployment to pair a mechanistic feedforward core with data-driven corrections. The objectives were to improve ammonia setpoint tracking, assess performance of the mechanistic model when enhanced with data-driven corrections, and document full-scale operation. The hybrid model incorporates two data-driven components: (1) a Mechanistic Error Forecasting Engine (MEFE), consisting of a multivariate linear regressor and a long short-term memory (LSTM) ensemble. Defying expectations, low-parameter models outperformed more complex alternatives, reducing the mechanistic error by 71%. (2) A Residual Oscillation Forecasting Engine (ROFE), based on Fast Fourier Transform, reduced the remaining error by another 35%. Two proportional–integral (PI) feedback loops further (i) trim the feedforward output and (ii) eliminate residual controller error in the final aerobic zone. In full-scale operation, the controller reduced mean-squared error (MSE) by 94% over the baseline and produced more stable dissolved oxygen (DO) setpoints. Overall, it was proven that layering multi-timescale data-driven models on a mechanistic core can yield reliable ABAC performance at WRRFs.

54 ENVIRONMENTAL SCIENCES↗