Search NASA⌕ Search

SEARCH · Search NASA

Results for “Variability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

The rate of extreme coronal line emitting galaxies in the Sloan Digital Sky Survey and their relation to tidal disruption events

High-ionization iron coronal lines (CLs) are a rare phenomenon observed in galaxy and quasi-stellar object spectra that are thought to be created by high-energy emission from active galactic nuclei and certain types of transients. In cases known as extreme coronal line emitting galaxies (ECLEs), these CLs are strong and fade away on a time-scale of years. The most likely progenitors of these variable CLs are tidal disruption events (TDEs), which produce sufficient high-energy emission to create and sustain the CLs over these time-scales. To test the possible connection between ECLEs and TDEs, we present the most complete variable ECLE rate calculation to date and compare the results to TDE rates from the literature. To achieve this, we search for ECLEs in the Sloan Digital Sky Survey (SDSS). We detect sufficiently strong CLs in 16 galaxies, more than doubling the number previously found in SDSS. Using follow-up spectra from the Dark Energy Spectroscopic Instrument and Gemini Multi-Object Spectrograph, Wide-field Infrared Survey Explorer mid-infrared observations, and Liverpool Telescope optical photometry, we find that none of the nine new ECLEs evolve in a manner consistent with that of the five previously discovered variable ECLEs. Using this sample of five variable ECLEs, we calculate the galaxy-normalized rate of variable ECLEs in SDSS to be $R_\mathrm{G}=3.6~^{+2.6}_{-1.8}~(\mathrm{statistical})~^{+5.1}_{-0.0}~(\mathrm{systematic})\times 10^{-6}~\mathrm{galaxy}^{-1}~\mathrm{yr}^{-1}$. The mass-normalized rate is $R_\mathrm{M}=3.1~^{+2.3}_{-1.5}~(\mathrm{statistical})~^{+4.4}_{-0.0}~(\mathrm{systematic})\times 10^{-17}~\mathrm{M_\odot ^{-1}}~\mathrm{yr}^{-1}$ and the volumetric rate is $R_\mathrm{V}=7~^{+20}_{-5}~(\mathrm{statistical})~^{+10}_{-0.0}~(\mathrm{systematic})\times 10^{-9}~\mathrm{Mpc}^{-3}~\mathrm{yr}^{-1}$. Our rates are one to two orders of magnitude lower than TDE rates from the literature, which suggests that only 10–40 per cent of all TDEs produce variable ECLEs. Additional uncertainties in the rates arising from the structure of the interstellar medium have yet to be included.

79 ASTRONOMY AND ASTROPHYSICS↗

Fast Active-Set Thresholding Method for Nonnegative Least Squares

Nonnegative Least Squares (NNLS) is a fundamental constrained optimization problem encountered in many applications such as image deblurring, signal processing, nonnegative matrix factorization, magnetic microscopy, and hyperspectral imaging. Active-set based methods are a common class of algorithms for solving NNLS which identify the optimal variable set of the NNLS solution. They do so by iteratively solving a series of unconstrained least squares problems, identifying which variables violate the nonnegativity constraints, and then swapping variables in/out of consideration until the optimal set of variables is found. Several variations improving upon this method exist in the literature. In this work, we propose an active-set swap heuristic which further improves upon existing active-set based methods for NNLS. Our optimizations are based upon adding multiple variables to the passive set within a threshold of the smallest gradient value and removing variables within a similar threshold of the closest boundary constraint. We leverage these optimizations to yield a Fast Active-Set Thresholding NNLS (FAST-NNLS) algorithm which significantly outperforms the existing state-of-the-art NNLS algorithms for a wide range of problems. Rigorous convergence guarantees are proven for the proposed method. We demonstrate the effectiveness of our proposed method on multiple synthetic datasets and two realworld text analysis applications. In doing so, we present the most comprehensive NNLS solver comparison in the literature to date.

Cobb, Benjamin [Georgia Institute of Technology]↗

HydroFish: freshwater fish co-occurrence with hydropower plants and non-powered dams in conterminous United States sub-basins

The HydroFish dataset lists all existing hydropower plants (EHAs) and non-powered dams (NPDs; ≥ 0.001 MW potential nominal capacity), delineates the hydrologic sub-basins in which they are situated, and then lists all freshwater fish species reported to occur in those sub-basins. This dataset was compiled using the HydroBio dataset (https://hydrosource.ornl.gov/data/datasets/hydrobio/) and contains 24 total variables that describe hydrologic sub-basins, each unique EHA (plant ID and name, geographic coordinates, permit type, capacity, etc.) and NPD (ID value, known names, geographic coordinates, and estimated potential nominal capacity), and freshwater fish species in the sub-basin (common and scientific name, origin, and migratory and threat status). The HydroBio dataset was built using Oak Ridge National Laboratory’s Existing Hydropower Assets Dataset (2024 version) and Non-powered Dam Technical Potential Dataset (2024 version), and NatureServe’s fish species distribution dataset (2023 version). The dataset also contains summary variables that report the unique number of EHAs, NPDs, and freshwater fish species per sub-basin. The HydroFish dataset contains two unique data files: 1) a .csv metadata file describing the dataset variables, and 2) a .csv data file containing the actual dataset. Note that there may be many rows per unique existing hydropower plant or non-powered dam given that distinct species are listed per existing plant or NPD per sub-basin. The dataset is downloadable as a zip file containing the metadata and dataset files.

Bozeman, Bryan [Oak Ridge National Laboratory (ORN↗

Using feature importance as an exploratory data analysis tool on Earth system models

Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.

Ries, Daniel (ORCID:0000000250294647)↗

Global teleconnections influencing large-scale drought in the United States using SVDI

Understanding recent large-scale drought patterns and the mechanisms producing extreme drought events is vital for future drought forecasts and understanding future drought risks. Increasingly, vapor pressure deficit (VPD) has been used as an important measure of evaporative demand and proxy for drought detection. In this study, VPD is used to calculate the new Standardized VPD Drought Index (SVDI) with NASA North American Land Data Assimilation System (NLDAS) data. Previous studies have shown that SVDI accurately identifies the timing and magnitude short-term droughts in the United States (U.S). In the present study, SVDI is now used to identify large-scale drought patterns between 1980 and 2021 and drought variability driven by selected global teleconnections originating in the Pacific and Atlantic Oceans. Spatial drought characteristics were extracted from SVDI using empirical orthogonal function (EOF) analysis. Then a k-means clustering algorithm was applied to both EOF principal components and primary teleconnections, including the El Nino-Southern Oscillation (ENSO) and Pacific Decadal Oscillation (PDO) to identify drought events driven by the Pacific Ocean. Results show that the SVDI is useful in evaluating large-scale drought variability in the U.S. related to global teleconnections, and that mechanisms influencing summer drought patterns in the Western and Southwestern U.S. are driven by a tropical-extratropical interactions originating in the equatorial Pacific Ocean related to ENSO dynamics with interdecadal variability modulated by PDO. The large-scale droughts in the Central and Southern U.S., like those in 2011 and 2012, on the other hand, are driven by the North Pacific Ocean warm pool during a strong negative PDO, which subsequently influenced variability in the Bermuda-Azores High in the Atlantic Ocean. In summer 2011, the Bermuda-Azores High weakened, reducing the onshore winds and moisture transport along the eastern Gulf of Mexico and contributing to ongoing drought in the region. The Northern Pacific and Atlantic Ocean sea surface temperatures (SSTs) have increased between 1980 and 2021. In conclusion, as SSTs continue to rise in the Northern Pacific Ocean, one consequence of the coupled North Pacific warm pool and atmospheric dynamics, is to increase summer drought variability over a large region in the southern and midwestern U.S. under global warming.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of Autoconversion Representation in E3SMv2 Using an Ensemble of Large-Eddy Simulations of Low-Level Warm Clouds

In numerical atmospheric models that treat cloud and rain droplet populations as separate condensate categories, precipitation initiation in warm clouds is often represented by an autoconversion rate (Au), which is the rate of formation of new rain droplets through the collisions of cloud droplets. Being a function of the cloud droplet size distribution (DSD), the local Au is commonly parameterized as a function of DSD moments: cloud droplet number (n c ) and mass (q c ) concentrations. When applied in a large-scale model, the grid-mean Au must also include a correction, or enhancement factor, to account for the horizontal variability of the cloud properties across the model grid. In this study, we evaluate the Au representation in the Energy Exascale Earth System Model version 2 (E3SMv2) climate model using large-eddy simulations (LES), which explicitly resolve cloud droplet spectra, and therefore the local Au, as well as its spatial variability. The analysis of an ensemble of warm low-level cloud cases shows that the E3SMv2 formulation represents the Au reasonably well compared to the horizontally averaged explicitly computed rate from LES. The agreement, however, comes from a combination of an underestimated E3SM-tuned local Au rate and an overestimated subgrid cloud variability enhancement factor. The latter bias is traced to neglecting the horizontal variability of n c and its co-variability with q c in parameterizing the grid-mean Au.

54 ENVIRONMENTAL SCIENCES↗

Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP)

Anthropogenic climate change is unfolding rapidly, yet its regional manifestation can be obscured by internal variability. A primary goal of climate science is to identify the externally forced climate response from among the noise of internal variability. Separating the forced response from internal variability can be addressed in climate models by using a large ensemble to average over different possible realizations of internal variability. However, with only one realization of the real world, it is a major challenge to isolate the forced response directly in observations. In the Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP), contributors used existing and newly developed statistical and machine learning methods to estimate the forced response over 1950–2022 within individual realizations of the climate system. Participants used neural networks, linear inverse models, fingerprinting methods, and low-frequency component analysis, among other approaches. These methods were trained using large ensembles from multiple climate models and then applied to observations. Here, we evaluate method performance within large ensembles and investigate the estimates of the forced response in observations. Our results show that many different types of methods are skillful for estimating the forced response in climate models, though the relative skill of individual methods varies depending on the variable and evaluation metric. Methods with comparable skill in models can give a wide range of estimates of the forced response pattern in observations, illustrating the epistemic uncertainty in forced response estimates. ForceSMIP gives new insights into the forced response in observations, its uncertainty, and methods for its estimation.

Climate attribution↗

Evapotranspiration Partitioning Using Flux Tower Data in a Semi-Arid Ecosystem

Information about evapotranspiration (ET) and its components, that is, evaporation and transpiration, is crucial for a wide range of water and ecosystem management applications. However, partitioning ET into its two components is often challenging because of their spatiotemporal variabilities and lack of process understanding. This study developed a machine learning (ML) framework to shed light on ET processes and assess the relative importance of different drivers by incorporating hydrometeorology and biomass productivity variables. The Shapley Additive Explanations (SHAP) approach was applied to enhance explainability and rank the importance of ET drivers and their components. A total of 62 variables covering hydrometeorological and biomass productivity dimensions were considered from the Reynolds Creek Critical Zone Observatory (CZO) station in Idaho. The variable importance assessment identified the leading drivers individually for evaporation, transpiration and ET (soil water content for evaporation, vapour pressure deficit for transpiration and soil water content for ET). The results further highlighted the value of combining hydrometeorological and biomass productivity variables to achieve better predictability of ET processes.

54 ENVIRONMENTAL SCIENCES↗

Characterizing climate pathways using feature importance on echo state networks

The 2022 National Defense Strategy of the United States listed climate change as a serious threat to national security. Climate intervention methods, such as stratospheric aerosol injection, have been proposed as mitigation strategies, but the downstream effects of such actions on a complex climate system are not well understood. The development of algorithmic techniques for quantifying relationships between source and impact variables related to a climate event (i.e., a climate pathway) would help inform policy decisions. Data-driven deep learning models have become powerful tools for modeling highly nonlinear relationships and may provide a route to characterize climate variable relationships. In this paper, we explore the use of an echo state network (ESN) for characterizing climate pathways. ESNs are a computationally efficient neural network variation designed for temporal data, and recent work proposes ESNs as a useful tool for forecasting spatiotemporal climate data. However, ESNs are noninterpretable black-box models along with other neural networks. The lack of model transparency poses a hurdle for understanding variable relationships. We address this issue by developing feature importance methods for ESNs in the context of spatiotemporal data to quantify variable relationships captured by the model. We conduct a simulation study to assess and compare the feature importance techniques, and we demonstrate the approach on reanalysis climate data. In the climate application, we consider a time period that includes the 1991 volcanic eruption of Mount Pinatubo. This event was a significant stratospheric aerosol injection, which acts as a proxy for an anthropogenic stratospheric aerosol injection. Furthermore, we are able to use the proposed approach to characterize relationships between pathway variables associated with this event that agree with relationships previously identified by climate scientists.

black-box models↗

Adaptive immersed isogeometric level-set topology optimization

Here, this paper presents for the first time an adaptive immersed approach for level-set topology optimization using higher-order truncated hierarchical B-spline discretizations for design and state variable fields. Boundaries and interfaces are represented implicitly by the iso-contour of one or multiple level-set functions. An immersed finite element method, the eXtended IsoGeometric Analysis, is used to predict the physical response. The proposed optimization framework affords different adaptively refined higher-order B-spline discretizations for individual design and state variable fields. The increased continuity of higher-order B-spline discretizations together with local refinement enables direct control over the accuracy of the representation of each field while simultaneously reducing computational cost compared to uniformly refined discretizations. A flexible mesh adaptation strategy enables local refinement based on geometric measures or physics-based error indicators. These adaptive discretization and analysis approaches are integrated into gradient-based optimization schemes, evaluating the design sensitivities using the adjoint method. Numerical studies illustrate the features of the proposed framework with static, linear elastic, multi-material, two- and three-dimensional problems. The examples provide insight into the effect of refining the design variable field on the optimization result and the convergence rate of the optimization process. Using coarse higher-order B-spline discretizations for level-set fields promotes the development of smooth designs and suppresses the emergence of small features. Moreover, adaptive mesh refinement for state variable fields results in a reduction of overall computational cost. Higher-order B-spline discretizations are especially interesting when evaluating gradients of state variable fields due to their higher inter-element continuity.

36 MATERIALS SCIENCE↗

On Gibbs Equilibrium and Hillert Nonequilibrium Thermodynamics

During his time at Royal Institute of Technology (Kungliga Tekniska högskolan) in Sweden, the present author learned nonequilibrium thermodynamics from Mats Hillert. The key concepts are the separation of internal and external variables of a system and the definitions of potentials and molar quantities. In equilibrium thermodynamics derived by Gibbs, the internal variables are not independent and can be fully evaluated from given external variables. While irreversible thermodynamics led by Onsager focuses on internal variables though often mixed with external variables. Hillert integrated them together by first emphasizing their differences and then examining their connections. His philosophy was reflected by the title of his book “Phase Equilibria, Phase Diagrams and Phase Transformations” that puts equilibrium, nonequilibrium, and internal processes on equal footing. Here, in the present paper honoring Hillert, the present author reflects his experiences with Hillert and his work in last 40 years and expresses his gratitude for all the wisdom and support from him in terms of “Hillert nonequilibrium thermodynamics” and discusses some recent topics that the present author has been working on.

36 MATERIALS SCIENCE↗

Explainable machine learning to quantify the value of proximal remote sensing in latent energy flux estimation

Proximal remote sensing has the potential to provide critical information on vegetation biophysical factors that can predict land-atmosphere exchange of water and energy. Latent energy (LE) flux is traditionally estimated using process-based models which rely on vegetation parameters that change during the growing season. Data-driven models have the potential to address these issues by offering flexible predictor selection and more efficient utilization of the information in predictor sets. These models require careful choice of predictors to avoid redundancy and allow robust cross-validation. In this study we present a systematic and comprehensive evaluation of machine learning (ML) models to assess the capability of meteorological and proximal sensing data for predicting LE at a half-hourly temporal resolution across multiple growing seasons for an agricultural system. The results presented here demonstrate that a model using four environmental predictors in combination with two proximal sensing variables can capture 88 % of the variability in LE. ML models using only three predictors (one meteorological and two proximal remote sensing) captured 81 % of LE variability, offering the best trade-off between performance and complexity. An ML model utilizing only two predictors, one proximal remote sensing variable and downwelling radiation, captured 77 % of LE variability. These results demonstrate the power of proximal remote sensing and meteorological observations to estimate land-atmosphere water vapor exchange, providing a solution where more direct methods such as eddy covariance are not available and for evaluations of agronomic management and genotypic variations.

60 APPLIED LIFE SCIENCES↗

Neural chaos: A spectral stochastic neural operator

Building surrogate models for operators with uncertainty quantification capabilities is essential for many engineering applications where randomness–such as variability in material properties, boundary conditions, and initial conditions–is unavoidable. Polynomial Chaos Expansion (PCE) is widely recognized as a go-to method for constructing stochastic surrogates in both intrusive and non-intrusive ways, and it has recently been used in the context of operator learning. However, its application becomes challenging for complex or high-dimensional processes, as achieving accuracy requires higher-order polynomials, which can increase computational demand and/or the risk of overfitting. Furthermore, PCE requires specialized treatments to manage random variables that are not independent, and these treatments may be problem-dependent or may fail with increasing complexity. Here, in this work, we adopt the same formalism as the spectral expansion used in PCE; however, we replace the classical polynomial basis functions with neural network (NN) basis functions to leverage their expressivity. To achieve this, we propose an algorithm that identifies NN-parameterized basis functions in a purely data-driven manner, without any prior assumptions about the joint distribution of the random variables involved, whether independent or dependent, or about their marginal distributions. The proposed algorithm identifies each NN-parameterized basis function sequentially, ensuring they are orthogonal with respect to the data distribution. The basis functions are constructed directly on the joint stochastic variables without requiring a tensor product structure or assuming independence of the random variables. This approach may offer greater flexibility for complex stochastic models, while simplifying implementation compared to the tensor product structures typically used in PCE to handle random vectors. This is particularly advantageous given the current state of open-source packages, where building and training neural networks can be done with just a few lines of code and extensive community support. We demonstrate the effectiveness of the proposed scheme through several numerical examples of varying complexity and provide comparisons with classical PCE.

Polynomial chaos expansion↗

Systematic Evaluation of Atmospheric Forcing, Surface Datasets, and Mesh Effects on Kilometer-Scale Land Surface and River Modeling

Earth system models are advancing toward kilometer-scale resolution to capture local climate impacts and extremes. High-resolution land and river modeling depends on multiple factors, including mesh, surface datasets, and atmospheric forcing, but their relative effects at kilometer scales remain unquantified. We evaluated five Energy Exascale Earth System Model land and river configurations over the Mid-Atlantic region using two mesh (1/8° structured versus variable-resolution unstructured mesh), two surface datasets (default versus newly developed), and three atmospheric forcings (NLDAS2, MSWX, GSWP). Evaluation against satellite, reanalysis, and in situ benchmarks across water, energy, and carbon cycles quantifies how these factors affect model performance. Forcing selection produces the largest bias reductions (12-99% across variables), followed by surface datasets (7-75%) and mesh (up to 21%). Forcing effects vary by variable, with MSWX reducing biases for snow water equivalent, evapotranspiration, albedo, temperature, and gross primary productivity, GSWP for snow cover and runoff, and NLDAS for soil moisture and streamflow. The use of newly developed surface datasets improves gross primary productivity (58% bias reduction) and evapotranspiration but increase soil moisture and albedo biases due to current modeling limitations. Variable-resolution unstructured mesh improves the simulation of small-basin streamflow through better capturing drainage networks, though mesh minimally affects other land variables. These findings provide important guidance for high-resolution modeling development and actionable science.

Land and River modeling↗

Analysis of biokinetic parameters reveals patterns in mercury accumulation across aquatic species

Mercury (Hg) is a potent neurotoxicant and poses a risk to human health through the ingestion of Hg-contaminated fish. Mercury, especially in its organic form methylmercury (MeHg), biomagnifies up food chains such that even small aqueous concentrations of Hg can result in significant concentrations of total Hg in fish. Understanding the ecological and human health risks associated with Hg and MeHg exposure requires an understanding of the factors that affect its bioaccumulation in aquatic species. We compiled estimates of three biokinetic parameters: uptake rate (k u ), assimilation efficiency (AE), and efflux rate (k e ). These parameters describe contaminant uptake from aqueous (k u ) and dietary (AE) exposure and the rate of excretion (k e ). We found parameter values for 38 and 34 different species of fish and aquatic invertebrates, respectively, and collected 502 parameter values in total. Here, we used a machine learning technique to establish the relationships between experimental and physiological variables and these parameter values. We found differences in which variables were associated with biokinetic parameter values for fish and aquatic invertebrates. The form of Hg was the most impactful variable, influencing values of all parameters except k u for invertebrates, for which aqueous exposure time was the only significant predicator variable. The parameter k e were the only values significantly influenced by more than one variable, with water type (freshwater, brackish, or marine), organism weight, and form of Hg significantly impacting parameter values for fish and/or invertebrates. To our knowledge, this study represents the most extensive review of biokinetic parameters of Hg and MeHg accumulation in aquatic organisms. Environmental parameters found to significantly impact Hg and MeHg bioaccumulation in past studies were not identified as important in our analyses across aquatic ecosystems and species. Our dataset and analysis reveal novel patterns that may help us better understand and manage Hg bioaccumulation.

54 ENVIRONMENTAL SCIENCES↗

Performance evaluation of CMIP6 models on the Arctic-Siberian Plain teleconnection affecting the East Asian heat waves

The frequency and intensity of summer heat waves in East Asia have increased sharply in recent decades, significantly impacting public health and the economy. The Arctic-Siberian Plain (ASP) teleconnection pattern has been identified as a key driver, with ASP warming amplifying atmospheric circulation patterns conducive to extreme temperatures. This study evaluates the ability of Coupled Model Inter-comparison Project phase 6 models to simulate the ASP pattern across interannual variability (IAV) and intra-seasonal variability (ISV) timescales using the Common Basis Function method. The multi-model mean shows statistically significant pattern correlations with ERA5 reanalysis, with correlation coefficients of 0.90 and 0.99 for IAV and ISV, respectively. While the ASP pattern is generally well captured, models exhibit substantial inter-model diversity in the intensity and position of anticyclonic anomalies over the ASP and East Asia. Models with ASP pattern variability similar to reanalysis better reproduce extreme East Asian temperatures, whereas those over- or underestimating ASP variability exhibit lower skill. These performance differences are related to differences in simulating key variables associated with the development of the ASP pattern. Our findings highlight the role of the ASP pattern in modulating extreme heat events, as models with improved ASP simulations align more closely with observed temperature extremes. Refining ASP representations in models could enhance seasonal heat wave predictions, improving climate adaptation strategies.

Arctic-Siberian Plain (ASP)↗

Informative and non-informative decomposition of turbulent flow fields

Not all the information in a turbulent field is relevant for understanding particular regions or variables in the flow. Here, we present a method for decomposing a source field into its informative Φ I (x, t) and residual Φ R (x, t) components relative to another target field. The method is referred to as informative and non-informative decomposition (IND). All the necessary information for physical understanding, reduced-order modelling and control of the target variable is contained in Φ I (x, t), whereas Φ R (x, t) offers no substantial utility in these contexts. The decomposition is formulated as an optimisation problem that seeks to maximise the time-lagged mutual information of the informative component with the target variable while minimising the mutual information with the residual component. The method is applied to extract the informative and residual components of the velocity field in a turbulent channel flow, using the wall shear stress as the target variable. We demonstrate the utility of IND in three scenarios: (i) physical insight into the effect of the velocity fluctuations on the wall shear stress; (ii) prediction of the wall shear stress using velocities far from the wall; and (iii) development of control strategies for drag reduction in a turbulent channel flow using opposition control. In case (i), IND reveals that the informative velocity related to wall shear stress consists of wall-attached high- and low-velocity streaks, collocated with regions of vertical motions and weak spanwise velocity. This informative structure is embedded within a larger-scale streak–roll structure of residual velocity, which bears no information about the wall shear stress. In case (ii), the best-performing model for predicting wall shear stress is a convolutional neural network that uses the informative component of the velocity as input, while the residual velocity component provides no predictive capabilities. Finally, in case (iii), we demonstrate that the informative component of the wall-normal velocity is closely linked to the observability of the target variable and holds the essential information needed to develop successful control strategies.

97 MATHEMATICS AND COMPUTING↗

Modeling neutral defects in III-V ternary alloys with a special quasirandom structure: Analysis of As- and III-site point defects in InGaAs

While first-principles density functional theory modeling has become a vital tool to investigate defect properties in semiconductors, the lack of crystalline periodicity in pseudobinary random composition alloys, such as In 1−𝑥 ⁢Ga 𝑥 ⁢As, complicates such analyses. We present a simulation strategy to systematically take into account the variability in the local defect environment in order to predict statistical properties of neutral intrinsic defects in In 1−𝑥⁢ Ga 𝑥 ⁢As. We use a comprehensive sampling from a modest-sized 64-atom special quasirandom structure (SQS) to define a statistically representative set of defects, and use a 512-atom hypercell, a 2 × 2 × 2 supercell of SQS supercells, to achieve cell-size convergence. We articulate an equivalent site principle and describe how it constrains atomic chemical reference energies in computation of defect formation energies in pseudobinary alloys. A simple protocol for estimating reference energies for the Ga and In atoms sharing the III site succeeds in obtaining the equivalence of defects at Ga-sites and In sites in the SQS supercell, (<30 meV differences in average formation energies). For III-site defects, such as the As antisite As III , the statistical variability in formation energies is modest, ≈ 0.1–0.2 eV. The variability in formation energy at As-site defects, such as the As vacancy 𝑣 As , can be much larger, >1 eV. The As antisite is shown to be a low-energy defect and the most likely to be present in as-grown materials, just as in GaAs. All other defects are higher-energy defects unlikely to be important in native material, but potentially important in radiation-damaged material. With a strong variability in defect energies, especially on the As-site, explicit consideration of statistical variability due to compositional randomness will be imperative for meaningful and quantitative comparisons to experiment.

Density functional theory↗