Search NASA⌕ Search

SEARCH · Search NASA

Results for “operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Learning functional priors and posteriors from data and physics

In this work, we develop a new Bayesian framework based on deep neural networks to be able to extrapolate in space-time using historical data and to quantify uncertainties arising from both noisy and gappy data in physical problems. Specifically, the proposed approach has two stages: (1) prior learning and (2) posterior estimation. At the first stage, we employ the physics-informed Generative Adversarial Networks (PI-GAN) to learn a functional prior either from a prescribed function distribution, e.g., Gaussian process, or from historical data and physics. At the second stage, we employ the Hamiltonian Monte Carlo (HMC) method to estimate the posterior in the latent space of PI-GANs. In addition, we use two different approaches to encode the physics: (1) automatic differentiation, used in the physicsinformed neural networks (PINNs) for scenarios with explicitly known partial differential equations (PDEs), and (2) operator regression using the deep operator network (DeepONet) for PDE-agnostic scenarios. We then test the proposed method for (1) meta-learning for one-dimensional regression, and forward/inverse PDE problems (combined with PINNs); (2) PDE-agnostic physical problems (combined with DeepONet), e.g., fractional diffusion as well as saturated stochastic (100-dimensional) flows in heterogeneous porous media; and (3) spatial-temporal regression problems, i.e., inference of a marine riser displacement field using experimental data from the Norwegian Deepwater Programme (NDP). The results demonstrate that the proposed approach can provide accurate predictions as well as uncertainty quantification given very limited scattered and noisy data, since historical data could be available to provide informative priors. In summary, the proposed method is capable of learning flexible functional priors, e.g., both Gaussian and non-Gaussian process, and can be readily extended to big data problems by enabling mini-batch training using stochastic HMC or normalizing flows since the latent space is generally characterized as low dimensional.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

MIONet: Learning Multiple-Input Operators via Tensor Product

As an emerging paradigm in scientific machine learning, neural operators aim to learn operators, via neural networks, that map between infinite-dimensional function spaces. Several neural operators have been recently developed. However, all the existing neural operators are only designed to learn operators defined on a single Banach space; i.e., the input of the operator is a single function. Here, for the first time, we study the operator regression via neural networks for multiple-input operators defined on the product of Banach spaces. We first prove a universal approximation theorem of continuous multiple-input operators. We also provide a detailed theoretical analysis including the approximation error, which provides guidance for the design of the network architecture. Based on our theory and a low-rank approximation, we propose a novel neural operator, MIONet, to learn multiple-input operators. MIONet consists of several branch nets for encoding the input functions and a trunk net for encoding the domain of the output function. Here, we demonstrate that MIONet can learn solution operators involving systems governed by ordinary and partial differential equations. In our computational examples, we also show that we can endow MIONet with prior knowledge of the underlying system, such as linearity and periodicity, to further improve accuracy.

97 MATHEMATICS AND COMPUTING↗

Airborne hyperspectral imaging of cover crops through radiative transfer process-guided machine learning

Cover cropping between cash crop growing seasons is a multifunctional conservation practice. Timely and accurate monitoring of cover crop traits, notably aboveground biomass and nutrient content, is beneficial to agricultural stakeholders to improve management and understand outcomes. Currently, there is a scarcity of spatially and temporally resolved information for assessing cover crop growth. Remote sensing has a high potential to fill this need, but conventional empirical regression operated with coarse-resolution multispectral data has large uncertainties. Therefore, this study utilized airborne hyperspectral imaging techniques and developed new process-guided machine learning approaches (PGML) for cover crop monitoring. Specifically, we deployed an airborne hyperspectral system covering visible to shortwave-infrared wavelengths (400–2400 nm) to acquire high spatial (0.5 m) and spectral (3–5 nm) resolution reflectance over 23 cover crop fields across Central Illinois in March and April of 2021. Airborne hyperspectral surface reflectance with high spectral and spatial resolution can be well matched with field data to quantify cover crop traits. Furthermore, the PGML models were pre-trained by synthetic data from soil-vegetation radiative transfer modeling (one million records), and then fine-tuned with field data of cover crop biomass and nutrient content. Results show that airborne hyperspectral data with PGML can achieve high accuracy to predict cover crop aboveground biomass (R 2 = 0.72, relative RMSE = 15.16%) and nitrogen content (R 2 = 0.69, relative RMSE = 16.59%) through leave-one-field-out cross-validation. Unlike the pure data-driven approach (e.g., partial least-squares regression), PGML incorporated radiative transfer knowledge and obtained higher predictive performance with fewer field data. Meanwhile, with field data for model fine-tuning, PGML predicted biomass more accurately than the inversion of radiative transfer models. Here we also found that the red edge has a high contribution in quantifying aboveground biomass and nitrogen content, followed by green and shortwave spectra. This study demonstrated the first attempt of utilizing hyperspectral remote sensing to accurately quantify cover crop traits. We highlight the strength of PGML in exploiting sensing data to quantify ecosystem variables to advance agroecosystem monitoring for sustainable agricultural management.

60 APPLIED LIFE SCIENCES↗

How Does Flow Alteration Propagate Across a Large, Highly Regulated Basin? Dam Attributes, Network Context, and Implications for Biodiversity

Abstract Large dams are a leading cause of river ecosystem degradation. Although dams have cumulative effects as water flows downstream in a river network, most flow alteration research has focused on local impacts of single dams. Here we examined the highly regulated Colorado River Basin (CRB) to understand how flow alteration propagates in river networks, as influenced by the location and characteristics of dams as well as the structure of the river network—including the presence of tributaries. We used a spatial Markov network model informed by 117 upstream‐downstream pairs of monthly flow series (2003–2017) to estimate flow alteration from 84 intermediate‐to‐large dams representing >83% of the total storage in the CRB. Using Least Absolute Shrinkage and Selection Operator regression, we then investigated how flow alteration was influenced by local dam properties (e.g., purpose, storage capacity) and network‐level attributes (e.g., position, upstream cumulative storage). Flow alteration was highly variable across the network, but tended to accumulate downstream and remained high in the main stem. Dam impacts were explained by network‐level attributes (63%) more than by local dam properties (37%), underscoring the need to consider network context when assessing dam impacts. High‐impact dams were often located in sub‐watersheds with high levels of native fish biodiversity, fish imperilment, or species requiring seasonal flows that are no longer present. These three biodiversity dimensions, as well as the amount of dam‐free downstream habitat, indicate potential to restore river ecosystems via controlled flow releases. Our methods are transferrable and could guide screening for dam reoperation in other highly regulated basins.

54 ENVIRONMENTAL SCIENCES↗

Mechanistic studies of small molecule ligands selective to RNA single G bulges

Abstract Small-molecule RNA binders have emerged as an important pharmacological modality. A profound understanding of the ligand selectivity, binding mode, and influential factors governing ligand engagement with RNA targets is the foundation for rational ligand design. Here, we report a novel class of coumarin derivatives exhibiting selective binding affinity towards single G RNA bulges. Harnessing the computational power of all-atom Gaussian accelerated molecular dynamics simulations, we unveiled a rare minor groove binding mode of the ligand with a key interaction between the coumarin moiety and the G bulge. This predicted binding mode is consistent with results obtained from structure-activity relationship studies and transverse relaxation measurements by nuclear magnetic resonance spectroscopy. We further generated 444 molecular descriptors from 69 coumarin derivatives and identified key contributors to the binding events, such as charge state and planarity, by lasso (least absolute shrinkage and selection operator) regression. Our work deepened the understanding of RNA-small molecule interactions and integrated a new framework for the rational design of selective small-molecule RNA binders.

Biochemistry & Molecular Biology↗

Optimization and Multimachine Learning Algorithms to Predict Nanometal Surface Area Transfer Parameters for Gold and Silver Nanoparticles

Interactions between gold metallic nanoparticles and molecular dyes have been well described by the nanometal surface energy transfer (NSET) mechanism. However, the expansion and testing of this model for nanoparticles of different metal composition is needed to develop a greater variety of nanosensors for medical and commercial applications. In this study, the NSET formula was slightly modified in the size-dependent dampening constant and skin depth terms to allow for modeling of different metals as well as testing the quenching effects created by variously sized gold, silver, copper, and platinum nanoparticles. Overall, the metal nanoparticles followed more closely the NSET prediction than for Förster resonance energy transfer, though scattering effects began to occur at 20 nm in the nanoparticle diameter. To further improve the NSET theoretical equation, an attempt was made to set a best-fit line of the NSET theoretical equation curve onto the Au and Ag data points. An exhaustive grid search optimizer was applied in the ranges for two variables, 0.1≤C≤2.0 and 0≤α≤4, representing the metal dampening constant and the orientation of donor to the metal surface, respectively. Three different grid searches, starting from coarse (entire range) to finer (narrower range), resulted in more than one million total calculations with values C=2.0 and α=0.0736. The results improved the calculation, but further analysis needed to be conducted in order to find any additional missing physics. With that motivation, two artificial intelligence/machine learning (AI/ML) algorithms, multilayer perception and least absolute shrinkage and selection operator regression, gave a correlation coefficient, R2, greater than 0.97, indicating that the small dataset was not overfitting and was method-independent. This analysis indicates that an investigation is warranted to focus on deeper physics informed machine learning for the NSET equations.

Demers, Steven M. E. (ORCID:0000000192213246)↗

Empirical radius formulas for canonical neutron stars from bidirectionally selecting features of equations of state in extended Bayesian analyses of observational data

Significant advancement in Bayesian inference of nuclear equation of state (EOS) from gravitational wave and x-ray observations of neutron stars (NSs) has been made by the nuclear astrophysics community especially since GW170817. By extending the traditional Bayesian analysis which normally ends at presenting the marginalized posterior probability distribution functions (PDFs) of individual EOS parameters and their correlations (or sometimes only the Pearson correlation coefficients which are only reliably useful when the variables are linearly correlated while they are actually often not), we search for a data-driven and robust empirical formula for the radius 𝑅 1.4 of canonical NSs in terms of the characteristic EOS parameters (features). We also identify the single most important but currently poorly known EOS parameter for determining the 𝑅 1.4 . Using three regression-model-building methodologies: bidirectional stepwise feature selection, least absolute shrinkage selection operator (LASSO) regression, and neural network regression on a large set of posterior EOSs and the corresponding 𝑅 1.4 values inferred from earlier comprehensive Bayesian analyses of NS observational data, we systematically and rigorously develop the most probable 𝑅 1.4 formulas with varying statistical accuracy and technical complexity. Here, the most important EOS parameters for determining 𝑅 1.4 are found consistently in each of the feature selection processes to be (in order of decreasing importance): curvature 𝐾 sym , slope 𝐿, skewness 𝐽 sym of nuclear symmetry energy, skewness 𝐽 0 , incompressibility 𝐾 0 of symmetric nuclear matter, and the magnitude 𝐸 sym ⁡(𝜌 0 ) of symmetry energy at the saturation density 𝜌 0 of nuclear matter.

Bayesian methods↗

A comprehensive and fair comparison of two neural operators (with practical extensions) based on $\mathrm{FAIR}$ data

Neural operators can learn nonlinear mappings between function spaces and offer a new simulation paradigm for real-time prediction of complex dynamics for realistic diverse applications as well as for system identification in science and engineering. Herein, we investigate the performance of two neural operators, which have shown promising results so far, and we develop new practical extensions that will make them more accurate and robust and importantly more suitable for industrial-complexity applications. The first neural operator, DeepONet, was published in 2019 (Lu et al., 2019), and its original architecture was based on the universal approximation theorem of Chen & Chen (1995). The second one, named Fourier Neural Operator or FNO, was published in 2020, and it is based on parameterizing the integral kernel in the Fourier space. DeepONet is represented by a summation of products of neural networks (NNs), corresponding to the branch NN for the input function and the trunk NN for the output function; both NNs are general architectures, e.g., the branch NN can be replaced with a CNN or a ResNet. According to Kovachki et al. (2021), FNO in its continuous form can be viewed conceptually as a DeepONet with a specific architecture of the branch NN and a trunk NN represented by a trigonometric basis. In order to compare FNO with DeepONet computationally for realistic setups, we develop several extensions of FNO that can deal with complex geometric domains as well as mappings where the input and output function spaces are of different dimensions. We also develop an extended DeepONet with special features that provide inductive bias and accelerate training, and we present a faster implementation of DeepONet with cost comparable to the computational cost of FNO, which is based on the Fast Fourier Transform. Here we consider 16 different benchmarks to demonstrate the relative performance of the two neural operators, including instability wave analysis in hypersonic boundary layers, prediction of the vorticity field of a flapping airfoil, porous media simulations in complex-geometry domains, etc. We follow the guiding principles of FAIR (Findability, Accessibility, Interoperability, and Reusability) for scientific data management and stewardship. The performance of DeepONet and FNO is comparable for relatively simple settings, but for complex geometries the performance of FNO deteriorates greatly. We also compare theoretically the two neural operators and obtain similar error estimates for DeepONet and FNO under the same regularity assumptions.

42 ENGINEERING↗

Leveraging hyperspectral phenotyping for accurate, non-destructive prediction of metabolite profiles in poplar under drought stress

Accurately predicting drought tolerance in woody perennial bioenergy crops is critical for sustainable biomass production under fluctuating precipitation. Hyperspectral imaging (HSI) in the visible-near-infrared (VNIR) and shortwave-infrared (SWIR) ranges offers a promising approach for predicting plant biochemical traits, yet its application in metabolite profiling remains underexplored. We integrated VNIR+SWIR HSI with untargeted metabolomics to investigate drought-induced metabolic shifts in Populus leaves from eight Populus genotypes. Metabolite profiling identified 127 compounds, with 73 showing significant drought responses spanning amino acids (AA), carbohydrates (CHO), phenolic glycosides (PG), organic acids (OA), fatty acids and alcohols (FA), terpenes (T), phenolic metabolites (P), and unclassified metabolites. Spectral analysis revealed consistently higher reflectance across VNIR and SWIR wavelengths in drought-stressed plants, corresponding with increased accumulation of AA and reduced CHO and PG levels. Least absolute shrinkage and selection operator (LASSO) regression modeling identified robust spectral predictors of metabolite concentrations, associating VNIR wavelengths (500–700 nm) predominantly with AA and P, whereas SWIR wavelengths (1680–1700 nm) reliably predicted CHO, OA, and T. Several stable spectral-metabolite associations persisted across the two watering regimes (drought vs. well-watered), highlighting their potential as spectral biomarkers for non-destructive stress monitoring. Minimal genotype-specific variation suggests that observed spectral and metabolic responses were driven primarily by environmental factors, likely reflecting limited genetic diversity among the commercial Populus genotypes examined. This work establishes VNIR+SWIR hyperspectral imaging as a powerful, non-destructive phenotyping tool for precision monitoring and targeted improvement of drought resilience in bioenergy crops.

Biochemical trait prediction↗

Deep Neural Network Algorithm for CMC Microstructure Characterization and Variability Quantification

Microstructure characterization and variability quantification are crucial for understanding ceramic matrix composites (CMCs) mechanical behavior and deformation mechanisms across length scales. Traditionally, analyses of the micrographs obtained from microscopy are labor-intensive. However, with the vast improvement in computer vision (CV) and deep learning (DL), an automated algorithm can be designed to extract essential microstructure variability from micrographs which can then be used to construct a statistically representative volume element (SRVE). The DL-based algorithm spans the taxonomy of microstructure analyses, including semantic segmentation of microstructure constituents, secondary phases, matrix/fiber interface, and defects, and quantifying the microstructure variability in terms of probability distributions. In this work, C/SiNC and SiC/SiNC CMCs microstructures are semantically segmented through a deep convolutional neural network, followed by variability quantification through the implementation of a fully connected regression layer, hence forming a deep regression network. The deep regression network operates in a feedforward regime, in which the neuron output signal traverses through the network in a unidirectional manner. The weight tensor associated with each layer is updated through a backpropagation stochastic gradient descent approach. The input gray-scale image obtained through in-house scanning electron microscope and confocal microscope micrographs is augmented through affine transformations to increase the training set size, which is then processed through four strided convolutional layers. This compresses the image resolution by half at each layer while increasing the image depth by applying different filters (image encoding). The class activation maps (CAMs) corresponding to the applied filters highlight the key architectural features and assist with the semantic segmentation of the microstructure.

Hamza, Mohamed H.↗

Intermediate Molecular Phenotypes to Identify Genetic Markers of Anthracycline-Induced Cardiotoxicity Risk

Cardiotoxicity due to anthracyclines (CDA) affects cancer patients, but we cannot predict who may suffer from this complication. CDA is a complex trait with a polygenic component that is mainly unidentified. We propose that levels of intermediate molecular phenotypes (IMPs) in the myocardium associated with histopathological damage could explain CDA susceptibility, so variants of genes encoding these IMPs could identify patients susceptible to this complication. Thus, a genetically heterogeneous cohort of mice (n = 165) generated by backcrossing were treated with doxorubicin and docetaxel. We quantified heart fibrosis using an Ariol slide scanner and intramyocardial levels of IMPs using multiplex bead arrays and QPCR. We identified quantitative trait loci linked to IMPs (ipQTLs) and cdaQTLs via linkage analysis. In three cancer patient cohorts, CDA was quantified using echocardiography or Cardiac Magnetic Resonance. CDA behaves as a complex trait in the mouse cohort. IMP levels in the myocardium were associated with CDA. ipQTLs integrated into genetic models with cdaQTLs account for more CDA phenotypic variation than that explained by cda-QTLs alone. Allelic forms of genes encoding IMPs associated with CDA in mice, including AKT1, MAPK14, MAPK8, STAT3, CAS3, and TP53, are genetic determinants of CDA in patients. Two genetic risk scores for pediatric patients (n = 71) and women with breast cancer (n = 420) were generated using machine-learning Least Absolute Shrinkage and Selection Operator (LASSO) regression. Thus, IMPs associated with heart damage identify genetic markers of CDA risk, thereby allowing more personalized patient management.

60 APPLIED LIFE SCIENCES↗

Serum bile acid and unsaturated fatty acid profiles of non-alcoholic fatty liver disease in type 2 diabetic patients

The understanding of bile acid (BA) and unsaturated fatty acid (UFA) profiles, as well as their dysregulation, remains elusive in individuals with type 2 diabetes mellitus (T2DM) coexisting with non-alcoholic fatty liver disease (NAFLD). Investigating these metabolites could offer valuable insights into the pathophy-siology of NAFLD in T2DM. Our aim is to identify potential metabolite biomarkers capable of distinguishing between NAFLD and T2DM. A training model was developed involving 399 participants, comprising 113 healthy controls (HCs), 134 individuals with T2DM without NAFLD, and 152 individuals with T2DM and NAFLD. External validation encompassed 172 participants. NAFLD patients were divided based on liver fibrosis scores. The analytical approach employed univariate testing, orthogonal partial least squares-discriminant analysis, logistic regression, receiver operating characteristic curve analysis, and decision curve analysis to pinpoint and assess the diagnostic value of serum biomarkers. Compared to HCs, both T2DM and NAFLD groups exhibited diminished levels of specific BAs. In UFAs, particular acids exhibited a positive correlation with NAFLD risk in T2DM, while the ω-6:ω-3 UFA ratio demonstrated a negative correlation. Levels of α-linolenic acid and γ-linolenic acid were linked to significant liver fibrosis in NAFLD. The validation cohort substantiated the predictive efficacy of these biomarkers for assessing NAFLD risk in T2DM patients. This study underscores the connection between altered BA and UFA profiles and the presence of NAFLD in individuals with T2DM, proposing their potential as biomarkers in the pathogenesis of NAFLD.

60 APPLIED LIFE SCIENCES↗

Practical Guide to Chemometric Analysis of Optical Spectroscopic Data

The methodology and mathematical treatment of several classic multivariate methods for the analysis of spectroscopic data is demonstrated in a straightforward way that can be used as a basis for teaching an undergraduate introductory course on chemometric analysis. The multivariate techniques of classical least squares (CLS), principal component regression (PCR), and partial least squares (PLS), as well as the univariate Beer’s law method have been described and compared, building students’ understanding by starting with the univariate method and progressing step by step into the multivariate methods. Equations for the production of regression vectors from training set spectral data is described and their use demonstrated for the prediction of constituent concentrations on a separate validation set of spectra. Extreme care is taken to ensure consistency in variable formatting of data matrices. This provides a key foundation to understanding how spectral data are manipulated using these different mathematical approaches for building quantitative regression models. Each method is applied to a real-world data set, and the results are discussed to show students the types of information that can be gleaned from each method. A training set comprised of 20 infrared absorbance spectra containing 3 constituents (benzene, polystyrene, and gasoline) of known composition are used to demonstrate the matrix operations for each regression method. A separate set of 12 real-world napalm samples (containing benzene, polystyrene and gasoline) are used as a validation set to demonstrate the ability to utilize the regression models on an unknown dataset. A toolbox (PNNL Chemometric Toolbox) written in MATLAB language is supplied in the Supplemental Information file and can be used as a companion for understanding the development and deployment of the chemometric algorithms described in this paper. The datasets of the infrared spectra are also supplied, allowing users to build and inspect the chemometric models on their own. Finally, the Toolbox includes scripts to assist users in loading their own datasets into MATLAB and performing CLS, PCR, and PLS on their data.

Upper-Division Undergraduate, Analytical Chemistry↗

NCAPH drives breast cancer progression and identifies a gene signature that predicts luminal a tumour recurrence

Luminal A tumours generally have a favourable prognosis but possess the highest 10-year recurrence risk among breast cancers. Additionally, a quarter of the recurrence cases occur within 5 years post-diagnosis. Identifying such patients is crucial as long-term relapsers could benefit from extended hormone therapy, while early relapsers might require more aggressive treatment. We conducted a study to explore non-structural chromosome maintenance condensin I complex subunit H’s (NCAPH) role in luminal A breast cancer pathogenesis, both in vitro and in vivo, aiming to identify an intratumoural gene expression signature, with a focus on elevated NCAPH levels, as a potential marker for unfavourable progression. Our analysis included transgenic mouse models overexpressing NCAPH and a genetically diverse mouse cohort generated by backcrossing. A least absolute shrinkage and selection operator (LASSO) multivariate regression analysis was performed on transcripts associated with elevated intratumoural NCAPH levels. We found that NCAPH contributes to adverse luminal A breast cancer progression. The intratumoural gene expression signature associated with elevated NCAPH levels emerged as a potential risk identifier. Transgenic mice overexpressing NCAPH developed breast tumours with extended latency, and in Mouse Mammary Tumor Virus (MMTV)-NCAPH ErbB2 double-transgenic mice, luminal tumours showed increased aggressiveness. High intratumoural Ncaph levels correlated with worse breast cancer outcome and subpar chemotherapy response. A 10-gene risk score, termed Gene Signature for Luminal A 10 (GSLA10), was derived from the LASSO analysis, correlating with adverse luminal A breast cancer progression. The GSLA10 signature outperformed the Oncotype DX signature in discerning tumours with unfavourable outcomes, previously categorised as luminal A by Prediction Analysis of Microarray 50 (PAM50) across three independent human cohorts. This new signature holds promise for identifying luminal A tumour patients with adverse prognosis, aiding in the development of personalised treatment strategies to significantly improve patient outcomes.

60 APPLIED LIFE SCIENCES↗