Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Chemistry↗

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Arasteh, Kourosh [Lawrence Livermore National Labo↗

PV Generation and Load Forecasting for Adjuntas PR Community Microgrids

Existing frameworks to forecast time-series photovoltaic (PV) output power and consumer load for microgrid operations and controls assume a near-continuous availability of real-time input features from the field assets such as PV inverters, energy meters, and weather station. These incoming data points are used to periodically retrain models and update forecast snapshots over a moving horizon window, be it one hour-ahead, one-day ahead, or one-week ahead. However, such frameworks are not resilient to disruptions in data availability caused by losses in communications between the field sensors and data loggers. Hence, there is a need for programs that assume no availability of real-time microgrid asset data and still make reliable forecasts that can be used for decision-making. Such programs would be apt to function in extreme weather events such as hurricanes and would use lightweight recursive time-series models to independently forecast solar irradiance and ambient temperature, then compute PV power from those forecasts, as well as independently forecast consumer load. The codebase performs forecasting for the scenario of when the microgrid does not have a reliable access to forecasts or real-time observations of solar irradiance (I) and ambient temperature (AT) and load (Load) to be able to adequately forecast, in real-time, the PV power production or a business' load. In this case, using historical values of PV power and load, a univariate forecasting of generation and consumption are respectively made. The use-case in particular has two sub-scenarios: one, a normal 7-day ahead forecast where the unavailability of real-time data is assumed due to infrastructure issues such as loss of communication or sensor maintenance or service downtimes. Whereas a hurricane-caused unavailability of real-time data requires a second model trained specifically on historical hurricane days to be able to capture the extreme day behavior of generation in particular, and load if applicable. A gradient boosted regression tree comprises an ensemble of additive models that map between the input of historical values (be it irradiance, temperature, or load) and their corresponding output forecasts of a given horizon such that the individual learner predictions are summed up over the total number of such learners in the ensemble to produce an aggregate forecast. A weighting mechanism is applied to the training data in each iteration, where actual and forecast values are compared to penalize incorrect forecasts by increasing the weight and reducing it to reward correct forecasts. The code's benefits are that it: (a) accounts for a contingency where communication loss renders newly measured real-time data unavailable for model tuning and snapshot updates; (b) presents blind forecasting that recursively determines the next time-step value in a horizon using the forecast of the same attribute from a prior step; and (c) employs lightweight models that, once trained, can reliably generalize for different horizons, which make them suitable for enhancing the resilience of field microgrids prone to extreme events that encounter disruptions to data availability.

Sundararajan, Aditya [Oak Ridge National Laborator↗

Catalytic Reduction of Esters over Zirconia-Supported Metal Catalysts

Esters are often produced as unwanted byproducts during the catalytic upgrading of ethanol to diesel fuel precursors through Guerbet coupling. Removal of esters from the product stream is important to prevent the loss of downstream catalyst activity from ester-derived carboxylic acids. In this work, we studied ester hydrogenolysis to the parent alcohols as a viable route for enhanced diesel fuel production. Specifically, we investigated the reduction of hexyl acetate in butanol over ZrO 2 -supported Ni, Co, Cu, Rh, Pd, and Pt catalysts, where Cu/ZrO 2 was the most selective catalyst for the hydrogenolysis of hexyl acetate into hexanol and ethanol. Thermodynamic analysis reveals that a 90% alcohol yield can be obtained at 200 °C, 30 bar, and a relatively high H 2 :hexyl acetate molar ratio of 480:1. Experimentally, an alcohol yield of 88% yield was obtained with a 10 wt % Cu/ZrO 2 catalyst at these conditions with a residence time of 5.4 h kg cat kmol gas –1 . Catalytic tests on the support revealed that ZrO 2 catalyzes the transesterification reaction between hexyl acetate and butanol. However, only the Cu sites can catalyze the hydrogenolysis of the esters into the final alcohols. We developed a kinetic model for our experimental results, which shows that the transesterification and hydrogenolysis reactions run at two different timescales, the former being 10 times faster than the latter. Data regression has been used to develop a model to predict the mole fraction distribution of ester hydrogenolysis products over a wide range of contact times. Cu/ZrO 2 loses half its catalytic activity after 80 h of time on stream. Modeling of deactivation data reveals that the ZrO 2 support conserves a residual activity due to external active sites, while active sites over the Cu surface deactivate at different rates. Furthermore, the catalytic conversion of esters into their parent alcohols is relevant to the production of surrogate liquid fuels since alcohols can be bimolecularly dehydrated to produce a blend of ethers with diesel fuel-like properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HighDimMixedModels.jl: Robust high-dimensional mixed-effects models across omics data

High-dimensional mixed-effects models are an increasingly important form of regression in which the number of covariates rivals or exceeds the number of samples, which are collected in groups or clusters. The penalized likelihood approach to fitting these models relies on a coordinate descent algorithm that lacks guarantees of convergence to a global optimum. Here, we empirically study the behavior of this algorithm on simulated and real examples of three types of data that are common in modern biology: transcriptome, genome-wide association, and microbiome data. Our simulations provide new insights into the algorithm’s behavior in these settings, and, comparing the performance of two popular penalties, we demonstrate that the smoothly clipped absolute deviation (SCAD) penalty consistently outperforms the least absolute shrinkage and selection operator (LASSO) penalty in terms of both variable selection and estimation accuracy across omics data. To empower researchers in biology and other fields to fit models with the SCAD penalty, we implement the algorithm in a Julia package, HighDimMixedModels.jl .

Gorstein, Evan↗

UNCERTAINTY-AWARE DEEP LEARNING FRAMEWORK FOR FORECASTING COASTAL WATER LEVEL IN VIRGINIA BEACH

Coastal areas like Virginia Beach, USA, are increasingly vulnerable to flooding. To mitigate the impact of flooding, it is crucial for the City of Virginia Beach to have reliable 72-hour-ahead (3 days) forecasts of water levels at key gauge locations. To support this effort, several sensors have been installed throughout the city to monitor water levels and other environmental parameters such as wind speed, precipitation, and atmospheric pressure. Leveraging sensor data from one of these locations, we developed an uncertainty-aware deep learning model to forecast water levels. We employed deep quantile regression (DQR) to quantify variability in the predictions and examined the performance of three different model architectures. In addition to exclusively including historical data, we investigated the improvement wind forecasts provide to the accuracy of 72-hour-ahead water level predictions. The results show a twelvefold improvement in the flood forecast for a real flooding event.

Hasan, Mahmud [Thomas Jefferson National Accelerat↗

Bayesian Adaptive Polynomial Chaos Expansions

Polynomial chaos expansions (PCEs) are widely used for uncertainty quantification (UQ) tasks, particularly in the applied mathematics community. However, PCE has received comparatively less attention in the statistics literature, and fully Bayesian formulations remain rare—especially with implementations in R. Motivated by the success of adaptive Bayesian machine learning models such as BART, BASS and BPPR, we develop a new fully Bayesian adaptive PCE method with an efficient and accessible R implementation: khaos. Our approach includes a novel proposal distribution that enables data-driven interaction selection and supports a modified g-prior tailored to PCE structure. Through simulation studies and real-world UQ applications, we demonstrate that the Bayesian adaptive PCE provides competitive performance for surrogate modeling, global sensitivity analysis and ordinal regression tasks.

97 MATHEMATICS AND COMPUTING↗

Modeling ethanol/water adsorption in all-silica zeolites using the real adsorbed solution theory

A comprehensive set of single-component and binary isotherms were collected for ethanol/water adsorption into the siliceous forms of 185 known zeolites using grand-canonical Monte Carlo simulations. Using these data, a systematic analysis of ideal/real adsorbed-solution theory (IAST/RAST) was conducted and activity coefficients were derived for ethanol/water mixtures adsorbed in different zeolites based on RAST. It was found that activity coefficients of ethanol are close to unity while activity coefficients of water are larger in most zeolites, indicating a positive excess free energy of the mixture. This observation can be attributed to water/ethanol interactions being less favorable than water/water interactions in the single-component adsorption of water at comparable loadings. The deviation from ideal behavior can be highly structure-dependent but no clear correlation with pore diameters was identified. Furthermore, our analysis also demonstrates the following: (1) accurate unary isotherms in the low-loading regime are critical for obtaining physically sensible activity coefficients; (2) the global regression scheme to solve for activity model parameters performs better than fitting activity models to activity coefficients calculated locally at each binary state point; and (3) including the dependence on adsorption potential offers only a minor benefit for describing binary adsorption at the lowest fugacities. Finally, the Margules activity model was found incapable of capturing the non-ideal adsorption behavior over the entire range of fugacities and compositions in all zeolites, but for conditions typical of solution-phase adsorption, RAST predictions using zeolite-specific or even bulk Margules parameters provide an improved description compared to IAST.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identifying Differential Equations in Fourier Domain (FourierIdent)

We investigate identifying differential equations in the frequency domain. Fourier analysis is an important tool in theoretical analysis and numerical solvers of differential equations, yet there is limited work in exploring this connection in the identification of differential equations. This paper aims to identify the underlying differential equation in the frequency domain, from a given single realization of the differential equation perturbed by noise. Such setting imposes difficulties which are different from other identification methods where computation is carried out in the physical domain. We propose several ways to mitigate the challenges arising from noise in data and large differences in the magnitudes of frequency responses. The main takeaways are that identifying differential equations solely in the frequency domain is challenging, the method we propose is based on a form of domain partitions in the frequency domain, and this method shows benefits for complex data even with high level of noise. We introduce a Fourier feature denoising, and define the meaningful data region and the core regions of features to reduce the effect of noise in the frequency domain and to enhance the accuracy in coefficient identification. The proposed method is tested on various differential equations with linear, nonlinear, and high-order derivative feature terms, and shows advantages on complex data with many frequency modes, even under high level of noise.

97 MATHEMATICS AND COMPUTING↗

Development of real-time density feedback control on MAST-U in L-mode

In this paper we report on the development and demonstration of density feedback control for MAST-U. Sinusoidal perturbations are used to measure the frequency response from a deuterium gas valve (actuator) to line-integrated core electron density measured by the interferometer (sensor). In the frequency range relevant for control design, only two system-identification experiments were needed to regress a first-order dynamic model. This control-oriented model informs the offline design of a proportional integral controller with the established loop-shaping controller design method. After offline verification of the controller implementation, control is demonstrated by experimentally tracking a staircase reference for the line-integrated electron density. This paper demonstrates the efficiency of controller design using system-identification and loop-shaping, providing reliable density control for MAST-U.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Analysis of biokinetic parameters reveals patterns in mercury accumulation across aquatic species

Mercury (Hg) is a potent neurotoxicant and poses a risk to human health through the ingestion of Hg-contaminated fish. Mercury, especially in its organic form methylmercury (MeHg), biomagnifies up food chains such that even small aqueous concentrations of Hg can result in significant concentrations of total Hg in fish. Understanding the ecological and human health risks associated with Hg and MeHg exposure requires an understanding of the factors that affect its bioaccumulation in aquatic species. We compiled estimates of three biokinetic parameters: uptake rate (k u ), assimilation efficiency (AE), and efflux rate (k e ). These parameters describe contaminant uptake from aqueous (k u ) and dietary (AE) exposure and the rate of excretion (k e ). We found parameter values for 38 and 34 different species of fish and aquatic invertebrates, respectively, and collected 502 parameter values in total. Here, we used a machine learning technique to establish the relationships between experimental and physiological variables and these parameter values. We found differences in which variables were associated with biokinetic parameter values for fish and aquatic invertebrates. The form of Hg was the most impactful variable, influencing values of all parameters except k u for invertebrates, for which aqueous exposure time was the only significant predicator variable. The parameter k e were the only values significantly influenced by more than one variable, with water type (freshwater, brackish, or marine), organism weight, and form of Hg significantly impacting parameter values for fish and/or invertebrates. To our knowledge, this study represents the most extensive review of biokinetic parameters of Hg and MeHg accumulation in aquatic organisms. Environmental parameters found to significantly impact Hg and MeHg bioaccumulation in past studies were not identified as important in our analyses across aquatic ecosystems and species. Our dataset and analysis reveal novel patterns that may help us better understand and manage Hg bioaccumulation.

54 ENVIRONMENTAL SCIENCES↗

Discovery of Pyridopyrimidinones that Selectively Inhibit the H1047R PI3Kα Mutant Protein

The H1047R mutation of PIK3CA is highly prevalent in breast cancers and other solid tumors. Selectively targeting PI3Kα H1047R over PI3Kα WT is crucial due to the role that PI3Kα WT plays in normal cellular processes, including glucose homeostasis. Currently, only one PI3Kα H1047R -selective inhibitor has progressed into clinical trials, while three pan mutant (H1047R, H1047L, H1047Y, E542K, and E545K) selective PI3Kα inhibitors have also reached the clinical stage. Herein, we report the design and discovery of a series of pyridopyrimidinones that inhibit PI3Kα H1047R with high selectivity over PI3Kα WT , resulting in the discovery of compound 17. When dosed in the HCC1954 tumor model in mice, 17 provided tumor regressions and a clear pharmacodynamic response. X-ray cocrystal structures from several PI3Kα inhibitors were obtained, revealing three distinct binding modes within PI3Kα H1047R including a previously reported cryptic pocket in the C-terminus of the kinase domain wherein we observe a ligand-induced interaction with Arg1047.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

D–MOPH–25: diverse MOF–molecule pairs for Henry’s constants prediction

Computational methods like grand-canonical Monte Carlo simulations and machine learning (ML) have accelerated metal–organic frameworks (MOF) exploration but are typically limited to a narrow range of adsorbates due to data availability and force field constraints. In this study, we introduce a dataset of diverse MOF–molecule pairs for Henry’s constant prediction, D–MOPH–25, which systematically explores a diverse chemical space by combining 113 molecular adsorbates with over 5000 MOF structures through an active learning process. D–MOPH–25 constitutes the most diverse adsorbate dataset used in any ML study of molecular adsorption in MOFs to date. Our workflow builds a benchmark for predicting Henry’s constants at 300 K, leveraging conformal prediction for uncertainty quantification. Assessment through Shannon entropy and uniform manifold approximation and projection confirms the comprehensiveness of D–MOPH–25 while highlighting the importance of robust classification to filter out unphysical data points in regression tasks. Although future enhancements in model architecture and sampling criteria could improve predictive performance, our dataset already spans the target space using only 2.31% of total possibilities. This comprehensive dataset facilitates assessment of model generalizability across adsorbate species and can establish a foundation for high-throughput MOF screening and ML-driven separation processes.

active learning↗

Discovery of BBO-11818, a Potent and Selective Noncovalent Inhibitor of (ON) and (OFF) KRAS with Activity against Multiple Oncogenic Mutants

Although KRAS G12C -specific inhibitors have been introduced, no approved targeted therapies exist for other clinically significant KRAS mutants, including KRAS G12D and KRAS G12V . We discovered BBO-11818, a potent, selective, orally bioavailable noncovalent pan-KRAS inhibitor capable of targeting multiple KRAS mutants in both the inactive GDP-bound (OFF) and active GTP-bound (ON) states. BBO-11818 binds in the Switch-II/Helix 3 pocket, inducing conformational changes incompatible with effector binding, and demonstrates high-affinity binding to mutant KRAS with strong selectivity over NRAS and HRAS. BBO-11818 potently inhibited MAPK signaling and cellular viability specifically in KRAS-driven lines and produced tumor regressions in KRAS-mutant xenograft models. Combination studies with anti–PD-1, anti-EGFR antibodies, and a RAS:PI3Kα breaker compound showed enhanced efficacy. BBO-11818 has entered phase I clinical trials for patients with various KRAS mutations in colorectal, pancreatic, and lung cancers (NCT06917079).

Stahlhut, Carlos [BridgeBio Oncology Therapeutics,↗

Substitution or Shared Utilization? Intrahousehold Vehicle Use in Mixed-Powertrain Households

While previous research has focused heavily on understanding the factors deriving alternative fuel vehicle adoption rates, there remains a significant gap in understanding how households distribute mileage across different powertrains. This study utilizes data from the 2022 Next Generation National Household Travel Survey to investigate vehicle miles traveled within a sample of 150 plug-in electric vehicle (PEV)-owning households (in which at least one battery electric vehicle is present), characterizing how different powertrains are integrated into daily mobility. Leveraging a Seemingly Unrelated Regression (SUR) framework the study jointly models the utilization of PEVs, hybrid electric vehicles (HEV), and internal combustion engine vehicles (ICEVs) while accounting for household-level substitution effects. The results provide evidence of an asymmetric substitution effect. In households with mixed-powertrain configurations, the ICEV captures a substantially higher share of household miles (compared with the PEV), acting as a utility sponge. Conversely, the model identifies specific socioeconomic and geographic cohorts that prioritize PEV as the primary household workhorse, indicating a systematic sorting effect. Although the sample size limits broader generalizability, these findings suggest that PEVs are used for frequent, specific routine-intensive roles, whereas the ICEV remains a specialized utility vehicle. These insights highlight distinct intrahousehold vehicle use behaviors that are often obscured by aggregate fleetwide statistics.

25 ENERGY STORAGE↗

Modeling dynamics of acute HIV infection incorporating density-dependent cell death and multiplicity of infection

Understanding the dynamics of acute HIV infection can offer valuable insights into the early stages of viral behavior, potentially helping uncover various aspects of HIV pathogenesis. The standard viral dynamics model explains HIV viral dynamics during acute infection reasonably well. However, the model makes simplifying assumptions, neglecting some aspects of HIV infection. For instance, in the standard model, target cells are infected by a single HIV virion. Yet, cellular multiplicity of infection (MOI) may have considerable effects in pathogenesis and viral evolution. Further, when using the standard model, we take constant infected cell death rates, simplifying the dynamic immune responses. Here, we use four models—1) the standard viral dynamics model, 2) an alternate model incorporating cellular MOI, 3) a model assuming density-dependent death rate of infected cells and 4) a model combining (2) and (3)—to investigate acute infection dynamics in 43 people living with HIV very early after HIV exposure. We find that all models qualitatively describe the data, but none of the tested models is by itself the best to capture different kinds of heterogeneity. Instead, different models describe differing features of the dynamics more accurately. For example, while the standard viral dynamics model may be the most parsimonious across study participants by the corrected Akaike Information Criterion (AICc), we find that viral peaks are better explained by a model allowing for cellular MOI, using a linear regression analysis as analyzed by R 2 . These results suggest that heterogeneity in within-host viral dynamics cannot be captured by a single model. Depending on the specific aspect of interest, a corresponding model should be employed.

60 APPLIED LIFE SCIENCES↗

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT↗

Evaluation of the Planetary Boundary Layer Height From ERA5 Reanalysis With MOSAiC Observations Over the Arctic Ocean

The planetary boundary layer height (PBLH) is a crucial indicator reflecting the region of the atmosphere characterized by continuous turbulence. Here, we use radiosonde and surface meteorological observations (4–7 times per day, year-round measurements) during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition to derive the PBLH (PBLH MOSAiC ), and further evaluate the PBLH from the ERA5 reanalysis (PBLH ERA5 ). Comparisons between PBLH MOSAiC and PBLH ERA5 from different perspectives reveal that: (a) The overestimation of PBLH ERA5 when the sea ice concentration is >90% is significant with the centered root mean squared error reaching up to 201 m; (b) The difference between the two products is notably pronounced in cold seasons, while it is comparatively diminished in warm seasons; (c) In neutral boundary layers, differences in PBLH ERA5 are larger compared with stable and convective boundary layers. In addition, the analysis of error sources indicates that the bias of PBLH ERA5 is sensitive to the bias of vertical thermal structure and wind speed profiles in ERA5 data sets in all conditions. Finally, we find a Random Forest model effectively reduces the bias of PBLH ERA5 with the index of agreement reaching up to 0.71 in the test data set, while a multiple linear regression demonstrates comparable performance to the Random Forest model.

54 ENVIRONMENTAL SCIENCES↗