Simultaneous estimation by partial totals for compartmental models
Simultaneous estimation procedure for parameters in multiple equation regression model
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Simultaneous estimation procedure for parameters in multiple equation regression model
Early battery life prediction models are most useful for R&D if they help us understand the early changes in battery electrochemical response that correspond with long-term degradation and failure. Linear regression models such as Fused lasso and Partial Least Squares can fit coefficients directly to high-dimensional electrochemical data like capacity-voltage and ΔV–state-of-charge, i.e., Q(V) and ΔV(SOC) curves, learning coefficients that can be physically interpreted. We leverage the ISU-ILCC battery aging data set to learn high-dimensional coefficients for early battery life prediction from traditional slow-rate capacity check data, demonstrating learning on Q(V), d Q· d V −1 , and ΔV(SOC) curves. A thorough study on the dependence of coefficient values on train/test size and data preprocessing methods is made, demonstrating the reliability of high-dimensional regression approaches unless very small amounts of data are used for model training. For this data set, coefficients from Q(V) and d Q· d V −1 models highlight changes in electrode stoichiometry due to lithium loss, while ΔV(SOC) coefficients highlight changes in positive electrode diffusivity due to particle cracking as well as electrode stoichiometry shifts. By directly interpreting the coefficients of a regression model, we make physical insights into battery degradation mechanisms without requiring the assumptions of traditional battery data analysis methods.
A simulation can stand its ground against an experiment only if its prediction uncertainty is known. The unknown accuracy of interatomic potentials (IPs) is a major source of prediction uncertainty, severely limiting the use of large-scale classical atomistic simulations in a wide range of scientific and engineering applications. Here we explore covariance between predictions of metal plasticity, from 178 large-scale (~10 8 atoms) molecular dynamics (MD) simulations, and a variety of indicator properties computed at small-scales (≤10 2 atoms). All simulations use the same 178 IPs. In a manner similar to statistical studies in public health, we analyze correlations of strength with indicators, identify the best predictor properties, and build a cross-scale “strength-on-predictors” regression model. This model is then used to estimate regression error over the statistical pool of IPs. Small-scale predictors found to be highly covariant with strength are computed using expensive quantum-accurate calculations and used to predict flow strength, within the statistical error bounds established in our study.
This paper presents a discussion of the applicability of neural networks in the identification and control of dynamic systems. Emphasis is placed on the understanding of how the neural networks handle linear systems and how the new approach is related to conventional system identification and control methods. Extensions of the approach to nonlinear systems are then made. The paper explains the fundamental concepts of neural networks in their simplest terms. Among the topics discussed are feed forward and recurrent networks in relation to the standard state-space and observer models, linear and nonlinear auto-regressive models, linear, predictors, one-step ahead control, and model reference adaptive control for linear and nonlinear systems. Numerical examples are presented to illustrate the application of these important concepts.
Astronauts are exposed to a unique set of stressors in spaceflight. Microgravity, isolation, confinement, and environmental and operational hazards: all of these can impact sleep, vigilant attention, and alertness, which are critical to mission success. In this paper, we seek to understand the most important predictors of alertness over the course of a space mission, using self-reported, cognitive, and environmental data collected from 24 astronauts on 6-month missions to the International Space Station (ISS). Alertness was repeatedly and objectively assessed on the ISS with a brief 3-minute Psychomotor Vigilance Test (PVT) that is highly sensitive to sleep deprivation. To relate PVT performance to time-varying and sparsely-measured environmental, operational, and psychological covariates, we propose a n ensemble prediction model comprising of linear mixed effects regression, random forest, and functional concurrent regression models. An extensive cross-validation procedure reveals that this ensemble outperforms any one of its components alone. We also discover that a participant’s past performance, reported fatigue and stress, and temperature and radiation exposure were among the most important variables associated with alertness. This method is broadly applicable to environmental studies where the main goal is accurate, individualized prediction involving a mixture of person-level traits and irregularly measured time series.
BACKGROUND The Privacy Act of 1974 regulates the use a nd disclosure of personally identifiable information by US Federal agencies. The Act applies to biographical, financial, a nd other identity-linked information, a s well a s personal health information (PHI). As such, the use of astronaut PHI is limited to authorized personnel for preapproved uses, with data reporting often limited to aggregated information about groups. These limitations on the use a nd reporting of astronaut PHI complicates surveillance efforts, wherein epidemiologists a t the National Aeronautics and Space Administration (NASA)monitor the incidence of targeted health conditions in the astronaut population, or to discover emerging trends of aging and disease. Stratification on one or more covariates –particularly time-period, sex, a nd mission participation –can lead to extremely small datasets such that the reporting of results is potentially attributable to individuals. An additional challenge is the small size of the astronaut population, both in terms of numbers of individuals a s well a s in terms of density of exposure time. Such small datasets yield volatile rate estimates that are difficult to interpret. To a id the epidemiological surveillance efforts, a surveillance tool is required that can (a) satisfy the need for rapid computation of condition-specific incidence and mortality rates; (b) improve the statistical estimates of these estimated rates; and (c) maintain astronaut privacy. Here we describe a nd demonstrate such a tool. METHODS We devised a system that models incidence a nd mortality rates rather than calculating them directly. This ha s the advantage of using all the available data to derive the estimates, lea ding to rates that a re not attributable to any one individual, a nd a re a s numerically stable a s they can be given the extremely limited data. The system models disease endpoints using a Poisson regression model with exposure density (measured in person-years) a s a n offset term. By doing so the model is estimating event counts per person-year, equivalent to modeling the rates directly. It uses a standard (pre-specified)set of covariates; the system does not engage in “model-building” as model parsimony is not the goa l. Instead, it is explicitly recognized that if a covariate is not statistically significant a nd not a confounder then it will likely have very little effect on the estimate of the incidence a nd mortality rates. Users are able to specify the disease endpoint of interest and the covariates over which they would like to stratify. The system then uses the resulting model to compute the estimated rates for the user-chosen configuration of variables as visualizes those either over an age range within a specified time-period, or over time for astronauts with a specified age range. RESULTS The first iteration of the tool computes incidence a nd mortality rates for cardiovascular conditions and cancers. Code ha s been developed to retrieve the appropriate data from the IMPALA analysis platform, compute the models for incidence a nd mortality, a nd then use those models to generate the corresponding rate curves. A companion graphical user interface allows the user to specify the curves and visualize the results. CONCLUSIONS It is important to note that the rapid surveillance tool described here is neither meant to be a definitive assessment of the incidence or mortality of any particular disease or condition in the astronaut population, nor is it meant to be used for research purposes. Rather, it is meant as an early indicator that in-depth investigation may be warranted. By automating a repetitive process and leveraging carefully curated astronaut health outcomes, the tool makes possible a rapid “first look” into known areas of concern, and, if used judiciously, may surface new areas of concern for long-term astronaut health. This work is supported in part by the Translational Research Institute for Space Health (TRISH) through NASA Cooperative Agreement NNX16AO69A.
The static response of sea level to the forcing of atmospheric pressure, the so-called inverted barometer (IB) effect, is investigated using TOPEX/POSEIDON data. This response, characterized by the rise and fall of sea level to compensate for the change of atmospheric pressure at a rate of -1 cm/mbar, is not associated with any ocean currents and hence is normally treated as an error to be removed from sea level observation. Linear regression and spectral transfer function analyses are applied to sea level and pressure to examine the validity of the IB effect. In regions outside the tropics, the regression coefficient is found to be consistently close to the theoretical value except for the regions of western boundary currents, where the mesoscale variability interferes with the IB effect. The spectral transfer function shows near IB response at periods of 30 degrees is -0.84 +/- 0.29 cm/mbar (1 standard deviation). The deviation from = 1 cm /mbar is shown to be caused primarily by the effect of wind forcing on sea level, based on multivariate linear regression model involving both pressure and wind forcing. The regression coefficient for pressure resulting from the multivariate analysis is -0.96 +/- 0.32 cm/mbar. In the tropics the multivariate analysis fails because sea level in the tropics is primarily responding to remote wind forcing. However, after removing from the data the wind-forced sea level estimated by a dynamic model of the tropical Pacific, the pressure regression coefficient improves from -1.22 +/- 0.69 cm/mbar to -0.99 +/- 0.46 cm/mbar, clearly revealing an IB response. The result of the study suggests that with a proper removal of the effect of wind forcing the IB effect is valid in most of the open ocean at periods longer than 20 days and spatial scales larger than 500 km.
Estimation of total body water (T) from bioelectrical resistance (R) is commonly done by stepwise regression models with height squared over R, H(exp 2)/R, age, sex, and weight (W). Polynomials of H(exp 2)/R have not been included in these models. We examined the validity of a model with third order polynomials and W. Methods: T was measured with oxygen-18 labled water in 27 subjects. R at 50 kHz was obtained from electrodes placed on the hand and foot while subjects were in the supine position. A stepwise regression equation was developed with 13 subjects (age 31.5 plus or minus 6.2 years, T 38.2 plus or minus 6.6 L, W 65.2 plus or minus 12.0 kg). Correlations, standard error of estimates and mean differences were computed between T and estimated T's from the new (N) model and other models. Evaluations were completed with the remaining 14 subjects (age 32.4 plus or minus 6.3 years, T 40.3 plus or minus 8 L, W 70.2 plus or minus 12.3 kg) and two of its subgroups (high and low) Results: A regression equation was developed from the model. The only significant mean difference was between T and one of the earlier models. Conclusion: Third order polynomials in regression models may increase the accuracy of estimating total body water. Evaluating the model with a larger population is needed.
Coal fly ash is a high volume waste material that is discarded in landfills and surface water impoundments across the U.S. and is also widely recycled for a variety of applications. The leaching of potential of contaminants of concern, such as arsenic (As) and selenium (Se), is often the driver of risk assessments for coal ash disposal and reuse. The extent of leachable As and Se depends on several factors related to environmental conditions and fly ash characteristics. Previous studies employed various methods to delineate the concentration, chemical form, and distribution of As and Se in fly ash materials. However, few studies have attempted to directly correlate these properties to mobilization parameters relevant to disposal and reuse. Instead, the coal residuals industries often rely upon standardized leaching protocols that can be laborious or involve hazardous chemicals. The goals of the project were to: 1) Develop and evaluate a characterization protocol that can be used to screen fly ash samples for leachability of As and Se; 2) Characterize As, Se, and associated constituents of fly ash particles at multiple length scales (nanometer to micrometer) to determine if elemental associations differ as a function of the resolution of characterization; and 3) Establish a predictive model for the chemical composition of coal ash produced annually at major U.S. coal fired power facilities on 50-year national coal supply records. For the first objective, we performed leaching experiments with 52 fly ash samples collected from 15 different U.S. power plants and representing coal feedstocks from the three major domestic coal regions. For this work, we assessed the mobilization potential of As and Se in fly ash based on standardized leaching protocols and performed multivariate and lasso regression analyses to explore correlations of leachable As and Se contents with characteristics such as major element contents, loss on ignition (LOI) and pH. The results of regression models indicated that major elements (Fe, Ca, Al) for a wide range of fly ashes can serve as predictor variables for the leaching potential of As, but not for Se. LOI and pH were not important predictive variables in the models. Both regression approaches resulted in relatively strong fits for leachable As (correlation coefficient R 2 = 0.78 for both models) compared to models for leachable Se (R 2 = 0.49). Overall, these results suggest that correlation models combined with on-site elemental analysis with portable analyzers may enable a screening method for leachable As in coal ash. For the second objective, we utilized nanoscale 2-D imaging (30-50 nm spot size) with the Hard X-ray Nanoprobe (HXN) in combination with microprobe X-ray capabilities (~5 µm resolution) to determine As and Se elemental associations in fly ash particles. Speciation of As and Se was also measured at the nano- to microscale with X-ray absorption spectroscopy. The enhanced resolution of HXN showed As and Se that were diffusely located around or comingled with Ca- and Fe-rich particles. The results also showed nanoparticles of Se attached to the surface of fly ash grains. Overall, a comparison of As and Se species across scales highlights the heterogeneity and complexity of chemical associations for these trace elements of concern in coal fly ash. For the final objective, we developed a predictive model for major element composition of coal ash in reserve at disposal sites of major U.S. coal fired power plants. This model was constructed from coal purchase records of 705 power stations from 1973-2022 and was trained on coal ash composition data showing that coal ash elemental composition is strongly associated with the source of feedstock coal. The model showed regional shifts in the major element contents of ash produced by power plants in the last 50 years, particularly for calcium and iron (expressed as %CaO and %Fe 2 O 3 ), as coal-fired power stations changed their source of coal over this time frame. Our approach enables an estimation of coal ash chemical composition that is stored in waste impoundments at individual power stations. Such information can help delineate the regional market potential for material applications that would utilize coal ash harvested from disposal sites across the U.S.
Laser powder bed fusion (LPBF) Ti-6Al-4V is widely studied for use in structural applications in aerospace and medical industries, but mechanical anisotropy and microstructural inhomogeneity prohibits its wider adoption. Although successful microstructure prediction models have been developed, a remaining challenge is their limited integration across length/time scales and validation by experimental studies. Here, this work proposes a physics-augmented machine learning surrogate model to unite predictions of LPBF temperature, β phase morphology and texture, and α/α’ formation into a single framework that is calibrated and validated with experiments. First, a phase field (PF) model of the martensitic β→α’ transformation is developed and calibrated using data from in-situ synchrotron cyclic heating/cooling studies quantifying the variation of α phase fraction with time. In parallel, an established finite difference-Monte Carlo (FDMC) model predicts the part-scale temperature profile and β grain formation during solidification. A dataset is developed using LPBF cyclic temperature descriptors from the FDMC model as inputs and corresponding α/α’ phase fraction and width from the PF model as outputs. Five machine learning (ML) regression models are tested and optimized, having mean absolute error in testing ≤ 4 %, and the k-nearest neighbors (KNN) model is selected as the best performing. The KNN model is called at the nodal level during post-processing of the FDMC model to replace and downscale the response of the PF model. The combined agility and accuracy of the hybrid FDMC-ML model enables part-scale microstructure predictions that can be further used for property predictions to accelerate AM process optimization.
Data from the first Earth Resource Technology Satellite (LANDSAT-1) multispectral scanner (MSS) were used to develop three plant canopy models (Kubelka-Munk (K-M), regression, and combined K-M and regression models) for extracting plant, soil, and shadow reflectance components of cropped fields. The combined model gave the best correlation between MSS data and ground truth, by accounting for essentially all of the reflectance of plants, soil, and shadow between crop rows. The principles presented can be used to better forecast crop yield and to estimate acreage.
Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.
Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.
In recent years, the occurrence and impact of inland and coastal flood events have become more frequent and damaging, especially within agricultural fields, due to the global climate change and consistent sea level rise. Monitoring and measuring the magnitude of flood events in a timely manner and assessing the subsequent crop damages accurately are precursors in minimizing detrimental consequences that could potentially lead to a global food security crisis. Traditional gauge-based measurements with sophisticated hydrological models are capable of monitoring flood events precisely but limited within the smaller spatial extent, time-consuming, and costly. In recent decades, advancement in airborne- and satellite-based remote sensing technologies offering products at a daily global spatial extent with various spectral resolution helps address the shortcomings of the traditional in situ approaches in flood monitoring. Furthermore, the methods such as classification and band ratioing using remote sensing products are simple and effective in assessing flood-induced agricultural damages. The combination of remote sensing products and geographic information systems along with the current development in web mapping, users now can get near real-time flood monitoring and crop damage assessments, albeit dependent upon the quality of available data. A case study to quantify the impact of the 2011 Missouri Mississippi River flooding on the surrounding cornfield was performed through a regression model. The model was trained using historical daily NDVI and corn yield across Nebraska and Missouri, and the overall accuracy in estimating corn yield was about 90%. The method implemented in this localized case study could be extended at a larger geographical scale.
Biofibers serve as effective reinforcements for neat polylactic acid (PLA) in biocomposites, offering an attractive opportunity to decarbonize the manufacturing sector of the United States by displacing fossil-based reinforcement fibers such as carbon fibers. Also, biofiber production can stimulate economic growth in rural economies, fueling sustainable development. PLA resins are commonly compounded with biofibers to create biocomposites suitable for additive manufacturing. PLA-biofiber composites often exhibit better overall material properties than neat (pure) PLA, but the associations between biofiber properties and the material properties of their biocomposites remain largely unexplored. Hence, this research delves into a comprehensive exploration of diverse biofibers, scrutinizing their physical and chemical attributes, including size, shape, ash content and biochemical composition. The study meticulously analyzes the flow properties of each biofiber and elucidates the ultimate tensile strengths and Young's modulus of corresponding biocomposite samples. Noteworthy correlations between biofiber and biocomposite tensile properties are uncovered, shedding light on critical interrelationships. The study introduces an approach employing regression models to predict the ultimate tensile strength and Young's modulus of biocomposites. These models, validated with a cross-validation technique, exhibit remarkable predictive accuracy, particularly in estimating ultimate tensile strength. © 2024 Oak Ridge National Laboratory managed by UT-Battelle, LLC and The Author(s). Polymer International published by John Wiley & Sons Ltd on behalf of Society of Chemical Industry.
This paper summarizes recent results on applying the method of partial least squares (PLS) in a reproducing kernel Hilbert space (RKHS). A previously proposed kernel PLS regression model was proven to be competitive with other regularized regression methods in RKHS. The family of nonlinear kernel-based PLS models is extended by considering the kernel PLS method for discrimination. Theoretical and experimental results on a two-class discrimination problem indicate usefulness of the method.
With growing freshwater scarcity, direct potable reuse (DPR) systems that reclaim wastewater for drinking are becoming increasingly important for sustainable water supply. Reliable operation requires minimizing downtime in ultrafiltration (UF) units, where membrane fouling leads to elevated trans-membrane pressure (TMP). This study develops data-driven regression models based on random forest (RF) and autoregressive (AR) approaches to forecast the initial TMP at the start of each UF filtration cycle in a pilot-scale DPR system. The RF model consistently outperforms baseline methods, including historical mean, last observation carried forward, and AR models, across multiple forecast horizons, achieving the lowest root mean square error. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent input variables across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is assessed for both direct and recursive RF modelling approaches. The proposed RF framework establishes a robust foundation for predictive monitoring and real-time optimization of UF operations, supporting sustainable and reliable water reuse.
Wheat production in Brazil is insufficient to meet domestic demand and falls drastically in response to adverse climate events. Multiple, agro-climate-specific regression models, quantifying regional production variability, were combined to estimate national production based on past climate, cropping area, trend-corrected yield, and national commodity prices. Projections with five CMIP6 climate change models suggest extremes of low wheat production historically occurring once every 20 years would become up to 90% frequent by the end of this century, depending on representative concentration pathway, magnified by wheat and in some cases by maize price fluctuations. Similar impacts can be expected for other crops and in other countries. This drastic increase in frequency in extreme low crop production with climate change will threaten Brazil's and many other countries progress toward food security and abolishing hunger.