Search NASA⌕ Search

SEARCH · Search NASA

Results for “Cross-validation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Multi-trait multi-environment genomic prediction strategies for Miscanthus sacchariflorus

Genomic selection holds the potential to serve as a strategic tool to enhance the genetic gain of complex traits in Miscanthus breeding programs. The development of improved cultivars requires their assessment for various traits across diverse environments to ensure suitable overall performance. Hence, the multi-trait multi-environment (MTME) genomic prediction (GP) models offer an opportunity to improve selection accuracy. This study aims to evaluate the potential of five GP models: (1) three MTME models including genotype-by-trait-by-environment interaction (G×E×T) and (2) two single-trait multi-environment (STME) models (with and without G×E interaction). A Miscanthus sacchariflorus population comprising 336 genotypes evaluated in three environments and scored for four traits (biomass yield YDY, total culm number TCM, average internode length AIL, and culm node number CNN) was analyzed. The predictive ability of the models was evaluated considering three cross-validation schemes resembling realistic scenarios (CV1: predicting new genotypes, CVP: predicting missing traits in a given environment, and CV2: predicting partially observed genotypes). On average, in all cross-validation schemes compared to the STME the predictive ability of the MTME models was 10% to 70% higher for TCM and AIL. On the other hand, for YDY and CNN, both STME models performed similarly or slightly better (between 5 to 64%) than the MTME models in most environments. While the MTME models were not successful for all traits when compared to their STME counterparts, MTME models improved the prediction of the performance of genotypes that were untested across environments or lacked trait information in a specific environment. Overall, our study suggests that MTME GP models can be implemented in Miscanthus breeding programs to improve the predictive ability of the complex traits, shorten breeding cycles, and accelerate selection decisions.

genomic prediction (GP)↗

Quantum AI Based Enhanced Detection of Dementia

Quantum computing has the potential to significantly improve the early detection of Alzheimer's Disease and Related Dementias (ADRD). Quantum-enhanced machine learning can be used to perform an early screening of Alzheimer's disease using brain imaging data based on dataset of MRI scans from both healthy individuals and those diagnosed with Alzheimer's. This study aims to demonstrate the potential of quantum transfer learning to enhance the performance of the classical deep learning model for dementia detection. Using the MRI sagittal images available in the OASIS-2 (64 demented and 72 non-demented subjects between 60 and 96 years), we show how quantum techniques can transform a suboptimal classical model into a more effective solution for dementia detection, highlighting their potential impact on advancing healthcare technology. We begin with a simple classical deep learning model with a significantly smaller number of parameters, which gives suboptimal performance on the problem. Then, we apply different configurations of quantum transfer learning based on the pre-trained weak classifier (Figure 1). We fix the weak classifier's initial convolutional layers at their fixed pre-trained parameters and replace the last set of dense layers with a dressed quantum circuit (DQN), which we train to enhance performance. We performed 4-fold cross-validation for both the classical and the hybrid quantum models and trained them using Pennylane's `default.qubit' simulator and IonQ's Aria-1 simulator (noisy simulation). We showed that with significantly fewer parameters, the quantum transfer learning-based hybrid models showed significant performance enhancement over the base weak classical deep learning model for dementia detection. To classify between a demented and non-demented subject, the accuracy of quantum-based AI methods improved by 6 to 14% compared to classical methods. The sensitivity of the models improved by 4 to 17%. This shows that there are fewer chances of misclassifying demented patients. Figure 2 compares the performance of the hybrid quantum models and their base classical model, and Table 1 summarizes the results. We illustrated that with assistance from quantum machine learning, it is possible to enhance detection for dementia based on brain images. This shows the potential for practical utility of quantum computing in ADRD research.

Bhowmik, Sounak [University of Tennessee, Knoxvill↗

Molten Halide Salt Surface Tension: Methods and Correlations

Here, this paper reviews various methods for studying surface tension and their applicability to fluoride and chloride molten salt systems, including a comparison of benefits and drawbacks. Such a comparison aids in experiment design based on desired factors such as scale, accuracy, and repeatability. A detailed review is presented for existing literature data regarding the surface tension of molten fluoride and chloride salts. These reference data were compiled and analyzed to determine cross-validated correlation equations for several alkali and alkaline earth fluoride and chloride salts as functions of temperature. These correlations are necessary for reliable multiphysics modeling approaches as well as accurate design and analysis of multiphase molten salt phenomena such as gas sparging and bubble formation/transport. This analysis supports the development of the thermophysical arm of the Molten Salt Thermal Properties Database (MSTDB-TP) managed by Oak Ridge National Laboratory.

Chloride↗

In-situ sensor monitoring of multi-class gas porosity formation in laser powder bed fusion using convolutional neural network

In-situ monitoring of defect formation remains a significant challenge in the laser powder bed fusion (LPBF) process. Recent advances have enabled real-time defect detection with machine learning and in-situ sensing technologies; however, most studies focus on binary classification of keyhole pores, limiting nuanced multi-class pore differentiation and formation mechanisms. This work introduces a multi-class pore detection framework (no pore, small pores < 15 µm, and large pores > 15 µm) by leveraging photodiode sensor data alongside high-fidelity synchrotron X-ray imaging. The 15 µm threshold is selected to distinguish between two fundamentally different defect mechanisms, following the physical size-mechanism boundary established by prior high-resolution synchrotron X-ray characterization of Al6061 LPBF. Distinguishing these classes is critical because large keyhole pores are structurally detrimental, whereas small gas pores are often benign, requiring different process control strategies. Thermal emission monitoring data collected simultaneously with high-speed X-ray imaging at the Stanford Synchrotron Radiation Lightsource (SSRL), are correlated with subsurface melt pool dynamics to establish ground truth. Continuous Wavelet Transform (CWT) with optimized parameters converts the photodiode time-series signals into time–frequency images, facilitating feature extraction. Convolutional Neural Networks (CNN) are then applied for real-time multi-class pore classification in an average inference time of 1 ms per signal window. It achieves 79% accuracy and an Area Under the Receiver Operating Characteristic curve (AUC ROC) score of 0.89 with five-fold cross-validation. The results demonstrate that coupling CWT-based feature engineering with CNN architecture enables reliable multi-class pore detection in Al6061 builds using affordable in-situ sensors. This approach advances scalable and affordable quality assurance in additive manufacturing by moving beyond binary defect detection toward more nuanced classification of porosity mechanisms with in-situ sensors and machine learning.

Laser powder bed fusion, Multi-class pores, In-sit↗

Projection-based multifidelity linear regression for data-scarce applications

Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. Here, the proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2% – 12% improvement in median accuracy and a higher R 2 score relative to single-fidelity methods at comparable computational cost.

data augmentation↗

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES↗

Toward equitable environmental exposure modeling through convergence of data, open, and citizen sciences: an example of air pollution exposure modeling amidst increasing wildfire smoke

Exposure modeling is critical in environmental epidemiology and human health but may face challenges (e.g., skewed data, unequal error, context-insensitive validation, and computational demands). Modeling decisions reflect the intended use of the models and the values that modelers prioritize. We aimed to provide a conceptual framework and machine learning (ML) modeling protocols that address these issues. With 500m-gridded hourly PM 2.5 and O 3 levels in Illinois before, during, and after the 2023 Canadian wildfire season as a motivating example, we conducted modeling experiments to evaluate modeling methods, guided by three domains we propose based on theories of science: 1) Data Diversity, leveraging open and citizen science data to enhance inclusivity, parsimony, and representativeness; 2) Equitable Accuracy, ensuring fairly distributed uncertainties across subpopulations; and 3) Sustainable Modeling, balancing accuracy with reducing computational demands to promote accessibility for under-resourced researchers. Here, we found that ML with publicly available data can achieve high accuracy. Depending on methods, performance may vary substantially, even with identical input data. Large but skewed data may reduce performance. Misuse of cross-validation protocols can underestimate prediction error; although we observed R 2 s of ∼98 %, the modeled estimates varied significantly, indicating the need for careful model validation. By using new modeling protocols including representativeness-considered training and validation data and a new loss function, we achieved high agreement between estimates and ground-based measurements (e.g., R 2 = ∼90 % for PM 2.5 ; ∼80 % for O 3 ), equally distributed errors across sociodemographic strata and urban–rural divides, and reduction in computation time—from several weeks or months to a few days.

Exposure assessment↗

Numerical simulation of vortex-induced vibration response of a single IEA 10-MW wind turbine blade

Three-dimensional simulation of vortex-induced vibration (VIV) of a single International Energy Agency (IEA) 10-MW reference wind turbine blade with a length of 97.325 m is performed using the ExaWind stack, an open-source suite of codes. This study aims to illustrate the spanwise VIV response characteristics and cross-validate the results with an existing commercial framework. Five near-body meshes and three time steps are selected for the convergence study. To improve computational efficiency, several VIV triggering methods are also compared to shorten the VIV development period. The ExaWind-based VIV simulation strategy for a single IEA 10-MW blade is determined. First, the modal shape is validated against published results. Then, spanwise VIV responses of four blade configurations under a fixed and varied incoming flow velocity are analyzed. Results show that the VIV response is dominated by the first edgewise (second overall) mode. Little first-mode contributions appear near the second-mode node, producing a pi phase jump, and a higher harmonics response occurs near the blade root. Rotational degrees of freedom are minor compared with translational motion. The response versus reduced velocity is analyzed, showing a two-branch behavior similar to that of VIV for a bluff cylinder. Across all tested cases, the dominant frequency remains locked to the natural frequency of the second mode with no observed desynchronization. A mild deviation is observed for the case of 90-degree pitch and 310-degree azimuth rotation near a reduced velocity of 6, which will be examined with additional cases in future work. These findings indicate that severe VIV responses can arise under specific configurations and flow conditions, thereby increasing the potential for VIV fatigue damage and requiring greater attention during operation.

17 WIND ENERGY↗

Optimization of well design and CO 2 injection strategy for risk reduction in Class VI geological carbon sequestration wells

The safety and durability of Class VI wells are critical for geological carbon sequestration (GCS). However, current GCS operations face unique challenges: unlike traditional Class II wells, Class VI CO 2 injection wells operate at rates up to 100 times higher, dramatically increasing the risk of wellbore leakage and structural compromise due to severe temperature drops and associated mechanical stresses. Despite existing guidelines on material selection, there remains a substantial gap in understanding how rapid CO 2 injection rates, low surface temperatures, and variable reservoir conditions interact to threaten long-term well integrity. This study presents a comprehensive, original workflow integrating advanced analytical and numerical models for both well flow and well integrity analysis. By systematically simulating a wide range of field-relevant scenarios—including variations in injection rate, CO 2 temperature, and reservoir pressure—this work provides the first cross-validated assessment of cooling effects on wellbore. The results reveal that extreme temperature drops, up to 60 °C, can occur under high injection rates, particularly in depleted reservoirs, significantly increasing the risk of cement failure. Building on these insights, the study proposes innovative, practical well design and operational strategies, including ductile cement formulations, pre-stressing techniques, advanced insulation coatings, and proactive management of injection rates. The safety of Class VI well extends beyond simply using CO 2 resistant materials. Cement materials should possess optimal thermo-hydraulic-mechanical-chemical properties for effective performance. This work provides a scientific basis for optimizing Class VI well designs, with direct benefits for minimizing environmental risk, lowering operational costs, and enhancing the long-term reliability of GCS.

25 ENERGY STORAGE↗

A comparative study of multimodal data fusion strategies for planetary spectroscopy

Integrating heterogeneous data sources can improve scientific inference when different modalities capture complementary information, but doing so is challenging in high-dimensional, small-sample settings. In spectroscopy for planetary exploration, Laser-Induced Breakdown Spectroscopy (LIBS), Raman Spectroscopy (Raman), Visible Infrared Spectroscopy (VISIR), and Mid-Infrared Spectroscopy (MIR) each examine different aspects of composition and mineralogy, raising fundamental questions about when and how data fusion improves predictive performance. Using a Mars-relevant set of geologic standards with measurements from all four modalities, we present a rigorous systematic evaluation of four data fusion strategies: low-level (data) fusion, mid-level (feature) fusion, high-level (decision) fusion, and residual-boosting (sequential) fusion. We assess performance in predicting oxide composition via nested cross-validation and corrected significance testing to evaluate whether data fusion improves upon single-modality baselines. We show that data fusion does not uniformly improve accuracy, and that observed gains are modest, oxide-dependent, and sensitive to modality and model structure. To move beyond aggregate accuracy metrics, we use model coefficients, permutation importance, and residual gain analysis to examine how the fusion models weight individual modalities and to identify patterns of apparent complementarity or redundancy. Though focused on spectroscopy for planetary exploration, our framework for data fusion evaluation and interpretation extends to other scientific domains with heterogeneous and scarce data and provides a principled approach evaluating data fusion strategies, interpreting modality contributions, and understanding tradeoffs among data fusion strategies.

97 MATHEMATICS AND COMPUTING↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Rigor and Reproducibility in Electrocatalysis: Best Practices for Operando Studies

Operando measurements have rapidly expanded the scope of electrocatalysis by enabling direct observation of catalytic interfaces under working conditions and by linking structural, compositional, and spectroscopic observables to activity and selectivity. However, the growth of operando methods has outpaced the adoption of broadly shared experimental standards, creating persistent challenges in reproducibility, interpretation, and comparison across laboratories and platforms. This perspective synthesizes discussions from the 2025 National Science Foundation Workshop on Rigor and Reproducibility in Electrocatalysis and outlines a practical framework for the rigorous use of operando measurements in electrocatalysis. We highlight three recurring needs: careful implementation of complex methods to avoid overinterpretation; recognition that (subtle) differences in reactor architecture, hydrodynamics, and electrical boundary conditions can alter apparent kinetics and selectivity; and transparent reporting standards that enable meaningful cross-comparison without constraining measurement-specific cell innovation. Focusing on widely used techniques (including X-ray and vibrational spectroscopies, mass spectrometry, and electron microscopy), we discuss technique-specific pitfalls, cross-validation strategies, and recurring platform-agnostic considerations such as mass transport, current distribution, temporal-resolution mismatches, and catalyst evolution. This Perspective aims to strengthen the mechanistic inference and improve the reproducibility, comparability, and predictive value of operando electrocatalysis research.

X-ray absorption spectroscopy↗

Autonomous Nanoparticle Synthesis Guided by In Situ Multiscale Structural Characterization

Autonomous synthesis platforms promise rapid exploration of vast parameter spaces; yet, integrating in situ structural characterization in closed-loop synthesis optimization remains challenging. We demonstrate a realization of such a closed-loop platform coupled with a droplet-flow microreactor, in situ X-ray scattering methods (SAXS/WAXS), and Gaussian process optimization to synthesize citrate-reduced Au nanoparticles with targeted characteristics. The system efficiently explored ∼19,000 synthesis recipes through 365 experiments, achieving precise control over size (4–60 nm) and polydispersity (σ < 0.11) across large citrate/gold ratios, exceeding traditional synthesis boundaries (1–10). Beyond confirming classical Turkevich–Frens trends, partial-dependence analysis revealed strong nonlinear coupling among precursor, citrate, and pH effects. Combining quantitative SAXS/WAXS analysis with electron microscopy characterization, we uncovered that crystallite size (d c ) and particle size (d) follow d c = 0.18d + β, where synthesis chemistry controls the intercept β while maintaining a universal slope. This parallel-band structure enables independent tuning of crystallite domain size at fixed particle diameter through a combination of chloride, gold precursor, citrate, and pH contributions (cross-validated Spearman ρ = 0.7 ± 0.1). High-resolution electron microscopy shows multiple lattice-fringe orientations within single particles, directly confirming polycrystalline domains and the ability to tune d c at the fixed d. The platform’s validation includes indistinguishable static versus flowing measurements, stable droplet transport at 100 °C, and <5% run-to-run variation, establishing a robust framework for mapping and controlling multiscale nanoparticle structure across expansive chemical spaces. In conclusion, the developed closed-loop platform can be applied to a borad range of nanosyntheis processes.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Taking three-dimensional x-ray diffraction (3DXRD) from the synchrotron to the laboratory scale

Three-dimensional x-ray diffraction (3DXRD), a rotating x-ray diffraction technique, is a powerful tool for studying the micromechanical behavior of polycrystalline materials, capable of measuring the volume, position, orientation, and strain of thousands of grains simultaneously. However, its application has been historically limited to synchrotron facilities. Here, we present the first demonstration of laboratory-scale 3DXRD (Lab-3DXRD) using a liquid-metal-jet source. Lab-3DXRD achieves accuracy comparable to synchrotron-based 3DXRD, as validated against laboratory diffraction contrast tomography (LabDCT) and synchrotron-3DXRD. Over 96% of the grains detected with Lab-3DXRD are cross-validated, particularly for coarse grains (> ~60 μm), while the results suggest that finer grains should be accessible by taking advantage of high-efficiency detectors. We further demonstrate that its sensitivity to finer grains is enhanced by incorporating pre-characterization into the analysis. This study establishes Lab-3DXRD as a practical alternative to synchrotron techniques, making 3DXRD accessible to a wider range of academic and industrial researchers.

characterization and analytical techniques↗

Multi-diagnostic characterization of laser-produced tin plasmas for EUV lithography

We present a comprehensive characterization of laser-produced tin (Sn) plasmas relevant to extreme ultraviolet (EUV) lithography using a multi-diagnostic suite integrated into the new experimental platform, “SparkLight.” Tin plasmas are generated by irradiating a continuously moving tin-coated wire with laser pulses (1064 nm, 10 ns, up to 5.7 × 10 10 W/cm 2 ) and probed via coherent Thomson scattering, laser interferometry, and EUV emission spectroscopy. Thomson scattering measurements reveal electron temperatures and densities that decay with distance from the target. Densities derived from Thomson scattering are cross-validated against laser interferometry, showing excellent agreement. Correlating the results of these laser diagnostics with spatially resolved EUV spectroscopy suggests that the bulk of useful EUV emission originates within 150 μm of the target and is generated under suboptimal plasma conditions. This work demonstrates a practical integrated approach for plasma characterization in EUV source development.

Musikhin, S. [Princeton Plasma Physics Laboratory ↗

Enhancing dimensionality prediction in hybrid metal halides via feature engineering and class-imbalance mitigation

We present a machine learning (ML) framework for predicting the structural dimensionality of hybrid metal halides (HMHs), including organic-inorganic perovskites, using a combination of chemically-informed feature engineering and advanced class-imbalance handling techniques. This study is motivated by the small and highly imbalanced nature of experimentally available HMH datasets, which limits the applicability and reliability of conventional ML approaches. The dataset, consisting of 494 HMH structures, is highly imbalanced across dimensionality classes (0D, 1D, 2D, 3D), posing significant challenges to predictive modeling. To mitigate this limitation, the dataset was augmented to 1336 samples using the synthetic minority oversampling technique, enabling improved learning of underrepresented dimensionality classes while preserving chemically meaningful feature relationships. We developed interaction-based descriptors designed to capture coupled steric and polarity effects relevant to dimensionality prediction, which are not readily captured by standard single-parameter or composition-only descriptors. These descriptors are integrated into a multi-stage workflow combining feature selection, ensemble stacking, and performance optimization. Our approach significantly improves F1-scores for underrepresented classes, achieving robust cross-validation performance across all dimensionalities. This work demonstrates a generalizable strategy for extracting reliable and interpretable structure–dimensionality relationships from limited experimental data, enabling pre-synthesis screening of organic cations and providing a practical blueprint for small-data ML in hybrid materials systems.

36 MATERIALS SCIENCE↗

Hybrid Dynamic Modeling of Smart Inverter

This letter proposes a novel hybrid method for assessing grid-connected three-phase converter interfaced resources (CIR) dynamics with the IEEE standard 1547-2018 grid support functions (GSFs), which blends physics and data-driven techniques. First, the letter derives an analytical model of a CIR to represent the internal physics and data-driven model (DDM) using a system identification algorithm to represent the rest of the dynamics, including the GSF. The derived hybrid model combines the analytical model of CIR and DDM, which balances accuracy and flexibility and is compared with the detailed switched model. Furthermore, the efficacy of the proposed approach to represent the advanced CIR dynamics is substantiated by power hardware-in-the-loop experiment data where real measurements from a commercial CIR are used to cross-validate the proposed approach. Furthermore, the results indicate that despite simple, the hybrid model accurately reproduces the dynamics of the detailed CIR model with an acceptable accuracy.

Data-driven model↗

DEPRECATED AI-Batt-OS (Autonomous Identification of Battery Life Models - Open Source) [SWR 21-17]

DEPRECATED. This repository was archived by the owner on Jun 30, 2026. It is now read-only. Open source implementation of some of the methods utilized by AI-Batt, a battery lifetime modeling and analysis toolkit provided by the National Laboratory of the Rockies (NLR). This software demonstrates the use of bi-level optimization and symbolic regression techniques to semi-autonomously identify algebraic models predicting the capacity fade of lithium-ion batteries during calendar aging. Modeling the degradation of batteries is a complex task, due to the difficulty in separating the time-dependent and time-independent factors impacting cell level degradation, across multiple data series with different numbers of measurements and/or data quality. Bi-level optimization enables model parameters to be optimized to either the entire data set or to individual data series, allowing statistical disambiguation of global behaviors (data series independent) and local behaviors (data series dependent). Symbolic regression is used to automatically search for optimal low-dimesional models predicting the variation of locally optimized parameters versus time-independent experimental variables from millions of possible models, resulting in a more accurate and repeatable model identification process than is possible by a manual search. The provided tools also implement cross-validation and bootstrap resampling schemes, empowering statistical model comparison/selection and quantification of model uncertainties. An example script replicates the results from the manuscript "Challenging Practices of Algebraic Battery Life Models through Statistical Validation and Model Identification via Machine-Learning", submitted to ECS. All code is written in MATLAB. Requires the Statistics and Machine Learning Toolbox. Contact Dr. Paul Gasper at Paul.Gasper@nlr.gov for any questions.

Gasper, Paul [National Renewable Energy Lab. (NREL↗