Search NASASearch

SEARCH · Search NASA

Results for “principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Enhancing physicochemical, bioactive, and nutritional properties of sweet potatoes: Ultrasonic contact drying with slot jet nozzles compared to hot-air drying and freeze drying

Sweet potatoes are a rich source of nutrients and bioactive compounds, but their quality can be impacted by the drying process. This study investigates the impact of slot jet reattachment (SJR) nozzle and ultrasound (US) combined drying (SJR + US) on sweet potato quality, compared to freeze-drying (FD), SJR drying, and hot air drying (HAD). SJR + US drying at 50 °C closely resembled FD in enhancing quality attributes and outperformed HAD and SJR in key areas such as rehydration, shrinkage ratios, and nutritional composition. Notably, SJR + US at 50 °C produced the highest total starch (36.84 g/100 g), total dietary fiber (8.48 g/100 g), total phenolic content (158.19 mg GAE/100 g), total flavonoid content (119.08 mg QE/g), DPPH antioxidant activity (6.44 μmol TE/g), β-carotene (31.98 mg/100 g), and vitamin C (5.27 mg/100 g). It also exhibited higher glass transition temperatures (Tg: 14.49 °C), indicating better stability at room temperature. The hardness values for SJR + US samples were similar to FD, while HAD samples had the highest hardness. SJR + US at 50 °C resulted in the lowest total color changes (ΔE), indicating minimal impact on appearance. Additionally, FTIR analysis revealed that peaks in specific spectral regions indicated superior preservation of bioactive compounds in SJR + US samples compared to other methods, which was also confirmed by principal component analysis (PCA) and heatmap visualization. Overall, these findings suggest that SJR + US is an effective alternative to conventional drying techniques, significantly improving the quality of dried sweet potatoes.

Color

Chemistry imaging and distribution analysis of rare earth elements in coal using LIBS and LA-ICP-MS instruments

Currently, demand for rare earth elements (REEs) increased significantly. Coal is actively evaluated as potential economic sources for extraction of REEs. Here, in this work, laser-induced breakdown spectroscopy (LIBS) was evaluated for rapid estimation of REEs content and their distribution in the natural coal samples. The results were compared with similar laser ablation–inductively coupled plasma–mass spectrometry (LA-ICP-MS) measurements. Thirteen coal samples (nine standard samples and five natural samples) were used in this study. Powder samples were pressed into pellets while coal chunks were directly ablated for data recording. Pellets of the powder standard samples were used to optimize the data acquisition system and then data recorded with this optimized system was used to identify the proper data acquisition and analysis models. After establishing the proper data acquisition system and analysis model using the standard samples, natural coal samples in powder form and their chunks were utilized to record LIBS and LA-ICP-MS spectra. Multivariate calibration models were developed using four of the natural samples, which were evaluated by predicting the REE content in the fifth sample. Principal component analysis was performed on the LIBS data obtained from the natural samples and it classified all the samples with high accuracy. Two-dimensional (2D) elemental mapping on coal chunk samples was also performed using both LIBS and LA-ICP-MS to study the distribution of REEs in the samples. The resulting elemental images and their correlations can be used to infer mineral distributions.

01 COAL, LIGNITE, AND PEAT

Transient uncertainty quantification and Global Sensitivity Analysis of the open-source Molten Chloride Reactor Experiment (MCRE) using GP-PCA surrogate models

Uncertainties in the thermophysical properties of molten salts impact both the steady-state and transient behavior of Molten Salt Reactors (MSRs). In this work, we aim to quantify the influence of such uncertainties on the transient operation of the Molten Chloride Reactor Experiment (MCRE), utilizing the open-source specifications provided for this reactor. Seven representative transient scenarios are considered. For each scenario, we evaluate the impact of thermophysical property uncertainties on four key multiphysics model output variables of interest (VoIs): maximum power density, maximum fuel temperature, maximum reflector temperature, and average fuel velocity magnitude. In addition, we perform a Global Sensitivity Analysis (GSA) by computing Sobol’ indices for the uncertain input parameters to determine their contribution to the variability of each VoI. Conducting GSA is computationally intensive due to the large number of required evaluations of the high-fidelity multiphysics model. To mitigate this cost, we develop a surrogate modeling framework that combines Gaussian Process (GP) regression with Principal Component Analysis (PCA), enabling efficient sample generation for the GSA. Our results show that for energy-related VoIs, thermal conductivity is the dominant contributor to uncertainty. In contrast, for flow-related VoIs, density and dynamic viscosity are the primary sources of uncertainty. The specific heat of the fuel salt was found to play a secondary role in the transient analyses.

42 - ENGINEERING

Data and Code for Understanding Generative AI Content with Embedding Models

This repository contains code for the experiments in the paper "Understanding Generative AI Content with Embedding Models". Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully hand-crafting data representations based on domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models -- which are trained to be useful across many contexts -- we demonstrate that simple and well-studied dimensionality-reduction techniques such as Principal Component Analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence (AI).

Vargas, Max [Pacific Northwest National Laboratory

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]

Novel Cell-Type-Specific Drought-Responsive Proteins in Root Tips of Field-Grown Perennial Switchgrass

The root-tip region of plants, including the root cap, forms the most basal terminal of the root and exhibits a high degree of cellular complexity in terms of morphology, cytological function, and interaction with environmental cues in the soil. Cells in this region follow a developmental trajectory, transitioning from stem cells to meristematic cells, and ultimately to fully differentiated cell types. However, our understanding of root-tip cell-type specific proteomic responses to abiotic stresses, such as drought, particularly under field conditions, remains limited. This study aimed to identify spatially resolved, cell type-specific proteomes in switchgrass (Panicum virgatum) root tips under drought stress. Root tips were collected from seven-year-old, field-grown switchgrass ‘Alamo’ plants excavated under both well-watered and long-term drought conditions. Cell type-specific proteins were identified using laser capture microdissection (LCM) coupled with nanoPOTS (Nanodroplet Processing in One Pot for Trace Samples) and nano-LC-MS proteomics analysis. Five distinct cell types were targeted: (1) cells in the quiescent center and stem cell niche (QuC), (2) protodermal epidermal cells (PEC) in the meristematic zone, (3) epidermal cells in the transition and elongation zones above the root cap (Epi), (4) peripheral root cap cells (PRC), forming 2–3 layers below the PEC and 1–2 layers above the root border cells, and (5) columella root cap cells (Col) comprising of the columella initials and a single underlying layer of cells undergoing active growth. Principal component analysis (PCA) revealed clear separation among the five targeted cell types, confirming distinct proteomic profiles. Proteins predominantly enriched in each cell type were linked to distinct cellular functions, with QuC cells showing involvement in chromosomal behavior, DNA replication, and mitosis—key processes for stem cell niche regulation. Drought stress resulted in alterations of proteostasis, as evidenced by significant decreases in ribosomal proteins and increases in protein synthesis inhibitors. Moreover, drought stress induced unique cell-type–specific proteins involved in phytohormone biosynthesis and signaling pathways, including auxin, cytokinin, and jasmonic acid. In particular, QuC cells were more highly enriched in proteins associated with DNA repair and mitotic processes. Metabolic pathways related to amino acids, carbohydrates, and lipids were differentially affected in a cell-type–dependent manner, whereas general stress-responsive proteins exhibited consistent changes across all five cell types. Overall, this study provides unique spatially resolved, cell-type-specific proteomic profiles in root tips, representing a significant advancement in our understanding of the cellular mechanisms underlying plant responses to drought stress in natural field conditions.

perennial grass

TOFHunter—unlocking rapid untargeted screening of inductively coupled plasma–time-of-flight–mass spectrometry data

This study provides an overview of a newly developed open source program written in Python, TOFHunter, which permits the rapid and untargeted screening of inductively coupled plasma (ICP)-time-of-flight (TOF)-mass spectrometry (MS) datasets. ICP-TOF-MS is an analytical tool capable of providing quasi simultaneous detection of all nuclides from Li to Pu. This capability has triggered an increase in studies investigating single-particle analysis in which the TOF-MS provides correlated elemental/isotopic signatures on a particle basis in time. Similarly, laser ablation mapping has seen rapid growth owing to ICP-TOF-MS's capacity to handle fast washout times (<10 ms) while providing a broad nuclide coverage. The caveat to this broad mass coverage and high time resolution comes in the form of large, overwhelming datasets. With datasets typically on the scale of gigabytes, it is easy for a user to only focus on very targeted analytes; however, this focus diminishes the opportunity offered by the TOF-MS detector. TOFHunter applies chemometric methods, principal component analysis (PCA), and interesting features finder (IFF) on ICP-TOF-MS data, allowing for investigation of correlations, major and minor variance sources, and sample screening. The unique spectra identified by the (IFF) are used to generate a list of mass peaks, which are then matched with both nuclides and potential interferences before being exported for the user to investigate. Several case studies are discussed herein, demonstrating TOFHunter's ability to screen aqueous injections, single-particle/single-cell analysis, and probe laser ablation mapping files for unique regions of interest.

47 OTHER INSTRUMENTATION

Revealing Phase Heterogeneity in Vertically Aligned Nanocomposites via Plan-View Electron Energy Loss Spectroscopy

Hydrogen utilization in clean energy technologies is challenged by limited storage and transport within materials, owing to the complex hydrogen kinetics at interfaces [1]. Understanding these interfacial mechanisms at the nanoscale is crucial for developing improved materials for hydrogen applications, particularly proton-conducting fuel cells (PCFCs). Vertically aligned nanocomposites (VANs) grown by pulsed laser deposition (PLD) offer a unique platform for investigating the interfacial effects on hydrogen transport due to their well-defined interfaces parallel to the direction of charge transport [2-4]. To investigate hydrogen transport, the two phases within the VANs were chosen as BaZr 0.9 Y 0.1 O 3-x (BZY), a known proton conductor, and Pr 0.1 Ce 0.9 O 2-x (PCO), a mixed ionic-electronic conductor [5]. This PCO-BZY VANs architecture allows the investigation of how the interface between a proton conductor and a mixed conductor influences hydrogen transport. However, because of the small size of hydrogen, it is difficult to discern the nature of its interactions with interfaces from bulk measurements at the macroscale, thus necessitating nanoscale measurements [6]. Electron energy loss spectroscopy (EELS) allows for nanometer-resolution probing of the local atomic structure and chemistry at the BZY/PCO interface. In this study, plan-view analysis of PCO-BZY VANs films was employed to characterize the structure and phase distribution of the VANs and investigate the interface between the nanostructures. The films were imaged using scanning electron microscopy (SEM) in the Hitachi S-4800 SEM, collecting secondary electron images using mixed upper and lower detectors. Then, plan-view transmission electron microscopy (TEM) and scanning transmission electron microscopy (STEM) EELS were employed using a JEOL ARM300 microscope operated at 300kV with a Gatan K3 GIF Continuum detector to study the distribution of the BZY and PCO phases through the film. As a result, spectrum images were acquired at a dispersion of 0.18eV per channel and denoised afterward by principal component analysis (PCA) method.

Griffin, Elizabeth [Northwestern University, Evans

Machine Learning-Driven Reliability Estimation of PV Inverters Considering Alert-Ambient Variability

Weather-induced spatio-temporal degradation limits outdoor PV inverter lifetime and reliability, necessitating advanced data analysis. This study employs a top-down, data-driven approach utilizing multiple machine learning (ML) algorithms to estimate inverter reliability in a 1.4 MW PV power plant, considering factors such as irradiance, humidity, temperature, time of day, and weather conditions. An extensive alert dataset from 17 identical inverters, including alert types, propagation, and frequency, reveals significant correlations with environmental factors and inverter output power, enabling the construction of a performance reliability model. Dual-stage supervised-ML models are evaluated for accuracy, with the ‘classification-regression’ model by an artificial neural network (ANN) tested on the averaged “Alert-Ambient” dataset, which is outperformed by ‘clustering-regression’ models using random forest (RF) and K-Nearest Neighbors (KNN) on individual inverter datasets. K-means clustering applies principal component analysis to reduce dimensions, achieving improved accuracy beyond the 80% achieved by ANN on the averaged dataset. Second-stage regression estimates inverter reliability with a mean square error of 0.0195 on the averaged dataset and as low as 0.002 on individual inverter datasets using RF. Furthermore, these findings highlight the method's suitability for estimating PV inverter output reliability under ambient conditions, essential for digital twin development and related applications.

14 SOLAR ENERGY

New Measurements of the Lyα Forest Continuum and Effective Optical Depth with LyCAN and DESI Y1 Data

Abstract We present the Ly α Continuum Analysis Network (LyCAN), a convolutional neural network that predicts the unabsorbed quasar continuum within the rest-frame wavelength range of 1040–1600 Å based on the red side of the Ly α emission line (1216–1600 Å). We developed synthetic spectra based on a Gaussian mixture model representation of nonnegative matrix factorization (NMF) coefficients. These coefficients were derived from high-resolution, low-redshift ( z < 0.2) Hubble Space Telescope/Cosmic Origins Spectrograph (COS) quasar spectra. We supplemented this COS-based synthetic sample with an equal number of DESI Year 5 mock spectra. LyCAN performs extremely well on testing sets, achieving a median error in the forest region of 1.5% on the DESI mock sample, 2.0% on the COS-based synthetic sample, and 4.1% on the original COS spectra. LyCAN outperforms principal component analysis (PCA) and NMF-based prediction methods using the same training set by 40% or more. We predict the intrinsic continua of 83,635 DESI Year 1 spectra in the redshift range of 2.1 ≤ z ≤ 4.2 and perform an absolute measurement of the evolution of the effective optical depth. This is the largest sample employed to measure the optical depth evolution to date. We fit a power law of the form τ ( z ) = τ 0 ( 1 + z ) γ to our measurements and find τ 0 = (2.46 ± 0.14) × 10 −3 and γ = 3.62 ± 0.04. Our results show particular agreement with high-resolution, ground-based observations around z = 2, indicating that LyCAN is able to predict the quasar continuum in the forest region with only spectral information outside the forest.

79 ASTRONOMY AND ASTROPHYSICS

Exploring Geothermal Potential of Great Basin Sub-Regions: Preprint

The INnovative Geothermal Exploration through Novel Investigations Of Undiscovered Systems (INGENIOUS) project aims to discover new, economically viable hidden geothermal systems in the Great Basin region by building on previous work in play fairway analysis and machine learning. A key objective of this project is to develop an exploration workflow to reduce geothermal exploration risks for hidden geothermal systems. A single preliminary play fairway workflow was developed from the assessment of the regional INGENIOUS geological, geophysical, and geochemical datasets. This workflow provided new preliminary predictive geothermal fairway maps for the INGENIOUS study area, which encompasses most of Nevada, western Utah, southern Idaho, southeastern Oregon, and easternmost California. However, a recent study (incorporating machine learning techniques) of a portion of Nevada identified four geologic domains and determined that the relative importance of individual datasets or features as indicators of geothermal potential may differ across these domains. The INGENIOUS study area includes a much larger and more geologically diverse region; therefore, additional geologic domains or sub-regions are expected. To assess the sub-regions in the INGENIOUS study area, principal component analysis and k-means clustering were applied. Preliminary results indicate that the INGENIOUS regional data cluster into groups that relate to different geologic domains in the Great Basin region. These include domains such as the Walker Lane, extensional western Great Basin region, broad lower strain region in the eastern Great Basin of western Utah and eastern Nevada, Quaternary volcanic fields, and the area adjacent to the Snake River Plain. These clusters are assessed to determine the key geologic drivers of the identified clusters. Understanding this variability can provide key insights for the exploration and characterization of hidden geothermal systems in the Great Basin region and could indicate the need to develop multiple geothermal conceptual models and play fairway workflows for the INGENIOUS study area.

GEOTHERMAL ENERGY

Exploring Geothermal Potential of Great Basin Sub-Regions

The INnovative Geothermal Exploration through Novel Investigations Of Undiscovered Systems (INGENIOUS) project aims to discover new, economically viable hidden geothermal systems in the Great Basin region by building on previous work in play fairway analysis and machine learning. A key objective of this project is to develop an exploration workflow to reduce geothermal exploration risks for hidden geothermal systems. A single preliminary play fairway workflow was developed from the assessment of the regional INGENIOUS geological, geophysical, and geochemical datasets. This workflow provided new preliminary predictive geothermal fairway maps for the INGENIOUS study area, which encompasses most of Nevada, western Utah, southern Idaho, southeastern Oregon, and easternmost California. However, a recent study (incorporating machine learning techniques) of a portion of Nevada identified four geologic domains and determined that the relative importance of individual datasets or features as indicators of geothermal potential may differ across these domains. The INGENIOUS study area includes a much larger and more geologically diverse region; therefore, additional geologic domains or sub-regions are expected. To assess the sub-regions in the INGENIOUS study area, principal component analysis and k-means clustering were applied. Preliminary results indicate that the INGENIOUS regional data cluster into groups that relate to different geologic domains in the Great Basin region. These include domains such as the Walker Lane, extensional western Great Basin region, broad lower strain region in the eastern Great Basin of western Utah and eastern Nevada, Quaternary volcanic fields, and the area adjacent to the Snake River Plain. These clusters are assessed to determine the key geologic drivers of the identified clusters. Understanding this variability can provide key insights for the exploration and characterization of hidden geothermal systems in the Great Basin region and could indicate the need to develop multiple geothermal conceptual models and play fairway workflows for the INGENIOUS study area.

exploration

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]

Hidden Features: How Subsurface and Landscape Heterogeneity Govern Hydrologic Connectivity and Stream Chemistry in a Montane Watershed

ABSTRACT Hydrologic connectivity is defined as the connection among stores of water within a watershed and controls the flux of water and solutes from the subsurface to the stream. Hydrologic connectivity is difficult to quantify because it is goverened by heterogeniety in subsurface storage and permeability and responds to seasonal changes in precipitation inputs and subsurface moisture conditions. How interannual climate variability impacts hydrologic connectivity, and thus stream flow generation and chemistry, remains unclear. Using a rare, four‐year synoptic stream chemistry dataset, we evaluated shifts in stream chemistry and stream flow source of Coal Creek, a montane, headwater tributary of the Upper Colorado River. We leveraged compositional principal component analysis and end‐member mixing to evaluate how seasonal and interannual variation in subsurface moisture conditions impacts stream chemistry. Overall, three main findings emerged from this work. First, three geochemically distinct end members were identified that constrained stream flow chemistry: reach inflows, and quick and slow flow groundwater contributions. Reach inflows were impacted by historic base and precious metal mine inputs. Bedrock fractures facilitated much of the transport of quick flow groundwater and higher‐storage subsurface features (e.g., alluvial fans) facilitated the transport of slow flow groundwater. Second, the contributions of different end members to the stream changed over the summer. In early summer, stream flow was composed of all three end members, while in late summer, it was composed predominantly of reach inflows and slow flow groundwater. Finally, we observed minimal differences in proportional composition in stream chemistry across all four years, indicating seasonal variability in subsurface moisture and spatial heterogeneity in landscape and geologic features had a greater influence than interannual climate fluctuation on hydrologic connectivity and stream water chemistry. These findings indicate that mechanisms controlling solute transport (e.g., hydrologic connectivity and flow path activation) may be resilient (i.e., able to rebound after perturbations) to predicted increases in climate variability. By establishing a framework for assessing compositional stream chemistry across variable hydrologic and subsurface moisture conditions, our study offers a method to evaluate watershed biogeochemical resilience to variations in hydrometeorological conditions.

Johnson, Keira [College of Earth, Ocean, and Atmos

Bayesian Calibration of Stochastic Agent Based Model via Random Forest

Agent-based models (ABM) provide an excellent framework for modeling outbreaks and interventions in epidemiology by explicitly accounting for diverse individual interactions and environments. However, these models are usually stochastic and highly parametrized, requiring precise calibration for predictive performance. When considering realistic numbers of agents and properly accounting for stochasticity, this high-dimensional calibration can be computationally prohibitive. This paper presents a random forest-based surrogate modeling technique to accelerate the evaluation of ABMs and demonstrates its use to calibrate an epidemiological ABM named CityCOVID via Markov chain Monte Carlo (MCMC). The technique is first outlined in the context of CityCOVID's quantities of interest, namely hospitalizations and deaths, by exploring dimensionality reduction via temporal decomposition with principal component analysis (PCA) and via sensitivity analysis. The calibration problem is then presented, and samples are generated to best match COVID-19 hospitalization and death numbers in Chicago from March to June in 2020. Further, these results are compared with previous approximate Bayesian calibration (IMABC) results, and their predictive performance is analyzed, showing improved performance with a reduction in computation.

60 APPLIED LIFE SCIENCES

Robust Spectral Anomaly Detection in EELS Spectral Images via 3D Convolutional Variational Autoencoders

Abstract A 3D Convolutional Variational Autoencoder (3D‐CVAE) is introduced for automated anomaly detection in electron energy‐loss spectroscopy spectrum imaging (EELS‐SI) data. This approach leverages the full 3D structure of EELS‐SI data to detect subtle spectral anomalies while preserving both spatial and spectral correlations across the datacube. By employing cross‐entropy loss and training on bulk spectra, the model learns to reconstruct bulk features characteristic of the defect‐free material. In exploring methods for anomaly detection, both the 3D‐CVAE approach and principal component analysis (PCA) are evaluated, testing their performance using FeL‐edge ΔEpeak shifts designed to simulate material defects. These results show that 3D‐CVAE achieves superior anomaly detection and maintains consistent performance across various shift magnitudes. The method demonstrates clear bimodal separation between bulk and anomalous spectra, enabling reliable classification. Further analysis verifies that lower‐dimensional representations are robust to anomalies in the data. While performance advantages over PCA diminish with decreasing anomaly concentration, our method maintains high reconstruction quality even in challenging, noise‐dominated spectral regions. This approach provides a robust framework for unsupervised automated detection of spectral anomalies in EELS‐SI data, particularly valuable for analyzing complex material systems.

Chemistry

Queue wait time prediction in high performance computing (HPC) systems

High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.

Okafor, Nwamaka

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy