Search NASA⌕ Search

SEARCH · Search NASA

Results for “well performance ranking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Predicting future well performance for environmental remediation design using deep learning

Here in this study, we developed a deep learning (DL) framework with a multi-channel three-dimensional convolutional neural network (MC3D-CNN) to predict well performance and thereby assist future environmental remediation design. Such prediction of extraction well performance at designated locations is critical for configuring pump-and-treat (P&T) well network design and operation, setting reasonable target closure dates for overall remedying, and estimating remedy costs. The framework is developed with operational and monitoring data routinely collected during P&T remedy operations, including well extraction and injection rates as well as in situ contaminant concentrations. Traditionally, the collected data were rarely used for purposes other than assessing past well performance and the accuracy of the conceptual site model. However, recent advances in data-driven computational approaches enable better use of the large datasets to inform future well performance, enhance site characterization, and improve remediation planning. In this study, we established a DL framework to integrate transient three-dimensional contaminant plumes and multiple aquifer properties (e.g., hydraulic conductivity and hydrostratigraphic maps) to identify characteristic patterns controlling and representing extraction well mass recovery, aiming at providing future mass recovery estimates for existing wells and candidate wells at any proposed locations. We evaluated our framework by using a realistic synthetic dataset generated from a well-calibrated flow and transport model used in the 200 West Area of the U.S. Department of Energy’s Hanford Site in southeastern Washington state. The multi-channel feature in our framework allows integration of various types and temporal densities of training datasets for DL model development. Overall, we found that the trained DL model achieved an accuracy of over 90% in ranking extraction well performance in validation datasets, and over 80% in predicting high-performance-ranking well locations. This data-informed approach provides a flexible tool to support adaptive site management, streamline decision-making, and potentially reduce remediation time and costs. Our DL framework can be used as a filtering tool to improve the current P&T network optimization design by reducing the number of candidate well locations.

54 ENVIRONMENTAL SCIENCES↗

Field studies of PERC and Al-BSF PV module performance loss using power and I-V timeseries

We have studied the degradation of both full-sized modules and minimodules with PERC and Al-BSF cell variations in fields while considering packaging strategies. We demonstrate the implementations of data-driven tools to analyze large numbers of modules and volumes of timeseries data to obtain the performance loss and degradation pathways. This data analysis pipeline enables quantitative comparison and ranking of module variations, as well as mapping and deeper understanding of degradation mechanisms. The best performing module is a half-cell PERC, which shows a performance loss rate ( PLR ) of −0.27 ± 0.12% per annum (%/ a ) after initial losses have stabilized. Minimodule studies showed inconsistent performance rankings due to significant power loss contributions via series resistance, however, recombination losses remained stable. Overall, PERC cell variations outperform or are not distinguishable from Al-BSF cell variations.

Curran, Alan J.↗

A Workflow for Characterizing Legacy Wells as Potential Leakage Pathways for Integration to NRAP-Open-IAM

Carbon capture and storage is a crucial component of climate change mitigation strategies, involving the capture of carbon dioxide (CO2) from point sources and its injection into permeable subsurface formation. Many suitable CO2 storage sites coincide with legacy wells since the conditions that kept hydrocarbons in-situ for thousands of years are also ideal for storage of carbon dioxide. To protect underground sources of drinking water (USDW) during greenhouse gas injection, the Environmental Protection Agency (EPA) mandates area of review evaluations. These evaluations ensure that drinking water sources would not be contaminated by injected fluids. They include identification of legacy wellbores, integrity assessments, and implementing any necessary corrective action. Previous assessment approaches of legacy wells include high-level scoring of regional data and well construction and abandonment evaluation. This work describes a novel methodology that evaluates well construction and abandonment, ranks them based on complexity, and performs a risk assessment with NRAP-Open-IAM. A workflow of the methodology is presented, highlighting its capabilities and limitations.

Wise, Jarrett↗

Developing a Prototype Methodology to Rank CO2-EOR Wells and Assess Their Reuse Potential for Geologic Carbon Storage

This paper presents a prototype methodology to assess the possible transition of Class II carbon dioxide-enhanced oil recovery (CO2-EOR) wells to Class VI wells. The focus is on wellbore construction materials—casing, cement, tubing, and the packer—and includes comprehensive workflows to evaluate these materials, with primary emphasis on compliance with Environmental Protection Agency (EPA) Class VI well construction and conversion guidelines. These workflows systematically assess material properties and performance criteria to ensure regulatory compliance and optimize long-term wellbore integrity and functionality. Utilizing Python scripts and JavaScript Object Notation (JSON) representations, the study automates checks on digitized Texas Railroad Commission (TRRC) data to rank wells based on workflow criteria. By emphasizing critical factors such as casing integrity, cementing techniques, tubing compatibility, and packer selection, the methodology helps well owners and operators prioritize wells for potential reuse as CO2 injection wells. Given limitations in digitized data, manual user verification is required in some sections. Future improvements include integrating non-digitized data through web scraping and machine learning techniques. This research serves as a practical guide for stakeholders, supporting environmental compliance and sustainable well operations.

geologic carbon sequestration↗

Laboratory-Scale Coal-Derived Graphene Process (Final Report)

The Energy & Environmental Research Center (EERC) conducted a laboratory-scale coal-derived graphene (CDG) project focused on developing a technological process for making graphene from four U.S. domestic coal or coal wastes, including lignite from North Dakota, subbituminous coal from Wyoming, bituminous coal from Utah, and anthracite from Pennsylvania. The project was divided into two performance or budget periods (BPs), with BP1 comprising the up-front laboratory experiments to make graphene materials from coal beginning on May 1, 2020, to April 30, 2022. BP2 was conducted from May 1, 2022, to April 30, 2023, and was focused on analyzing the CDG process economic feasibility and the technical gaps for technological scale-up and commercialization. During this project, a few different coal-derived high-value products have been demonstrated, including graphite, graphene oxide (GO), reduced graphene oxide (rGO), and graphene quantum dots (GQDs). A new graphite microstructure was discovered and named “croissant graphite” because of the exterior morphological and textural resemblance to croissant food items sold in commercial groceries stores. The new graphite structure and the associated preparation from coal or coal waste feedstocks has been the subject of a U.S. patent application. The systematic experimental processes involving coal cleaning, upgrading, and conversion to high-value carbon products culminated into a developed upgraded coal-to-products (UCP) technology that is being pursued for potential fast-track commercialization, if funding is available. It is envisioned that commercialization of the UCP technology would increase consumption of U.S. domestic coals or coal wastes to make environmentally sustainable high-value products for the electronics industry, high-energy-storage applications, and clean energy technologies such as electric vehicle (EV) lithium-ion batteries (LIBs), for which graphite has become a critical mineral commodity. Croissant graphite microstructures, when observed by field emission scanning electron microscopy (FESEM), display wavy surface morphology and often grow from a base that is made of graphitized particles with honeycomb-like layers, which are believed to be graphene layers. While more studies are needed to fully ascertain the mechanisms of the croissant graphite microstructure formation, it is postulated that their growth may begin from curling of the graphene sheets into ribbon-like structures, and continuous growth and densification of the ribbon-like structures forms croissant microstructures. Additional studies are ongoing to evaluate the electrochemical performance of croissant graphite for LIB applications and to determine the experimental conditions necessary to tune on/off croissant formation so that it can be either optimized or suppressed depending on performance evaluation results. In addition to the discovery of croissant graphite, the graphitization process from the four coal ranks in general was successful. X-ray diffraction (XRD) analysis showed that the degree of graphitization (DoG) ranged from 12% to 80% in an early sample set, and further optimization on lignite coal produces a DoG of about 92%, which was spectacular to see as lignite is the lowest-rank coal. Thus, it is expected that the graphitization performance for higher-rank coals will be similar or better when optimized as well. The coal-derived graphite was used to make GO and rGO. Analytical characterization, e.g., by methods such as Raman spectroscopy, XRD, Fourier transform infrared (FTIR) spectroscopy and FESEM, showed that the sequence of converting the coal to graphite, exfoliating it to GO, and then chemically reducing the GO to rGO was successful. Although coal naturally contains aromatic compounds and some relatively small-sized condensed aromatic units, it does not contain graphene sheets. In the UCP process, the aromatic domains in the coals, particularly low-rank coals, are concentrated and condensed further into graphene sheets, which are ordered into a 3D stack during graphitization. The synthesized graphite is then unpacked by methods such as exfoliation to various graphene products. GQDs were synthesized from all four coal types, and their optical properties were demonstrated to be tunable by the coal precursor preprocessing treatments. In all four coal types, enhanced optical properties were observed for the produced GDQs with incremental improvements made to the coal precursors. GQDs produced from raw coal samples displayed lower ultraviolet–visible (UV–Vis) spectroscopy absorbance intensity compared to those obtained from cleaned and upgraded coal residues. The photoluminescence (PL) intensities also varied with pretreatment conditions and with the concentration of GQDs in aqueous solutions. GQDs obtained from anthracite show longer emission wavelengths and can be excited by visible light as opposed to GQDs derived from the other coal ranks. UV fluorescence 3D maps and spectra revealed that the emission wavelength at which the GQDs solutions display the highest intensity was slightly redshifted based on the coal precursor pretreatments. In low-rank coal (lignite and subbituminous) samples, two clusters were observed in the maps for GQDs, which may suggest that there are potentially two types of fluorophores in solution or two main size populations. The ability to tune the properties of GQDs based on processing methods can be exploited to make GQDs for various optical display or optoelectronics applications. The results also highlight the importance of removing coal-borne impurities to improve the quality of the coal precursor for preparation of graphene products. Coal and/or coal wastes preprocessing methods were developed and applied to clean and upgrade the coal precursors prior to graphitization and subsequent conversion to graphene products. The preprocessing methods involve high specific-gravity separations, mineral acid cleaning (no hydrofluoric acid), and subsequent upgrading by reducing the coal-borne heteroatom (nitrogen, sulfur, and oxygen) content using proprietary chemical agents. Analytical characterization revealed that the preprocessing steps were successful, with ash reductions that range from 38% to 80% and residual ash content that was below the 5 wt% initial target. Based on proximate and ultimate analysis, the heteroatom reduction reactions produced upgraded coal residues with the oxygen content reduced by 8% to 24%, with additional reductions in the nitrogen and sulfur contents. An initial assessment of the waste streams from the UCP process shows very small to negligible environmental impact due to CO 2 , NO x , and SO x because most process steps are performed under inert atmosphere with argon. Consequently, reactive oxygen environments that tend to create these species are avoided. The inorganic and potentially hazardous species are released into aqueous waste streams that are easy to handle for proper disposal. The liquid waste streams were found to contain low-level concentrations of rare-earth elements (REEs), which could be concentrated and recovered as value-added by-products. Additionally, the volatile and gaseous fractions from carbonization and heat treatment contain useful organic compounds that can also be recovered as potential valuable by-products. Thus, the UCP technology is considered an environmentally sustainable and promising emerging technology for making high-value products from coal and coal wastes, with potential additional value-added by-products. Analysis of potential markets for the coal-derived carbon products shows a strong demand in both niche market sectors and across a wide variety of other industrial sectors. Graphite is currently considered a critical mineral commodity that has a large and growing demand in the LIB industry for EV applications. Based on data from Fortune Business Insights (2022) and Marketwatch (2023) reports, the average global graphite market is projected to reach about 33 billion by 2028, growing at a compound annual growth rate (CAGR) of about 7%, with much of this growth expected to be in the LIB industry. GO and rGO have strong market potentials in various application areas, such as coatings for anticorrosion, anti-icing, and antimicrobial protection, thermal barriers, wear resistance, sensors, additive manufacturing such as 3D inks, and others. GQDs are the emerging key player in the bioimaging, photovoltaics, and light-emitting diodes (LEDs) applications, with the potential to replace traditional semiconductor quantum dots (SQDs), which are based on metallic systems that are more toxic and more expensive. Biomedical applications of GQDs are becoming more attractive because of low to no toxicity and extremely low cost compared to SQDs. The major challenges for scale-up and commercialization of coal-derived carbon products such as graphene vary from the inherent attributes of graphene itself to reluctance to accept graphene in new manufacturing processes because of the uncertainty of the unknown. The 2D nature of graphene materials with a thickness of one atom presents significant challenges to proper handling/processing, and process scale-up becomes difficult because it requires high-end, expensive equipment, even for routine handling and analysis for quality assurance and control. Pristine graphene can also be extremely difficult to work into other matrices, thus hindering downstream processibility, especially at large scale. Currently, the cost of graphene and graphene products is still high and presents an economic risk that tends to slow down investment in scaling up emerging technologies. The lack of a standard for graphene materials for quality assurance and quality control poses a great challenge not only for the markets but also for commercialization efforts. A first-look economic feasibility analysis of the UCP technology provided valuable information that suggests the UCP process would be feasible, especially when it is scaled to a pilot scale and could be more competitive at the full scale. Graphitization was found to be the most energy-consuming and most capital-intensive step in the overall process. In small laboratory- and bench-scale experiments, labor is a significant contributor to the total process costs. Although these energy, capital, and labor constraints contribute to a higher selling price for the product, a preliminary economic model suggests that the process would be feasible at large scale when the process is fully integrated, optimized, and automated.

01 COAL, LIGNITE, AND PEAT↗

Why is the winner the best?

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multi- center study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), im- age preprocessing (97%), data curation (79%), and post- processing (66%). The “typical” lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work.

Eisenmann, Matthias↗

Comprehensive framework for assessing and optimizing existing research networks

Conservation, monitoring, and research networks, or collections of ecological research sites unified under a common mission of data collection or a research mission, are essential infrastructure for understanding large landscapes. However, most networks developed opportunistically over decades rather than through systematic design, creating potential limitations in the ability to address conservation challenges across entire regions. We developed a framework to evaluate how well an existing research network represents the environmental conditions its members study and devised an approach to rank sites of priority for strategic expansion. Our approach measures performance through environmental representativeness, geographic coverage, and adequacy for scientific inference and thus optimizes limited monitoring resources to maximize scientific impact. We demonstrated this approach with the U.S. Department of Agriculture (USDA) Forest Service Experimental Forests and Ranges Network (EFRN), a 79‐site network across the United States that grew opportunistically over a century. At the national scale, the network effectively captured high‐biomass forests important for carbon cycle research; 82% of forest biomass was in well‐represented areas. Some areas in Texas, Florida, the Rocky Mountains, and the West Coast had no relevant EFRN sites, which limits the ability to make regional inferences. A fundamental challenge for the EFRN was that sites improving regional extent coverage sometimes provided minimal national benefits, which can create conflicts between local and global priorities. Adding the highest‐ranked candidate site provided a relevant site for 17% of currently poorly represented 1‐km pixel cells nationally, but regional and national site rankings varied considerably due to nested spatial inference. This framework provides quantitative tools for strategic infrastructure decision‐making, ensures that limited monitoring resources maximize conservation impact, and can be applied broadly to address the widespread challenge of optimizing conservation and monitoring networks worldwide.

additional site↗

Machine learning and deep learning for mineralogy interpretation and CO 2 saturation estimation in geological carbon Storage: A case study in the Illinois Basin

Carbon capture and storage (CCS) is a promising approach to simultaneously maintaining energy security and reducing carbon dioxide (CO 2 ) emissions under the current energy portfolio that is dominated by fossil fuel energy. Pre-injection formation characterization and post-injection CO 2 monitoring are two critical tasks to guarantee storage efficiency in CCS. The CCS projects in the Illinois Basin, the first large-scale CO 2 injection into saline aquifers in the United States, employed conventional and the latest pulsed neutron logging (PNL) tools for mineralogy interpretation and CO 2 saturation estimation, which provide valuable references for future CCS projects. Because of the inherent fuzziness of petrophysical measurements and complex subsurface heterogeneity, interpreting well-logging data is time-consuming, and its accuracy can be user-biased. In recent years, data-driven methods have been widely used to capture the non-linear patterns between input features and interpretation results. This work applied and evaluated four commonly used machine learning (ML) models, including ridge regression (RR), random forest (RF), gradient boosting regression (GBR), support vector regression (SVR), and one deep learning (DL) model, the artificial neural network (ANN). We optimized the hyperparameters of the four ML models and the DL model using the simulated annealing algorithm and the grid search strategy, respectively. The input features of the mineralogy interpretation models were eleven conventional well-logging parameters, and the label data (i.e., ground truth) were the porosity and volumetric fractions of six minerals, including quartz, feldspar, dolomite, calcite, clay, and iron minerals. The results demonstrated that the GBR and RF models were superior in predicting volumetric fractions of minerals and porosity; label data with low coefficient of variation (CV) values tended to yield better performance. For CO 2 saturation estimation, the RF was the best-performing model, followed by SVR, ANN, GBR, and RR. Furthermore, we conducted feature importance ranking using the permutation importance algorithm and found that the formation sigma and well pressure were the most important features in this study. In conclusion, the study of CCS projects in the Illinois Basin bridges the gap between the limited knowledge and understanding of geological carbon storage and the increasing demand for reliable, cost-effective, and sustainable energy solutions.

58 GEOSCIENCES↗

Time-dependent-bases with local CUR decomposition method for accelerating turbulent combustion simulations

Here, this study presents a novel reduced-order modeling framework, Time-Dependent Bases with Local CUR decomposition (TDB-L-CUR), designed to efficiently and accurately approximate the species transport equations in reacting flow simulations. The method extends the existing TDB-CUR approach for chemically reacting flows (Jung et al. Comput. Methods Appl. Mech. Engrg. 437 (2025) 117758), which leverages matrix decomposition techniques to form a global-in-space, time-dependent low-dimensional manifold. While TDB-CUR performs well in homogeneous systems, it may be less well-suited to spatially heterogeneous systems such as turbulent flames, where higher-rank approximations are typically required. The proposed TDB-L-CUR framework introduces two methodological extensions to the baseline approach. First, it applies unsupervised clustering to partition the physical domain into distinct regions, enabling spatially localized manifold construction, thereby reducing the rank required for the reduced-order representation. Second, it incorporates a computational singular perturbation (CSP)-based scheme for identifying and penalizing fast species, allowing for spatio-temporally adaptive mitigation of chemical stiffness. The proposed framework is validated on a hierarchy of test cases, including a one-dimensional premixed flame, a two-dimensional nonpremixed ignition case with vortex interaction, and a three-dimensional turbulent premixed flame. TDB-L-CUR significantly improves accuracy over TDB-CUR while further reducing computational cost. The fully on-the-fly formulation of TDB-L-CUR (i.e., requiring no offline training or prior knowledge) makes it a robust and scalable tool for reduced-order modeling of reactive flows.

Local manifold↗

Exploring Black-box Adversarial Attacks on Low-rank Constrained Neural Networks

Low-rank compression has been shown as an effective tool to reduce parameter counts of convolutional and vision transformer architectures; however, low-rank training often reduces model robustness to adversarial perturbations. In this work, we explore the effects of low-rank training on black-box attacks, where attacked images are generated without knowledge of the low-rank parameters. We find that low-rank training is not sufficient as a black-box defense and can sometimes produce worse than expected as compared to baseline models. Influencing the spectrum of the low-rank models during training, which is known to increase model robustness against white-box attacks, improves black-box performance as well.

Schnake, Stefan [ORNL] (ORCID:0000000215183538)↗

An Alternative Ensemble Streamflow Prediction Approach Using Improved Subseasonal Precipitation Forecasts from the North America Multi-Model Ensemble Phase II

In this article, streamflow forecasting at a subseasonal time scale (10–30 days into the future) is important for various human activities. The ensemble streamflow prediction (ESP) is a widely applied technique for subseasonal streamflow forecasting. However, ESP’s reliance on the randomly resampled historical precipitation limits its predictive capability. Available dynamical subseasonal precipitation forecasts provide an alternative to the randomly resampled precipitation in ESP. Prior studies found the predictive performance of raw subseasonal precipitation forecast is limited in many regions such as the central south of the United States, which raises questions about its effectiveness in assisting streamflow forecasting. To further assess the hydrologic applicability of dynamical subseasonal precipitation forecasts, we test the subseasonal precipitation forecast from North America Multi-Model Ensemble Phase II (NMME-2) at four watersheds in the central south region of the United States. The subseasonal precipitation forecasts are postprocessed with bias correction and spatial disaggregation (BCSD) to correct bias and improve spatial resolution before replacing the randomly resampled precipitation in ESP for streamflow predictions. The performance of the resulting streamflow predictions is benchmarked with ESP. Evaluation is conducted using Kling–Gupta Efficiency (KGE), continuous ranked probability score (CRPS), probability of detection (POD), false alarm ratios (FARs), as well as reliability diagrams. Our results suggest that BCSD-corrected subseasonal precipitation forecasts lead to overall improved streamflow predictions due to added skills in winter and spring. Our results also suggest that BCSD-corrected subseasonal precipitation forecasts lead to improved predictions on the occurrence of high-percentile streamflow values above 75%. Overall, BCSD-corrected subseasonal precipitation has shown promising performance, highlighting its potential broader applications for river and flood forecasting.

54 ENVIRONMENTAL SCIENCES↗

Advancing density functional tight-binding method for large organic molecules through equivariant neural networks

Semi-empirical quantum-mechanical (QM) methods have become valuable tools for studying complex (bio)molecular systems due to their balance between computational efficiency and accuracy. A key aspect of these methods is their parameterization, which not only governs the reliability of the results but also provides an opportunity to enhance their overall performance. In our previous work [J. Phys. Chem. Lett., 2021, 11, 16], we advanced the third-order semi-empirical density functional tight-binding (DFTB3) method for computing multiple properties of small molecules by developing the machine learning (ML) potential NN rep to bridge the gap between DFTB3 electronic components and those of the hybrid DFT-PBE0 functional. To overcome the limitations of NN rep , we introduce the EquiDTB framework, which leverages physics-inspired equivariant neural networks (NN) to parameterize scalable and transferable many-body Δ TB potentials, replacing the standard pairwise DFTB repulsive potential. This advancement extends the applicability of our ML-corrected DFTB approach to larger molecules and non-covalent systems (including only C, N, O, and H atoms), going beyond the chemical space represented in the training QM datasets. The enhanced performance of EquiDTB over the standard TB methods is demonstrated by the accurate computation of the atomic forces of S66x8 molecular dimers, as well as their interaction energies. Moreover, EquiDTB can be effectively employed to explore the potential energy surfaces of large and flexible drug-like molecules—for example, to determine the minimum energy path between isomers, analyze structural transitions during dynamical simulations, compute vibrational modes, and investigate energetic rankings. The performance for single molecules slightly decreases when the DFTB electronic energy is reduced to first-order but remains superior to standard TB methods. Our work thus demonstrates that an optimal integration of an equivariant NN with QM datasets can advance the DFTB method while maintaining high efficiency, paving the way for reliable (bio)molecular simulations.

Medrano Sandonas, Leonardo [Technische Universität↗

AEflow (Autoencoder fluid flow compression network) [SWR-22-29]

As the size of turbulent flow simulations continues to grow, in situ data compression is becoming increasingly important for visualization, analysis, and restart checkpointing. For these applications, single-pass compression techniques with low computational and communication overhead are crucial. In this paper we present a deep-learning approach to in situ compression using an autoencoder architecture that is customized for three-dimensional turbulent flows and is well suited for contemporary heterogeneous computing resources. The autoencoder is compared against a recently introduced randomized single-pass singular value decomposition (SVD) for three different canonical turbulent flows: decaying homogeneous isotropic turbulence, a Taylor-Green vortex, and turbulent channel flow. Our proposed fully convolutional autoencoder architecture compresses turbulent flow snapshots by a factor of 64 with a single pass, allows for arbitrarily sized input fields, is cheaper to compute than the randomized single-pass SVD for typical simulation sizes, performs well on unseen flow configurations, and has been made publicly available. The results reported here show that the autoencoder dramatically outperforms a randomized single-pass SVD with similar compression ratio and yields comparable performance to a higher-rank decomposition with an order of magnitude less compression in regard to preserving a number of important statistical quantities such as turbulent kinetic energy, enstrophy, and Reynolds stresses.

King, Ryan↗

Random projection using random quantum circuits

The random sampling task performed by Google's Sycamore processor gave us a glimpse of the “quantum supremacy era.” This has definitely shed some light on the power of random quantum circuits in this abstract task of sampling outputs from the (pseudo)random circuits. In this paper, we explore a practical near-term use of local random quantum circuits in dimensional reduction of large low-rank data sets. We make use of the well-studied dimensionality reduction technique called the random projection method. This method has been extensively used in various applications such as image processing, logistic regression, entropy computation of low-rank matrices, etc. We prove that the matrix representations of local random quantum circuits with sufficiently shorter depths [ ∼ O ( n ) ] serve as good candidates for random projection. We demonstrate numerically that their projection abilities are not far off from the computationally expensive classical principal components analysis on MNIST and CIFAR-100 image datasets. We also benchmark the performance of quantum random projection against the commonly used classical random projection in the tasks of dimensionality reduction of image data sets and computing von Neumann entropies of large low-rank density matrices. And finally, using variational quantum singular value decomposition, we demonstrate a near-term implementation of extracting the singular vectors with dominant singular values after quantum random projecting a large low-rank matrix to lower dimensions. All such numerical experiments unequivocally demonstrate the ability of local random circuits to randomize a large Hilbert space at sufficiently shorter depths with robust retention of properties of large data sets in reduced dimensions. Published by the American Physical Society 2024

Kumaran, Keerthi (ORCID:0009000949125721)↗

Combining pairwise structural similarity and deep learning interface contact prediction to estimate protein complex model accuracy in CASP15

Abstract Estimating the accuracy of quaternary structural models of protein complexes and assemblies (EMA) is important for predicting quaternary structures and applying them to studying protein function and interaction. The pairwise similarity between structural models is proven useful for estimating the quality of protein tertiary structural models, but it has been rarely applied to predicting the quality of quaternary structural models. Moreover, the pairwise similarity approach often fails when many structural models are of low quality and similar to each other. To address the gap, we developed a hybrid method (MULTICOM_qa) combining a pairwise similarity score (PSS) and an interface contact probability score (ICPS) based on the deep learning inter‐chain contact prediction for estimating protein complex model accuracy. It blindly participated in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022 and performed very well in estimating the global structure accuracy of assembly models. The average per‐target correlation coefficient between the model quality scores predicted by MULTICOM_qa and the true quality scores of the models of CASP15 assembly targets is 0.66. The average per‐target ranking loss in using the predicted quality scores to rank the models is 0.14. It was able to select good models for most targets. Moreover, several key factors (i.e., target difficulty, model sampling difficulty, skewness of model quality, and similarity between good/bad models) for EMA are identified and analyzed. The results demonstrate that combining the multi‐model method (PSS) with the complementary single‐model method (ICPS) is a promising approach to EMA.

59 BASIC BIOLOGICAL SCIENCES↗

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Grid Optimization Competition on Synthetic and Industrial Power Systems

This paper summarizes a grid optimization (GO) competition effort in the United States to find the best solution strategies for up to interconnect-scale power system networks with around 32,000 buses. The optimization problem is a mixedinteger, non-convex non-linear problem, (MINLP) and includes discrete variables such as unit commitment and line switching, control settings (transformer taps and phase shifters with impedance correction tables), and bus shunts. The case study includes six actual industry grids as well as 16 realistic synthetic grids created by three different dataset teams. The winners are selected and ranked based on scoring criteria, which consider the solution quality (such as objective functions) within time limits. Nine winner teams are selected from 26 competitor teams. The results achieved by different teams are described and the performance of different algorithms on synthetic grids and actual industry grids are compared and analyzed.

mixed-integer non-linear programming↗