Search NASA⌕ Search

SEARCH · Search NASA

Results for “Rank”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations

In this article, we focus on the communication costs of three symmetric matrix computations: (i) multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK) (ii) adding the result of the multiplication of a matrix with the transpose of another matrix and the transpose of that result, known as a symmetric rank-2k update (SYR2K) (iii) performing matrix multiplication with a symmetric input matrix (SYMM). All three computations appear in the Level 3 Basic Linear Algebra Subroutines (BLAS) and have wide use in applications involving symmetric matrices. We establish communication lower bounds for these kernels using sequential and distributed-memory parallel computational models, and we show that our bounds are tight by presenting communication-optimal algorithms for each setting. Our lower bound proofs rely on applying a geometric inequality for symmetric computations and analytically solving constrained nonlinear optimization problems. As a result, the symmetric matrix and its corresponding computations are accessed and performed according to a triangular block partitioning scheme in the optimal algorithms.

Al Daas, Hussam [Rutherford Appleton Laboratory, D↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Direct numerical simulations for hybrid rocket boundary layers: Performance modeling and scaling

This paper presents a comprehensive performance and scaling analysis of direct numerical simulations for reacting boundary layers, focusing on slab burner configurations. Using a PETSc-based finite volume CFD framework, the study evaluates the scalability and computational cost of flow, chemistry, and radiation evaluations across 2D and 3D simulations. Polymethyl methacrylate (PMMA) is the fuel with pure O 2 as the oxidizer, modeled using a detailed chemical kinetics mechanism with 113 species and 660 reactions. A ray-tracing-based radiation solver, designed for distributed memory applications, is implemented to model radiation heat transfer. Parallel scalability is analyzed for the coupled flow, chemistry, and radiation heat transfer processes. Weak and strong scaling studies are conducted on up to 15,000 computational ranks, revealing robust performance when flow cells exceed 200 per rank. Chemistry evaluations dominate the computational cost in large 3D simulations, accounting for approximately 40% of the total runtime, while flow processes contribute around 35%, and radiation solver contributions remain below 10% due to reduced evaluation frequencies. GPU accelerated chemistry evaluation, implemented with Zero-RK, demonstrates significant promise, achieving up to a 4x speedup for workloads exceeding 30,000 cells per GPU. However, diminishing returns are observed for smaller workloads due to CPU-GPU communication overhead. This study identifies key challenges, including memory bottlenecks and the effects of domain partitioning on flow scalability, while highlighting the potential of GPU-accelerated chemistry to reduce computational costs. In conclusion, these findings provide realizable run configurations for 2D, 3D, and GPU-accelerated cases, offering insights for optimizing reactive flow solvers.

CFD Scalability↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Applications of fuzzy logic and best-worst method for tritium sensor selection

Accurate assessment of tritium as a fuel source is critical in fusion reactions, necessitating effective sensor evaluation methods. This study investigates a multi-criteria decision-making framework for selecting tritium sensors, integrating fuzzy logic to enhance decision quality. Initial attempts at applying fuzzy logic were found to be too elementary and failed to capture the complexity of multi-criteria selection; this prompted a refined approach that incorporated expert insights and advanced ranking techniques for sensor evaluation. The research used a two-stage methodology. In the first stage, important criteria and sub-criteria for sensor performance were identified and defined. These criteria were then weighted and scored using a fuzzy best-worst method, drawing upon expert opinions to ensure relevance and validity. The second stage involved interpreting information about varying sensors to rank them based on their overall criteria scores, encouraging the selection of the most suitable options. The result of the study is a proposed method for effective sensor selection in fusion reactors, which in turn will significantly improve the reliability of tritium monitoring in fusion applications.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Alternating Direction Decomposition with Strong Bounding and Convexification (ADDSBC) for Solving Security Constrained AC Unit Commitment Problems

This project aims to develop efficient and robust computational methods for solving the security-constrained unit commitment and alternating current optimal power flow problem (SC-UC-ACOPF). The SC-UC-ACOPF problem is at the center of the short-term operation of the U.S. Power Grid. It is solved every week, every day, and every 10 minutes to plan for the optimal action of electricity generation and consumption by minimizing the generation cost and maintaining power system reliability against potential disruptions of equipment failures. In mathematical terms, SC-UC-ACOPF is a challenging large-scale mixed-integer nonlinear optimization model. This means that the decisions involve both discrete variables, e.g. the turning on and off of generators and switching of transmission lines and transformers, and continuous decisions, e.g. the amount of energy generated by each generator and the power flows in the power grid. The physics of the power flow is described by nonlinear equations involving real and reactive power and bus voltages. Another key feature is the large number of contingencies, i.e. the system needs to stay reliable in face of failure of any one equipment, such as transmission lines and generators. The U.S. power grids are extremely complicated and large scale with more than 5,000 generators, 50,000 buses, and 100,000 high-voltage transmission lines, making the SC-UC-ACOPF a very large-scale computation challenge. The research developed in this project aims to solve the SC-UC-ACOPF problems in the three timescales, i.e. weekly, daily, and every 10-min. The proposed computational methods are built on a principled algorithmic approach of decomposition and penalization. More specifically, the algorithm develops spatial and temporal decomposition by exploiting the strong temporal coupling and weak spatial coupling of the UC problem and the complementary feature, i.e. weak temporal coupling and strong spatial coupling of the ACOPF problem. The algorithm also leverages recent progresses in strong convex relaxation of ACOPF. A unique feature of the proposed approach is that it generates a valid, global upper bound on the optimal maximum profit. In this way, a global optimality gap is available to measure the quality of the solution. To further speed up computation, the research team has developed a plethora of effective heuristics to strengthen the iterative penalty-based decomposition framework. For instance, a heuristic is developed to construct inner approximations of the time coupling constraints within the time decoupled problems. Contingencies are pre-screened and low-rank matrix computation is exploited to find the almost unique solution to each contingency. A novel heuristic for line switching is proposed and tested with positive impacts on instances where line switching is beneficial. Taking a systematic approach and carefully handling every detail of the problem pays off. The TIM-GO’s performance throughout the trials and the final event was stellar. TIM-GO garnered the second highest total prize money and is ranked in the top three positions across all categories of comparison.

97 MATHEMATICS AND COMPUTING↗

The Average Spectrum Norm and Near-Optimal Tensor Completion

We propose the average spectrum norm to study the minimum number of measurements required to approximate a multidimensional array (i.e., sample complexity) via low-rank tensor recovery. Our focus is on the tensor completion problem, where the aim is to estimate a multiway array using a subset of tensor entries corrupted by noise. Our average spectrum norm-based analysis provides near-optimal sample complexities, exhibiting dependence on the ambient dimensions and rank that do not suffer from exponential scaling as the order increases.

97 MATHEMATICS AND COMPUTING↗

Workflow for Developing and Operating Subsurface Hydrogen Storage Facilities in Porous Reservoirs

Long-duration (seasonal) storage of natural gas (NG), which primarily consists of methane (CH 4 ), has been practiced for more than a hundred years at underground gas storage (UGS) facilities that use depleted hydrocarbon reservoirs, saline aquifers, and salt caverns. To enable hydrogen (H 2 ) to be used as a long-duration, energy-storage medium, similar facilities are envisioned for underground H 2 storage (UHS) of either H 2 or H 2 /NG mixtures. Experience with UGS can be used to guide recommended practices for developing and operating UHS facilities in porous reservoirs. The most important factors (formation/fluid properties and engineering choices) that influence the performance of UHS reservoirs have been identified and quantified in previous studies. These factors and choices influence phenomena that determine the sweep efficiency of the stored working gas. These phenomena include viscous fingering, hysteretic capillary trapping, and gravity override of the working gas, as well as the upconing of nonproductive fluid that determine the sweep efficiency of the stored working gas. This report describes initial recommended-practices and a project-development workflow for UHS facilities that utilize porous reservoirs, based on the current state-of-knowledge about H 2 behavior in the subsurface. The workflow sequentially addresses all aspects of UHS project development, including the identification of H 2 sources and users, site ranking and down-selection, geologic and reservoir-engineering characterization, reservoir design, testing, risk management, commissioning, operations, and monitoring for a UHS facility. The goal is to enable UHS facilities to be developed in an efficient and timely manner, while carefully managing project risks. This workflow is similar to that which has been developed for UGS facilities (see Figure 1 of API, 2022), with the addition of tasks and subtasks specific to H 2 and UHS. The project-development workflow is broken down into three major stages: (1) define the H 2 use case; (2) rank, down-select, and characterize potential, candidate UHS sites; and (3) reservoir design, integrity testing, risk assessment, commissioning, operations, and monitoring for selected UHS sites. Each major stage is further broken down into tasks and subtasks, which are described at a high level. This report also provides more detailed descriptions of all tasks and subtasks that involve reservoir analysis and testing.

08 HYDROGEN↗

Economic Evaluation of Modernization Expenditures for Electric Utility Distribution Systems: A Guide for Utility Regulators

Cost-effectiveness evaluation of potential grid modernization solutions is integral to integrated distribution system planning (IDSP). This report synthesizes and updates prior cost-effectiveness approaches identified by the U.S. Department of Energy for grid modernization investments, incorporating multi-objective decision-making. The structure of the report follows the IDSP process, from identifying objectives and priorities to prioritizing grid solutions. The report details two methods for conducting initial cost-effectiveness screening of individual and interdependent grid components: (1) Lowest Reasonable Cost and (2) Benefit-Cost Analysis. Drawing on solutions that passed the initial cost-effectiveness screen, the utility prioritizes grid solutions based on policy objectives and priorities, regulatory compliance, operational efficiencies, and other factors to develop a portfolio of distribution solutions. The utility uses multi-objective decision-making to assess each proposed expenditure against each objective and respective metric. The utility applies a weighting factor, reflecting the priority ranking of the objective, to the total numerical “score” based on the contribution of the proposed solution to addressing each objective. Ultimately, the utility ranks each expenditure from highest to lowest priority using the final score for each solution as well as its cost. The report includes examples of emerging best practices for states and utilities and a cost-effectiveness process checklist.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The Poisson tensor completion non-parametric differential entropy estimator

We introduce the Poisson tensor completion (PTC) estimator, a non-parametric differential entropy estimator. The PTC estimator leverages inter-sample relationships to compute a low-rank Poisson tensor decomposition of the frequency histogram. Our crucial observation is that the histogram bins are an instance of a space partitioning of counts and thus can be identified with a spatial Poisson process. The Poisson tensor decomposition leads to a completion of the intensity measure over all bins—including those containing few to no samples—and leads to our proposed PTC differential entropy estimator. A Poisson tensor decomposition models the underlying distribution of the count data and guarantees non-negative estimated values and so can be safely used directly in entropy estimation. Our estimator is the first tensor-based estimator that exploits the underlying spatial Poisson process related to the histogram explicitly when estimating the probability density with low-rank tensor decompositions for the purpose of tensor completion. Furthermore, we demonstrate that our PTC estimator is a substantial improvement over standard histogram-based estimators for sub-Gaussian probability distributions because of the concentration of norm phenomenon.

42 ENGINEERING↗

Dense Image Matching Uncertainty Estimation and Confidence Metrics

Dense stereo matching takes overlapping image pairs as input and outputs a disparity map which encodes pixel-by-pixel matches between the images. Recently, there has been an interest in ranking the quality, or even quantifying the accuracy, of disparity estimates. The proposed methods can be described as either uncertainty estimators or confidence metrics. Uncertainty estimators are a small minority of the research. However, they have the potential to be the most useful because they estimate disparity accuracy (in pixel units) that can be used to threshold matches or carried forward using error propagation. The majority of the research deals with confidence metrics which give an ordinal (or binary) ranking of a match’s quality relative to other matches. Confidence metrics do not have units and thus are useful primarily for thresholding matches from mismatches. The methods could also be described as handcrafted or deep-learning based. The majority of the research focused on outdoor driving scenes. Hence, our interest–application to a satellite semi-global matching pipeline–is a domain shift that may challenge deep-learning based methods. We conclude by recommending five handcrafted and two deep-learning based methods for evaluation in our pipeline.

97 MATHEMATICS AND COMPUTING↗

Advanced Processing of Coal and Coal Waste to Produce Graphite for Fast-Charging Lithium-Ion Battery Anode

The University of North Dakota (UND) Energy & Environmental Research Center (EERC), in collaboration with the UND Center for Process Engineering Research (CPER), conducted a project to validate two technologies capable of converting North Dakota lignite and lignite coal waste to high-quality graphite for fast-charging lithium-ion battery (LIB) anode. The project was conducted over about 3 years from April 7, 2022, to July 6, 2025. The two technological paths pursued in this project include path A – direct conversion of coal or coal waste to graphite by the upgraded carbon ores to products (UCOP) process being developed at the EERC and path B – lignite-derived coal tar pitch (CTP) conversion to graphite (CTP2G) process being developed at CPER. The results from this project validate the two technological approaches and are expected to be an integral part of a portfolio of emerging technologies for making high-quality graphite not only from North Dakota lignite, but from all ranks of U.S. domestic coal and coal waste resources. The quality of the graphite produced by these technologies is high enough for various applications, including batteries for the fast-growing electric vehicle industry, energy storage applications, electric arc furnace electrodes for steel production, and graphene production, among others. Although the two technologies can produce high-quality graphite, they are fundamentally different in that the UCOP technology provides a direct path to transform coal to graphite, while the CTP2G technology needs to go through a CTP intermediate and a coking process for the intermediate, which requires a special facility to accomplish. For application in the industry, the UCOP process is designed to be more flexible, with feedstock to include potentially any carbonaceous material such as all coal ranks and biochar, while the CTP2G process is designed to utilize CTP as the starting precursor. The key project accomplishments include the following: • Successful preparation of high-quality synthetic graphite from North Dakota lignite coal/coal wastes and lignite-derived CTP. • Patent application has been filed for the UCOP process and an internal invention disclosure has been filed for the CTP2G process. • The produced graphite performs better than a commercial battery-grade sample in LIB coin cells, especially fast-charging capability, stability, and long-duration cycling. • Coin-type Li-ion half-cells with CTP2G graphite showed excellent performance, with >370 mAh/g capacity, >90% initial coulombic efficiency, and 93%/67% retention at 1C/2C rate, which outperforms commercial graphite in charging speed, stability, and cycling. • Results of fabricated 18650 cells were consistent with the observations in coin cells. • Preliminary techno-economic analysis (TEA) estimates for the UCOP technology indicate a manufacturing cost of about $\$$39/kg based on 50-metric ton/year capacity. • Preliminary TEA estimates for the CTP2G technology indicate a market price of about $\$$7107/ton ($\$$7/kg) based on 22,000-ton/year production capacity.

01 COAL, LIGNITE, AND PEAT↗

In Situ Data Analysis Through Physics-informed Tensor Decompositions (LDRD Final Report)

We introduce a new low-dimensional model of high-dimensional numerical simulation data based on low-rank tensor decompositions. Our new model aims to minimize differences between the model data and simulation data as well as functions of the model data and functions of the simulation data. This novel approach to dimensionality reduction of simulation data provides a means of directly incorporating quantities of interests and invariants associated with conservation principles associated with the simulation data into the low-dimensional model, thus enabling more accurate analysis of the simulation without requiring access to the full set of high-dimensional data. Computational results of applying this approach to two standard low-rank tensor decompositions of data arising from simulation of combustion and plasma physics are presented.

97 MATHEMATICS AND COMPUTING↗

Approximate Quantum Codes From Long Wormholes

We discuss families of approximate quantum error correcting codes which arise as the nearly-degenerate ground states of certain quantum many-body Hamiltonians composed of non-commuting terms. For exact codes, the conditions for error correction can be formulated in terms of the vanishing of a two-sided mutual information in a low-temperature thermofield double state. We consider a notion of distance for approximate codes obtained by demanding that this mutual information instead be small, and we evaluate this mutual information for the SYK model and for a family of low-rank SYK models. After an extrapolation to nearly zero temperature, we find that both kinds of models produce fermionic codes with constant rate as the number, N , of fermions goes to infinity. For SYK, the distance scales as N 1 / 2 , and for low-rank SYK, the distance can be arbitrarily close to linear scaling, e.g. N .99 , while maintaining a constant rate. We also consider an analog of the no low-energy trivial states property which we dub the no low-energy adiabatically accessible states property and show that these models do have low-energy states that can be prepared adiabatically in a time that does not scale with system size N . We discuss a holographic model of these codes in which the large code distance is a consequence of the emergence of a long wormhole geometry in a simple model of quantum gravity.

Physics↗

Beyond microbial abundance: metadata integration enhances disease prediction in human microbiome studies

Multiple studies have highlighted the interaction of the human microbiome with physiological systems such as the gut, immune, liver, and skin, via key axes. Advances in sequencing technologies and high-performance computing have enabled the analysis of large-scale metagenomic data, facilitating the use of machine learning to predict disease likelihood from microbiome profiles. However, challenges such as compositionality, high dimensionality, sparsity, and limited sample sizes have hindered the development of actionable models. One strategy to improve these models is by incorporating key metadata from both the human host and sample collection/processing protocols. This remains challenging due to sparsity and inconsistency in metadata annotation and availability. In this paper, we introduce a machine learning-based pipeline for predicting human disease states by integrating host and protocol metadata with microbiome abundance profiles from 68 different studies, processed through a consistent pipeline. Our findings indicate that metadata can enhance machine learning predictions, particularly at higher taxonomic ranks like Kingdom and Phylum, though this effect diminishes at lower ranks. Our study leverages a large collection of microbiome datasets comprising 11,208 samples, therefore enhancing the robustness and statistical confidence of our findings. This work is a critical step toward utilizing microbiome and metadata for predicting diseases such as gastrointestinal infections, diabetes, cancer, and neurological disorders.

Mathematics and Computing↗

Housing-Performance Atlas of Baltimore Row Homes: Archetype-Based Multi-Hazard Baseline of Energy, Heat, Survivability, and Durability

Baltimore’s historic row-home neighborhoods face escalating risks to energy, heat, and durability under intensifying climate stress. This study develops a Housing-Performance Atlas that quantifies multi-hazard performance for eight representative archetypes using DesignBuilder/EnergyPlus Version 7.3.1.003, under Baltimore TMY3 boundary conditions. Performance is evaluated across the following four adaptation domains: energy use intensity, passive survivability during 72 h outage events, roof overheating exposure (>150 °F exceedance hours), and material service life derived from ISO 15686 and synthesized into Lean and Full Deficit Indices for comparative resilience ranking. Results show that EUI ranged from 46.7 to 67.6 kBtu ft −2 ·yr −1 , survivability from 0 to 23 h, and roof temperatures exceeded 150 °F for 150–210 h, shortening roof service life by up to 10 years. Composite Lean and Full Deficit Indices ranged 7.8–92.4, ranking Model 5 (end-unit, flat roof, two-story with basement) as the most resilient configuration and Model 8 (end-unit, pitched roof, three-story above-grade) as the least resilient due to compounded overheating and energy losses. Heat-related domains accounted for nearly 70% of overall resilience deficits, confirming thermal safety and roof reflectivity as retrofit priorities. The Housing-Performance Atlas establishes a reproducible diagnostic framework linking simulation, service life, and resilience metrics to guide cost-effective, climate-responsive retrofits in Baltimore’s aging urban housing stock.

Housing-Performance atlas↗

Optimizing Solar PV Deployment in Manufacturing: A Morphological Matrix and Fuzzy TOPSIS Approach

The growing energy demand of the industrial sector and the need for sustainable solutions highlight the importance of efficient decision making in solar photovoltaic (PV) implementation. Selecting optimal PV configuration is complex due to the interdependent technical, economic, environmental, and social factors involved. This study introduces an integrated decision-making method combining a morphological matrix and fuzzy TOPSIS to systematically select and rank optimal PV system configurations for manufacturing firms. While the morphological matrix exhaustively examines possible design solutions based on sensing, smart, sustainable, and social (S4) attributes, the fuzzy TOPSIS method ranks the alternatives by handling uncertainty in decision making. A case study conducted in a Mexican manufacturing company validates the methodology’s effectiveness. The optimal PV configuration identified comprehensively addresses operational and sustainability criteria, covering all lifecycle stages. This approach demonstrates quantitative superiority and greater robustness compared to existing fuzzy TOPSIS-based methods for solar PV applications. The findings highlight the practical value of data-driven, multi-criteria decision making for industrial solar energy adoption, enhancing project feasibility, cost efficiency, and environmental compliance. Future research will incorporate discrete event simulation (DES) to further refine energy consumption strategies in manufacturing.

Briceño, Citlaly Pérez↗