Search NASASearch

SEARCH · Search NASA

Results for “Generalizable model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Neighborhood sociome factors and pediatric asthma exacerbations: Protective role of tree crown density and importance of pharmacy access in Chicago's south side

Abstract Background Pediatric asthma exacerbations remain a critical public health concern, particularly in historically underserved urban settings. Objective This study investigates sociome factors—the social context of disease—associated with asthma exacerbations among children living in Chicago's South Side, leveraging clinical and publicly available generalizable census tract‐level datasets from agencies including ChiVes, the City of Chicago Data Portal, EPA, Census Bureau, HUD, NOAA, and more. The aim is to uncover novel hypotheses for potential new interventions. Methods A generalized linear model assessed associations with the outcome of asthma exacerbations while accounting for clustering at the patient level. Predictors included all variables from the Sociome Data Commons, including social, environmental, behavioral, economic, housing, and school variables. Results Predictors of decreased risk included patient age (+4.8 years, −22%), tree crown density (+6% coverage, −17%), parks per acre (+0.41, −8%), and labor market engagement (+0.8 points, −9%). Conversely, predictors of increased risk included increased distance to the nearest pharmacy (+0.28 miles, +12%), limited English skills (+2.3%, +10%), higher inequality (+0.08 points, +8%), and visits in the Spring (+11%) and Fall (+20%). Conclusion The results suggest that tree crown density, a novel finding in the context of asthma exacerbations, may play a protective role. Limited access to health care facilities such as pharmacies continues to complicate care. Clinical Implications These findings provide hypotheses for future interventions for long‐standing asthma disparities.

Allergy

An uncertainty visualization framework for large-scale cardiovascular flow simulations: A case study on aortic stenosis

We present a generalizable uncertainty quantification (UQ) and visualization framework for lattice Boltzmann method simulations of high Reynolds number vascular flows, demonstrated on a patient-specific stenosed aorta. The framework combines EasyVVUQ for parameter sampling with large-eddy simulation turbulence modeling in HemeLB, and executes ensembles on the Frontier exascale supercomputer. Spatially resolved metrics, including entropy and isosurface-crossing probability, are used to map uncertainty in pressure and wall shear stress fields directly onto vascular geometries. Two sources of model variability are examined: inlet peak velocity and the Smagorinsky constant. Inlet velocity variation produces high uncertainty downstream of the stenosis where turbulence develops, while upstream regions remain stable. Smagorinsky constant variation has little effect on the large-scale pressure field but increases WSS uncertainty in localized high-shear regions. In both cases, the stenotic throat manifests low entropy, indicative of robust identification of elevated WSS. By linking quantitative UQ measures to three-dimensional anatomy, the framework improves interpretability over conventional 1D UQ plots and supports clinically relevant decision-making, with broad applicability to vascular flow problems requiring both accuracy and spatial insight.

Hemodynamics

Developing multi-gene CRISPRa/i programs to accelerate DBTL cycles in ABF hosts engineered for chemical production

This project developed and implemented a modular CRISPR activation and interference (CRISPRa/i) platform to accelerate strain optimization and pathway development for industrially relevant microbial hosts. By integrating multiplexed transcriptional perturbation tools with data-driven Design–Build–Test–Learn (DBTL) workflows, the team achieved reductions in cycle time and enhanced production of industrial aromatics, particularly 4-aminocinnamic acid (4-ACA), in Pseudomonas putida. Key accomplishments included: ● Development of a robust, tunable CRISPRa/i system in P. putida that enabled efficient multi-target gene regulation via guide RNA (gRNA) programs ● Completion of two full DBTL cycles, guided by machine learning (ML) models trained on transcriptomic and performance data, reducing engineering time by over 30% ● Optimization of multi-gene regulatory programs to balance expression of host and pathway modules, improve 4-ACA titers, and resolve metabolic bottlenecks ● Demonstration of system portability through a limited proof-of-concept extension in Acinetobacter baylyi, underscoring the generalizability of the approach ● Evaluation of strain performance on lignocellulosic biomass-derived substrates, demonstrating the feasibility of converting renewable carbon into aromatic building blocks These results illustrate the feasibility of applying ML-guided CRISPRa/i perturbation strategies to accelerate strain development in complex microbial systems. The resulting tools and datasets contribute to DOE objectives by improving platform predictability, reducing development costs, and enabling broader access to sustainable, economically viable bioproduction technologies.

09 BIOMASS FUELS

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat

Maps Suggest Transport and Source Processes of PM2.5 at 1 km x 1 km for the Whole San Joaquin Valley, Winter 2011 (Generalizations from DISCOVER-AQ)

We present interpreted data analysis using MAIAC (Multiangle implementation of Atmospheric Correction) retrievals and appropriate RAPid Update Cycle (RAP) meteorology to map respirable aerosol (PM2.5) for the period January and February, 2011. The San Joaquin Valley is one of the unhealthiest regions in the USA for PM2.5 and related morbidity. The methodology evaluated can be used for the entire moderate-resolution imaging spectrometer (MODIS, VIIRS) data record. Other difficult areas of the West: Riverside, CA, Salt Lake City, UT, and Doa Ana County, NM share similar difficulties and solutions. The maps of boundary layer depth for 1116 hr local time from RAP allows us to interpret aerosol optical thickness as a concentration of particles in a nearly well-mixed box capped by clean air. That mixing is demonstrated by DISCOVER-AQ data and afternoon samples from the airborne measurements, P3B (on-board) and B200 (HSRL2 lidar). This data and the PM2.5 gathered at the deployment sites allowed us to estimate and then evaluate consistency and daily variation of the AOT to PM2.5 relationship. Mixed-effects modeling allowed a refinement of that relation from day to day; RAP mixed layers explained the success of previous mixed-effects modeling. Compositional, size-distribution, and MODIS angle-of-regard effects seem to describe the need for residual daily correction beyond ML depth.We report on an extension method to the entire San Joaquin Valley for all days with MODIS imagery using the permanent PM2.5 stations, evaluated for representativeness. Resulting map movies show distinct sources, particularly Interstate-5 (at approx. 1km x 1km resolution) and the broader Bakersfield area. Accompanying winds suggest transport effects and variable pathways of pollution cleanout. Such estimates should allow morbiditymortality studies. They should be also useful for actual model assimilations, where composition and sources are uncertain. We conclude with a description of new work to extend these insights to similar regions, e.g. interior valleys of California, the Po Valley, the Mediterranean litoral, and the Ganges Plain.This work show generalizable use of remote sensing, a major goal of DISCOVER-AQ, Deriving Information on Surface Conditions from COlumn and VERtically Resolved Observations Relevant to Air Quality.

Chatfield, R.

Combining ToF‐SIMS and Multivariate Analysis to Resolve Active Sites on Ni‐Based HER Catalysts

Unambiguous identification of active sites in heterogeneous catalysis remains a major challenge, particularly for materials with ultrathin, chemically mixed surface layers. Here, we demonstrate a generalizable approach that combines time-of-flight secondary ion mass spectrometry (ToF-SIMS) with multivariate statistical analysis (principal component analysis [PCA] and multivariate curve resolution [MCR]) to resolve catalytically relevant motifs at the nanoscale. Using Ni electrodes as a model system, PCA distinguished hydroxide-enriched domains from oxide- and metal-rich regions, while MCR decomposed depth profiles and 3D images into hydroxide, oxide, and metallic layers with nanometer resolution. A unique secondary-ion fragment, NiO 3 H 3 − (m/z 108.94), emerged as a marker of hydroxide-rich environments and correlated with hydrogen evolution reaction (HER) activity across a series of Ni electrodes. Complementary density functional theory (DFT) calculations revealed that Ni(OH) 2 clusters adjacent to metallic Ni offer the most favorable water dissociation energetics, establishing the structural origin of the marker. While illustrated here for Ni-based HER, this workflow provides a broadly applicable framework to isolate and rank near-surface patterns that govern catalytic activity, thereby extending ToF-SIMS from a qualitative probe to a predictive tool for active site identification.

HER active sites

Strangers in a foreign land: ‘Yeastizing’ plant enzymes

Abstract Expressing plant metabolic pathways in microbial platforms is an efficient, cost‐effective solution for producing many desired plant compounds. As eukaryotic organisms, yeasts are often the preferred platform. However, expression of plant enzymes in a yeast frequently leads to failure because the enzymes are poorly adapted to the foreign yeast cellular environment. Here, we first summarize the current engineering approaches for optimizing performance of plant enzymes in yeast. A critical limitation of these approaches is that they are labour‐intensive and must be customized for each individual enzyme, which significantly hinders the establishment of plant pathways in cellular factories. In response to this challenge, we propose the development of a cost‐effective computational pipeline to redesign plant enzymes for better adaptation to the yeast cellular milieu. This proposition is underpinned by compelling evidence that plant and yeast enzymes exhibit distinct sequence features that are generalizable across enzyme families. Consequently, we introduce a data‐driven machine learning framework designed to extract ‘yeastizing’ rules from natural protein sequence variations, which can be broadly applied to all enzymes. Additionally, we discuss the potential to integrate the machine learning model into a full design‐build‐test cycle.

59 BASIC BIOLOGICAL SCIENCES

Distinguishing Orbiting and Infalling Dark Matter Particles with Machine Learning

Dark matter halos are typically defined as spheres that enclose some overdensity, but these sharp, somewhat arbitrary boundaries introduce nonphysical artifacts such as backsplash halos, pseudo-volution, and an incomplete accounting of halo mass. A more physically motivated alternative is to define halos as the collection of particles that are physically orbiting within their potential well. However, existing methods to classify particles as orbiting or infalling suffer from trade-offs between accuracy, computational cost, and generalizability across cosmologies. We present an efficient, yet accurate, supervised machine learning approach using decision trees. The classification is based on only the particle radii and velocities at two epochs. Compared to detailed analysis of particle trajectories, we find that our model matches the classification of 97% of particles. Consequently, we are able to quickly and accurately reproduce the density profiles of the orbiting and infalling components out to many virial radii. We demonstrate that our model generalizes to a significantly different cosmology that lies outside the training data set. We make publicly available both our final model and the code to train similar models.

79 ASTRONOMY AND ASTROPHYSICS

Roadmap for transforming heterogeneous catalysis with artificial intelligence

Artificial intelligence (AI) is poised to transform heterogeneous catalysis, opening avenues for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental and chemical sectors. This promise, however, hinges on overcoming fundamental barriers, including limitations in data availability and quality, challenges in the generalizability and interpretability of data-augmented decisions, and the persistent gap between in silico predictions and experiments. Furthermore, we outline a forward-looking roadmap for deeply integrating AI into heterogeneous catalysis with an AI-ready data ecosystem, multimodal foundation models, and ultimately autonomous laboratories to accelerate the development of next-generation catalytic technologies via AI-empowered human–machine collaboration.

Computational methods

The unique architecture of umbrella toxins permits a two-tiered molecular bet-hedging strategy for interbacterial antagonism

Bacteria exist in competitive and rapidly changing environments in which the nature of future threats cannot be easily predicted. Streptomyces coelicolor produces three antibacterial umbrella particles that harbor distinct polymorphic toxin domains and an overlapping set of six diversified lectins. Here, we show that the exquisite specificity of umbrella particles derives from lectin-mediated species-specific binding to previously undescribed hypervariable surface glycoconjugates. A cryo-electron microscopy (cryo-EM) structure of one such lectin in complex with its oligosaccharide substrate defines the molecular basis for targeting through the coordinated recognition of multiple glycan features. Biochemical and genetic studies of several target species, in conjunction with lectin-swapping experiments, support a model whereby S. coelicolor umbrella toxin diversification at the levels of lectin composition and toxin polymorphism represents a unique, two-tiered bet-hedging strategy. Bioinformatic analyses support this as a means by which the unusual architecture of umbrella toxins offers Streptomyces a generalizable strategy to antagonize an unpredictable array of competitors.

59 BASIC BIOLOGICAL SCIENCES

The Second Skin: A Wearable Sensor Suite that Enables Real-Time Human Biomechanics Tracking Through Deep Learning

Objective: Real-time determination of human kinematics and kinetics could advance biomechanics research and enable valuable applications of biofeedback and generalizable exoskeleton control. Here, this work aims to investigate a taskindependent, user-independent method for obtaining precise realtime joint state estimation across lower-body joints during a wide variety of tasks. Methods: We developed a generalizable sensing approach using a suit comprised of inertial measurement units (IMUs) and pressure insoles. With the suit, we collected a dataset of 33 tasks commonly performed during construction and hazardous waste cleanup (N = 10). We then trained deep learning user-independent, task-agnostic models to estimate joint lowerbody kinematics and dynamics using only worn sensor data. We likewise computed joint kinematics and dynamics analytically from sensor data to serve as a comparison tool for model results. Results: Our models achieved overall angle estimation root-meansquared-errors (RMSE) of 6.56±.92°, 8.60±1.01°, 7.58±.89°, and 6.00±.73° compared to 13.9±.1.3°, 15.31±1.0°, 10.76±.70°, and 7.56±.48° via analytical methods at the lower back, hip, knee, and ankle, respectively. Likewise, our models achieved overall normalized moment estimation RMSEs of .207±.069 Nm/kg, .242±.044 Nm/kg, .202±.038 Nm/kg, and .193±.034 Nm/kg compared to .306±.036 Nm/kg, .407±.021 Nm/kg, 1.18 ±.022 Nm/kg, and 1.73±.071 Nm/kg via analytical methods at the lower back, hip, knee, and ankle, respectively. Conclusion: These results are comparable to other state-of-the-art wearable sensing systems, establishing deep learning as a viable sensing approach that generalizes to new users and tasks. Significance: This work shows promise for enabling accurate real-world biomechanical data collection and enhancement of biofeedback systems and wearable robot control.

Casey, Ryan T. F. [Georgia Institute of Technology

Generalizable, fast, and accurate DeepQSPR with fastprop

Abstract Quantitative Structure–Property Relationship studies (QSPR), often referred to interchangeably as QSAR, seek to establish a mapping between molecular structure and an arbitrary target property. Historically this was done on a target-by-target basis with new descriptors being devised to specifically map to a given target. Today software packages exist that calculate thousands of these descriptors, enabling general modeling typically with classical and machine learning methods. Also present today are learned representation methods in which deep learning models generate a target-specific representation during training. The former requires less training data and offers improved speed and interpretability while the latter offers excellent generality, while the intersection of the two remains under-explored. This paper introduces , a software package and general Deep-QSPR framework that combines a cogent set of molecular descriptors with deep learning to achieve state-of-the-art performance on datasets ranging from tens to tens of thousands of molecules. provides both a user-friendly Command Line Interface and highly interoperable set of Python modules for the training and deployment of feedforward neural networks for property prediction. This approach yields improvements in speed and interpretability over existing methods while statistically equaling or exceeding their performance across most of the tested benchmarks. is designed with Research Software Engineering best practices and is free and open source, hosted at github.com/jacksonburns/fastprop.

Burns, Jackson W. (ORCID:0000000206579426)

A Comparison of Rotor Disk Modeling and Blade-Resolved CFD Simulations for NASA's Tiltwing Air Taxi

A multi-fidelity computational fluid dynamics analysis is carried out for NASA’s tiltwing air taxi concept operating in airplane and helicopter mode. High-fidelity simulations are computationally expensive due to individual rotor blade modeling in a time-dependent computational domain with rotating grids. The mid-fidelity rotor disk option, in its source term implementation, is explored as a more affordable alternative. Computations are performed with NASA’s OVERFLOW flow solver loosely-coupled with the comprehensive code CAMRAD II for appropriate rotor trim. Detailed comparisons are shown for the trim solution, airloads, wake geometry, and rotor performance. While the rotor disk model is able to capture the flow field with satisfactory agreement in airplane mode, it faces difficulties in helicopter mode due to the three-dimensional effects of the wake. Although this study is limited to a specific vehicle geometry, it is expected that the results are somewhat generalizable to the analysis of multi-rotor configurations.

ARMD

Predicting Airport Runway Configurations for Decision-Support Using Supervised Learning

One of the most challenging tasks for air traffic controllers is runway configuration management (RCM). It deals with the optimal selection of runways to operate on (for arrivals and departures) based on current and forecast of traffic, surface wind speed, wind direction, other environmental variables, noise constraints, and several other airport-specific factors. In this paper, a methodology using supervised learning is developed to build a predictive model for RCM decision-support from large volumes of historical data. Data from two full years (2018 and 2019) related to current and forecast weather, demand/capacity, etc. is collected, analyzed, and fused together. A variety of supervised learning algorithms are tested for predicting runway configuration and hyperparameter tuning is carried out to select the best performing model. The validation process involves two airports of low (Charlotte Douglas International Airport, CLT) and high (Denver International Airport, DEN) complexity of configuration decision-making. The results show significant promise for the two airports with test accuracy of 93% (CLT) and 73% (DEN). The methodology is scalable and generalizable to other airports across the U.S. National Airspace System.

air traffic management

Predicting Airport Runway Configuration for Decision-Support Using Supervised Learning

One of the most challenging tasks for air traffic controllers is runway configuration management (RCM). It deals with the optimal selection of runways to operate on (for arrivals and departures) based on current and forecast of traffic, surface wind speed, wind direction, other environmental variables, noise constraints, and several other airport-specific factors. In this paper, a methodology using supervised learning is developed to build a predictive model for RCM decision-support from large volumes of historical data. Data from two full years (2018 and 2019) related to current and forecast weather, demand/capacity, etc. is collected, analyzed, and fused together. A variety of supervised learning algorithms are tested for predicting runway configuration and hyperparameter tuning is carried out to select the best performing model. The validation process involves two airports of low (Charlotte Douglas International Airport, CLT) and high (Denver International Airport, DEN) complexity of configuration decision-making. The results show significant promise for the two airports with test accuracy of 93% (CLT) and 73% (DEN). The methodology is scalable and generalizable to other airports across the U.S. National Airspace System.

air traffic management

AI-assisted rapid crystal structure generation towards a target local environment

In material design, traditional crystal structure prediction approaches are expensive as they require extensive structural sampling through expensive energy minimization methods. Emerging artificial intelligence (AI) generative models have shown great promise in rapidly generating realistic crystals, but they typically handle only a few tens of atoms per unit cell. To overcome this limitation, we introduce a symmetry-informed approach, the Local Environment Geometry-Oriented Crystal Generator (LEGO-xtal). Our method generates initial structures using AI models trained on an augmented dataset, and then optimizes them using structure descriptors rather than energy-based optimization. We demonstrate its effectiveness by expanding from 25 known low-energy sp2 carbon allotropes to over 1700, all within 0.5 eV/atom of the ground-state energy of graphite. This framework offers a generalizable strategy for the targeted design of materials with modular building blocks, such as metal-organic frameworks and battery materials.

Ridwan, Osman Goni [University of North Carolina a

Commutative Algebra Modeling in Materials Science – A Case Study on Metal–Organic Frameworks (MOFs)

Metal-organic frameworks (MOFs) are a class of important crystalline and highly porous materials whose hierarchical geometry and chemistry hinder interpretable predictions in materials properties. Commutative algebra is a branch of abstract algebra that has been rarely applied in data and material sciences. We introduce the first ever commutative algebra modeling and prediction in materials science. Specifically, category-specific commutative algebra (CSCA) is proposed as a new framework for MOF representation and learning. It integrates element-based categorization with multiscale algebraic invariants to encode both local coordination motifs and global network organization of MOFs. These algebraically consistent, chemically aware representations enable compact, interpretable, and data efficient modeling of MOF properties such as Henry’s constants and uptake capacities for common gases. Compared to traditional geometric and graph-based approaches, CSCA achieves comparable or superior predictive accuracy while substantially improving interpretability and stability across data sets. By aligning commutative algebra with the chemical hierarchy, the CSCA establishes a rigorous and generalizable paradigm for understanding structure and property relationships in porous materials and provides a nonlinear algebra-based framework for data-driven material discovery.

Khaemba, Caleb S.

Scenario storyline discovery for complex multi-actor human-natural systems

This poster was presented at the AGU Fall Meeting 2024. Abstract:Scenario analysis is a useful tool for assessing the impacts of future conditions or alternative strategies. However, the common practice of focusing on a small number of predetermined scenarios can limit our understanding of key uncertainties, and fail to represent diverse stakeholder impacts. Exploratory modeling approaches have been developed to address these issues by simulating a wide range of possible futures and system perspectives. A challenge with these approaches is that they often involve large ensemble experiments which limit interpretability and usability. We recently introduced the FRamework for Narrative Storylines and Impact Classification (FRNSIC; pronounced ``forensic''), a scenario discovery framework that helps users identify scenario storylines that capture key system dynamics and as well as important outcomes. In this poster presentation, we present training materials to support the generalizable application of the framework to other multi-actor systems with complex dynamics. Specifically, we will present a step-by-step methodological typology of tools and methods that can be used to generate and classify plausible states of the world on key metrics and consequential dynamics. The typology will also discuss potential implications of these choices and their applicability to different systems.

Colorado River