Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computational pipelines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

COMPILE: a GWAS computational pipeline for gene discovery in complex genomes

Abstract Background Genome-Wide Association Studies (GWAS) are used to identify genes and alleles that contribute to quantitative traits in large and genetically diverse populations. However, traits with complex genetic architectures create an enormous computational load for discovery of candidate genes with acceptable statistical certainty. We developed a streamlined computational pipeline for GWAS (COMPILE) to accelerate identification and annotation of candidate maize genes associated with a quantitative trait, and then matches maize genes to their closest rice and Arabidopsis homologs by sequence similarity. Results COMPILE executed GWAS using a Mixed Linear Model that incorporated, without compression, recent advancements in population structure control, then linked significant Quantitative Trait Loci (QTL) to candidate genes and RNA regulatory elements contained in any genome. COMPILE was validated using published data to identify QTL associated with the traits of α-tocopherol biosynthesis and flowering time, and identified published candidate genes as well as additional genes and non-coding RNAs. We then applied COMPILE to 274 genotypes of the maize Goodman Association Panel to identify candidate loci contributing to resistance of maize stems to penetration by larvae of the European Corn Borer ( Ostrinia nubilalis ). Candidate genes included those that encode a gene of unknown function, WRKY and MYB-like transcriptional factors, receptor-kinase signaling, riboflavin synthesis, nucleotide-sugar interconversion, and prolyl hydroxylation. Expression of the gene of unknown function has been associated with pathogen stress in maize and in rice homologs closest in sequence identity. Conclusions The relative speed of data analysis using COMPILE allowed comparison of population size and compression. Limitations in population size and diversity are major constraints for a trait and are not overcome by increasing marker density. COMPILE is customizable and is readily adaptable for application to species with robust genomic and proteome databases.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A computational pipeline to generate a synthetic dataset of metal ion sorption to oxides for AI/ML exploration

The charged mineral/electrolyte interfaces are ubiquitous in the surface and subsurface–including the surroundings of the geological disposal sites for radioactive waste. Therefore, understanding how ions interact with charged surfaces is critically important for predicting radionuclide mobility in the case of waste leakage. At present, the Surface Complexation Models (SCMs) are the most successful thermodynamic frameworks to describe ion retention by mineral surfaces. SCMs are interfacial speciation models that account for the effect of the electric field generated by charged surfaces on sorption equilibria. These models have been successfully used to analyze and interpret a broad range of experimental observations including potentiometric and electrokinetic titrations or spectroscopy. Unfortunately, many of the current procedures to solve and fit SCM to experimental data are not optimal, which leads to a non-transferable or non-unique description of interfacial electrostatics and consequently of the strength and extent of ion retention by mineral surfaces. Recent developments in Artificial Intelligence (AI) offer a new avenue to replace SCM solvers and fitting algorithms with trained AI surrogates. Unfortunately, there is a lack of a standardized dataset covering a wide range of SCM parameter values available for AI exploration and training–a gap filled by this study. Here, we described the computational pipeline to generate synthetic SCM data and discussed approaches to transform this dataset into AI-learnable input. First, we used this pipeline to generate a synthetic dataset of electrostatic properties for a broad range of the prototypical oxide/electrolyte interfaces. The next step is to extend this dataset to include complex radionuclide sorption and complexation, and finally, to provide trained AI architectures able to infer SCMs parameter values rapidly from experimental data. Here, we illustrated the AI-surrogate development using the ensemble learning algorithms, such as Random Forest and Gradient Boosting. These surrogate models allow a rapid prediction of the SCM model parameters, do not rely on an initial guess, and guarantee convergence in all cases.

Li, Chunhui↗

Deep-Learning-Derived Evaluation Metrics Enable Effective Benchmarking of Computational Tools for Phosphopeptide Identification

Tandem mass spectrometry (MS/MS)-based phosphoproteomics is a powerful technology for global phosphorylation analysis. However, applying four computational pipelines to a typical mass spectrometry (MS)-based phosphoproteomic dataset from a human cancer study, we observed a large discrepancy among the reported phosphopeptide identification and phosphosite localization results, underscoring a critical need for benchmarking. While efforts have been made to compare performance of computational pipelines using data from synthetic phosphopeptides, evaluations involving real application data have been largely limited to comparing the numbers of phosphopeptide identifications due to the lack of appropriate evaluation metrics. We investigated three deep learning-derived features as potential evaluation metrics: phosphosite probability, Delta RT and spectral similarity. Predicted phosphosite probability is computed by MusiteDeep, which provides high accuracy as previously reported; Delta RT is defined as the absolute retention time (RT) difference between RTs observed and predicted by AutoRT; and spectral similarity is defined as the Pearson’s correlation coefficient between spectra observed and predicted by pDeep2. Using a synthetic peptide dataset, we found that both Delta RT and spectral similarity provided excellent discrimination between correct and incorrect peptide-spectrum matches (PSMs) both when incorrect PSMs involved wrong peptide sequences and even when incorrect PSMs were caused by only incorrect phosphosite localization. Based on these results, we used all the three deep learning-derived features as evaluation metrics to compare different computational pipelines on diverse set of phosphoproteomic datasets and showed their utility in benchmarking performance of the pipelines. The benchmark metrics demonstrated in this study will enable users to select computational pipelines and parameters for routine analysis of phosphoproteomics data and will offer guidance for developers to improve computational methods.

59 BASIC BIOLOGICAL SCIENCES↗

Computer Vision Pipeline for Image Analysis for Freeze‐Fracture Electron Microscopy: Rosette Cellulose Synthase Complexes Case

In materials science, plant biology, agriculture, and environmental research, the automated analysis of high-magnification, complex microscopy images, such as those generated by freeze-fracture electron microscopy (FF-TEM), remains a critical challenge that limits the scalability of data interpretation. We present a deep learning computer vision pipeline for high-throughput detection and morphological characterization analysis of cellulose synthase complexes (CSCs, or rosettes) in FF-TEM images. The pipeline integrates preprocessing, detection, human-in-the-loop verification, and semantic segmentation to quantify features such as rosette diameter and inter-lobe spacing. The approach was trained and tested on a curated dataset of high-resolution FF-TEM micrographs of Physcomitrium patens, expanded via strategic tiling and augmentation to over 650 images. We compare YOLOv8 and YOLOv9 architectures and demonstrate that YOLOv9 achieves superior performance in both localization accuracy (mAP50-95 = 0.854) and inference speed. The resulting distributions revealed biological variability consistent with prior manual studies, validating the approach for high-throughput applications. Our results show that the pipeline achieves human-expert level accuracy while dramatically reducing analysis time, enabling scalable, reproducible structural characterization of intramembrane protein complexes. The pipeline is broadly applicable to other domains requiring precise interpretation of complex microscopy data and establishes a foundation for future artificial intelligence (AI)-assisted workflows in biological imaging.

59 BASIC BIOLOGICAL SCIENCES↗

Biosynth Pipeline v1.0

BioPKS Pipeline is a computational pipeline for retrosynthetic design of small molecule biosynthesis pathways (e.g. retrobiosynthesis). It combines capabilities by interfacing with existing retrobiosynthesis tools- RetroTide (developed at LBNL) and DORAnet to create pathways that combine multiple biosynthesis approaches- both megasynthase assembly line enzymes and single step enzymes.

Backman, Tyler [Lawrence Berkeley National Laborat↗

Learning epistatic polygenic phenotypes with Boolean interactions

Detecting epistatic drivers of human phenotypes is a considerable challenge. Traditional approaches use regression to sequentially test multiplicative interaction terms involving pairs of genetic variants. For higher-order interactions and genome-wide large-scale data, this strategy is computationally intractable. Moreover, multiplicative terms used in regression modeling may not capture the form of biological interactions. Building on the Predictability, Computability, Stability (PCS) framework, we introduce the epiTree pipeline to extract higher-order interactions from genomic data using tree-based models. The epiTree pipeline first selects a set of variants derived from tissue-specific estimates of gene expression. Next, it uses iterative random forests (iRF) to search training data for candidate Boolean interactions (pairwise and higher-order). We derive significance tests for interactions, based on a stabilized likelihood ratio test, by simulating Boolean tree-structured null (no epistasis) and alternative (epistasis) distributions on hold-out test data. Finally, our pipeline computes PCS epistasis p-values that probabilisticly quantify improvement in prediction accuracy via bootstrap sampling on the test set. We validate the epiTree pipeline in two case studies using data from the UK Biobank: predicting red hair and multiple sclerosis (MS). In the case of predicting red hair, epiTree recovers known epistatic interactions surrounding MC1R and novel interactions, representing non-linearities not captured by logistic regression models. In the case of predicting MS, a more complex phenotype than red hair, epiTree rankings prioritize novel interactions surrounding HLA-DRB1 , a variant previously associated with MS in several populations. Taken together, these results highlight the potential for epiTree rankings to help reduce the design space for follow up experiments.

59 BASIC BIOLOGICAL SCIENCES↗

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S↗

Unraveling design principles of protein landscapes in photosynthetic membranes in plant chloroplasts

The supramolecular organization of proteins within photosynthetic membranes is crucial for energy conversion in plants. Here, we introduce an analytical and computational pipeline that integrates high-resolution cryo–scanning electron microscopy, biochemical quantification, advanced Monte Carlo computer simulations, and statistical methods to elucidate the elusive protein landscapes of grana membranes in intact Arabidopsis leaves. Our integrated analysis challenges the prevailing view that particles on the exoplasmic fracture faces in freeze-fracture samples represent photosystem II exclusively. Instead, these particles also include cytochrome b 6 f complexes. Furthermore, our steric clash analysis demonstrates that stacked membranes contain a mixture of larger PSII supercomplexes (C 2 S 2 M 2 and C 2 S 2 ) in addition to a smaller complex (C 2 ). This suggests that in vivo PSII supercomplexes exist in an equilibrium distribution of differing sizes. Furthermore, we discovered that, although size exclusion effects govern the global protein arrangement, local packing exhibits orientational order indicative of lateral attractive protein-protein interactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

OES CO 2 Pipeline FEED Project Design Basis Memorandum

The OES CO₂ Pipeline project will move captured carbon dioxide from two ethanol facilities near Gibson City, Illinois, roughly 7.8 miles southeast to three injection wells outside Anchor, where it will be permanently stored underground. The system is designed to handle up to 4.5 million metric tonnes per year of dense-phase CO₂ at pressures up to 2,500 psig, using 16-inch mainline pipe and 10.750-inch laterals made from API 5L X-60 and X-65 steel. Wall thicknesses vary depending on location, with thinner pipe in open country, heavier wall at road crossings, and the heaviest where the pipe passes under highways or railroads via horizontal directional drill. The pipe gets a fusion-bonded epoxy coating, with an added abrasion-resistant layer wherever it's bored or drilled. Major water crossings will use HDD rather than open trenching. The pipeline will be cathodically protected, equipped with SCADA-compatible pressure and temperature instrumentation, and monitored for leaks using a computational pipeline monitoring system per API RP 1130. Hydrostatic testing will be performed at 1.25 times design pressure, and an ILI caliper run will follow to catch any construction defects. Several items, including fracture toughness requirements, specific NDE methods, and ILI tool selection, are left for the detailed design phase. The whole system falls under 49 CFR Part 195 and ASME B31.4, and Gulf Interstate Engineering prepared this document as the FEED-level design basis under the CarbonSAFE Phase III program.

09 BIOMASS FUELS↗

Predicting drug-metagenome interactions: Variation in the microbial β-glucuronidase level in the human gut metagenomes

Characterizing the gut microbiota in terms of their capacity to interfere with drug metabolism is necessary to achieve drug efficacy and safety. Although examples of drug-microbiome interactions are well-documented, little has been reported about a computational pipeline for systematically identifying and characterizing bacterial enzymes that process particular classes of drugs. The goal of our study is to develop a computational approach that compiles drugs whose metabolism may be influenced by a particular class of microbial enzymes and that quantifies the variability in the collective level of those enzymes among individuals. The present paper describes this approach, with microbial β-glucuronidases as an example, which break down drug-glucuronide conjugates and reactivate the drugs or their metabolites. We identified 100 medications that may be metabolized by β-glucuronidases from the gut microbiome. These medications included morphine, estrogen, ibuprofen, midazolam, and their structural analogues. The analysis of metagenomic data available through the Sequence Read Archive (SRA) showed that the level of β-glucuronidase in the gut metagenomes was higher in males than in females, which provides a potential explanation for the sex-based differences in efficacy and toxicity for several drugs, reported in previous studies. Our analysis also showed that infant gut metagenomes at birth and 12 months of age have higher levels of β-glucuronidase than the metagenomes of their mothers and the implication of this observed variability was discussed in the context of breastfeeding as well as infant hyperbilirubinemia. Overall, despite important limitations discussed in this paper, our analysis provided useful insights on the role of the human gut metagenome in the variability in drug response among individuals. Importantly, this approach exploits drug and metagenome data available in public databases as well as open-source cheminformatics and bioinformatics tools to predict drug-metagenome interactions.

59 BASIC BIOLOGICAL SCIENCES↗

EvoNet: A phylogenomic and systems biology approach to identify genes underlying plant survival in marginal, low‐N soils

The DOE‐BER “EvoNet” project investigates the genetic and molecular basis of plant resilience in extreme environments. We do this by identifying key genes that enable “extreme survivor” species to thrive in the nitrogen-poor soils of Chile’s hyper-arid Atacama Desert. Our collections focus on 32 Atacama extremophile species, including seven grass species with potential biofuel applications. To identify genes-of-importance to survival we compared genomic and transcriptomic profiles of extremophile species that thrive in the Atacama to those of closely related “sister” species from nitrogen-rich arid and mesic regions of California. Deep RNA sequencing and de novo transcriptome assembly across these triplet species sets supported a phylogenomic framework for identifying positively selected genes associated with adaptive divergence. Our integrative analysis combined ecological and environmental data, metagenomics, evolutionary and systems biology, and metabolomics. This enabled us to create an unprecedented framework for systematically understanding how non-model plants have adapted to survive in extreme conditions. Our resulting database of positively selected ortholog groups in the extremophile plants offers promising targets for engineering crop and biofuel species with enhanced resilience to drought and extreme weather. Additionally, our newest dataset explores and exploits a complementary metabolomic approach. This new aspect provides innovative strategies to manipulate plant cell metabolism, further supporting efforts to improve agricultural productivity in the face of extreme climates. Importantly, our combined evolutionary- and metabolomic-based strategies focused on convergent patterns of adaptation, providing a genetic and metabolomic toolkit for improving crop and biofuel resilience across diverse plant species. Finally, our novel exploration of ecological and evolutionary dynamics delivered to the community a phylogenomic computational pipeline called “PhyloGeneious.” Our continued adaptations of this pipeline are publicly available to expedite evolutionary genomic research for future scientific discoveries. In total, our DOE-BER has provided genomic, metabolomic, and computational strategies to understand how extremophile plants provide evolutionary and physiological targets for improving agricultural and biofuel production.

59 BASIC BIOLOGICAL SCIENCES↗

Computationally Guided and Experimentally Validated Design of Custom Chelators for Critical Mineral Recovery

Selective, high throughput separation of target critical metals from complex environments such as fly ash leachates and mining process streams presents a significant challenge for economical production. Custom chelators and sorbents are an attractive technology for selective metal extraction, however it can be difficult to predict their performance, and significant experimental efforts are often required to develop chelating technologies. Here, we present a computational strategy focused on modelling chelator-metal binding interactions and benchmark these results versus experimental data. A computational pipeline combining forcefield, semiempirical, and meta-GGA methods with a thermodynamic framework optimized for error cancellation has been developed to predict binding energies of chelator complexes towards critical mineral recovery applications. This approach, originally validated on [2.2.2] cryptates binding mono- and divalent cations, demonstrated robust predictive capabilities with an R2 of 0.850 against experimental aqueous binding energies. The workflow includes metadynamics for exploring high-dimensional potential energy surfaces and a cluster-continuum model for accurate yet computationally efficient solvation modeling. Error cancellation between solvation energies of free and chelator-coordinated ions enables faster convergence, even with finite cluster sizes. Initial studies on the cryptates revealed consistent metal-ligand coordination patterns, with systematic variations influenced by ion size and charge, highlighting key structural features linked to binding selectivity. Further studies of a proprietary chelator have resulted in identification of previously unreported selectivity towards economically significant metals, which in-house experiments have confirmed, demonstrating the feasibility of this approach. By applying this methodology to new chelators targeting critical minerals such as lithium, cobalt, nickel and other strategic metals, we aim to accelerate the discovery of next-generation chelators for efficient recovery, recycling, and separation processes. This computational framework serves as the backbone of a high-throughput design pipeline tailored for sustainable resource utilization and may be applied to a wide range of systems to meet experimental needs.

computational materials↗

A reusable neural network pipeline for unidirectional fiber segmentation

Abstract Fiber-reinforced ceramic-matrix composites are advanced, temperature resistant materials with applications in aerospace engineering. Their analysis involves the detection and separation of fibers, embedded in a fiber bed, from an imaged sample. Currently, this is mostly done using semi-supervised techniques. Here, we present an open, automated computational pipeline to detect fibers from a tomographically reconstructed X-ray volume. We apply our pipeline to a non-trivial dataset by Larson et al . To separate the fibers in these samples, we tested four different architectures of convolutional neural networks. When comparing our neural network approach to a semi-supervised one, we obtained Dice and Matthews coefficients reaching up to 98%, showing that these automated approaches can match human-supervised methods, in some cases separating fibers that human-curated algorithms could not find. The software written for this project is open source, released under a permissive license, and can be freely adapted and re-used in other domains.

79 ASTRONOMY AND ASTROPHYSICS↗

In‐Silico Device Performance Prediction of Cosensitizer Dye Pairs for Dye‐Sensitized Solar Cells

Abstract Endeavors in the field of dye‐sensitized solar cells (DSCs) have shown great promise when adopting a data‐driven approach to materials discovery, such as successful molecular‐scale predictions of light‐harvesting chromophores. However, predictions of DSC dyes would become much more sophisticated if a molecular‐to‐macroscopic DSC device prediction methodology existed. Thereby, a fully computational pipeline is presented that predicts device‐performance parameters of DSCs which contain varying dye combinations. Optimal pairing of complementary dyes is identified via a data‐driven workflow that affords cosensitized DSCs with maximum power‐conversion efficiencies. Six high‐performing DSC dyes are paired with partner dyes that are screened from a database of 8488 compounds using sequential heuristic filters. Existing models that predict short‐circuit‐current density ( J SC ) and open‐circuit voltage ( V OC ) parameters are adapted to predict singly sensitized and cosensitized DSC performance. The predictions for J sc values of singly sensitized devices match experimental literature values with comparable accuracy to more computationally costly methods. Five out of six dye pairings are predicted to have greater J SC values when cosensitized compared to their corresponding singly sensitized devices, including two pairs that show strong J sc boosts of +13% and +12% when cosensitized. Thus, the prospect of an entirely in‐silico prediction pipeline for DSC performance that can be used to realize the fully automated design of optimized cosensitized DSCs is demonstrated.

14 SOLAR ENERGY↗

Machine Learning for Synchrophasor Analysis

The report presents results from the development of a cloud-based, Big Data analysis framework for power systems. The computational pipeline uses the Apache Spark framework running in an OpenStack cloud infrastructure. A real-world phasor measurement unit (PMU) dataset has been used to carry out the analysis. Several Machine Learning (ML) methods have been developed and implemented for event and anomaly detection and classification. Actual examples of power system events detection and analysis using synchrophasor data are presented. It has been shown that applications of the cloud-based computing environment and the Apache Spark framework enable a significant increase in the computational efficiency of large-scale PMU data analysis.

20 FOSSIL-FUELED POWER PLANTS↗

Portable Parallel Algorithms and Frameworks for Exascale Graph Analytics

Graphs (or networks) are a tool used to model the interactions among various entities. Efficiently processing large graphs has recently attracted significant attention due to the applications of graphs in various domains, such as biology, chemistry, and cyber-security. Analyzing the structure and properties of these graphs is an important component of many scientific computing pipelines. With the explosion in the volume of data, graphs have become very large and can contain hundreds of billions of vertices and trillions of edges. Therefore, it is crucial to develop high-performance methods to enable graph analysis to be done quickly and energy-efficiently. Furthermore, these solutions should be highly parallel in order to take advantage of modern parallel machines. However, designing efficient solutions is not enough. With the wide variety of computing environments available, each with different programmability and performance characteristics, it is necessary to develop solutions that are portable in terms of both performance (i.e., provide theoretical guarantees) and programmability (i.e., provide high level abstractions).

97 MATHEMATICS AND COMPUTING↗

Tools Assessing Performance (TAP) 2.0

Dmitry Duplyakin will be presenting on the latest research and results in the Tools Assessing Performance (TAP) 2.0 project. This presentation will include updates on the latest data the group has produced, integration of obstacle models in the computational pipeline for distributed wind siting, and the plans for the near-term analysis and validation efforts. The talk will acknowledge the work of collaborators from NREL and three other national labs - ANL, LANL, and PNNL - all contributing to this multi-year project.

distributed wind↗