Search NASA⌕ Search

SEARCH · Search NASA

Results for “computational toxicology”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Structure–activity relationship-based chemical classification of highly imbalanced Tox21 datasets

Abstract The specificity of toxicant-target biomolecule interactions lends to the very imbalanced nature of many toxicity datasets, causing poor performance in Structure–Activity Relationship (SAR)-based chemical classification. Undersampling and oversampling are representative techniques for handling such an imbalance challenge. However, removing inactive chemical compound instances from the majority class using an undersampling technique can result in information loss, whereas increasing active toxicant instances in the minority class by interpolation tends to introduce artificial minority instances that often cross into the majority class space, giving rise to class overlapping and a higher false prediction rate. In this study, in order to improve the prediction accuracy of imbalanced learning, we employed SMOTEENN, a combination of Synthetic Minority Over-sampling Technique (SMOTE) and Edited Nearest Neighbor (ENN) algorithms, to oversample the minority class by creating synthetic samples, followed by cleaning the mislabeled instances. We chose the highly imbalanced Tox21 dataset, which consisted of 12 in vitro bioassays for > 10,000 chemicals that were distributed unevenly between binary classes. With Random Forest (RF) as the base classifier and bagging as the ensemble strategy, we applied four hybrid learning methods, i.e., RF without imbalance handling (RF), RF with Random Undersampling (RUS), RF with SMOTE (SMO), and RF with SMOTEENN (SMN). The performance of the four learning methods was compared using nine evaluation metrics, among which F 1 score, Matthews correlation coefficient and Brier score provided a more consistent assessment of the overall performance across the 12 datasets. The Friedman’s aligned ranks test and the subsequent Bergmann-Hommel post hoc test showed that SMN significantly outperformed the other three methods. We also found that a strong negative correlation existed between the prediction accuracy and the imbalance ratio (IR), which is defined as the number of inactive compounds divided by the number of active compounds. SMN became less effective when IR exceeded a certain threshold (e.g., > 28). The ability to separate the few active compounds from the vast amounts of inactive ones is of great importance in computational toxicology. This work demonstrates that the performance of SAR-based, imbalanced chemical toxicity classification can be significantly improved through the use of data rebalancing.

Idakwo, Gabriel↗

Application of Cell Painting for chemical hazard evaluation in support of screening-level chemical assessments

‘Cell Painting’ is an imaging-based high-throughput phenotypic profiling (HTPP) method in which cultured cells are fluorescently labeled to visualize subcellular structures (i.e., nucleus, nucleoli, endoplasmic reticulum, cytoskeleton, Golgi apparatus / plasma membrane and mitochondria) and to quantify morphological changes in response to chemicals or other perturbagens. HTPP is a high-throughput and cost-effective bioactivity screening method that detects effects associated with many different molecular mechanisms in an untargeted manner, enabling rapid in vitro hazard assessment for thousands of chemicals. Here, 1201 chemicals from the ToxCast library were screened in concentration-response up to ~100 μM in human U-2 OS cells using HTPP. A phenotype altering concentration (PAC) was estimated for chemicals active in the tested range. PACs tended to be higher than lower bound potency values estimated from a broad collection of targeted high-throughput assays, but lower than the threshold for cytotoxicity. In vitro to in vivo extrapolation (IVIVE) was used to estimate administered equivalent doses (AEDs) based on PACs for comparison to human exposure predictions. AEDs for 18/412 chemicals overlapped with predicted human exposures. Phenotypic profile information was also leveraged to identify putative mechanisms of action and group chemicals. Of 58 known nuclear receptor modulators, only glucocorticoids and retinoids produced characteristic profiles; and both receptor types are expressed in U-2 OS cells. Thirteen chemicals with profile similarity to glucocorticoids were tested in a secondary screen and one chemical, pyrene, was confirmed by an orthogonal gene expression assay as a novel putative GR modulating chemical. Most active chemicals demonstrated profiles not associated with a known mechanism-of-action. However, many structurally related chemicals produced similar profiles, with exceptions such as diniconazole, whose profile differed from other active conazoles. Overall, the present study demonstrates how HTPP can be applied in screening-level chemical assessments through a series of examples and brief case studies.

60 APPLIED LIFE SCIENCES↗

CATMoS: Collaborative Acute Toxicity Modeling Suite

Background: Humans are exposed to tens of thousands of chemical substances that need to be assessed for their potential toxicity. Acute systemic toxicity testing serves as the basis for regulatory hazard classification, labeling, and risk management. However, it is cost- and time-prohibitive to evaluate all new and existing chemicals using traditional rodent acute toxicity tests. In silico models built using existing data facilitate rapid acute toxicity predictions without using animals. Objectives: The U.S. Interagency Coordinating Committee on the Validation of Alternative Methods Acute Toxicity Workgroup organized an international collaboration to develop in silico models for predicting acute oral toxicity based on five different endpoints: LD50 value, U.S. Environmental Protection Agency hazard categories, Globally Harmonized System for Classification and Labelling hazard categories, very toxic chemicals (LD50 =50 mg/kg), and non-toxic chemicals (LD50 >2000 mg/kg). Methods: An acute oral toxicity data inventory for 11,992 chemicals was compiled, split into training and evaluation sets, and made available to 35 participating international research groups that submitted a total of 139 predictive models. Predictions that fell within the applicability domains of the submitted models were evaluated using external validation sets. These were then combined into consensus models to leverage strengths of individual approaches. Results: The resulting consensus predictions, which leverage the collective strengths of each individual model, form the Collaborative Acute Toxicity Modeling Suite (CATMoS). CATMoS demonstrated high performance in terms of accuracy and robustness when compared to in vivo results. Discussion: CATMoS is being evaluated by regulatory agencies for its utility and applicability as a potential replacement for in vivo rat acute oral toxicity studies. CATMoS predictions for over 800,000 chemicals have been made available via the NTP’s Integrated Chemical Environment. The models are also implemented in a free, standalone open-source tool, OPERA, which allows predictions of new and untested chemicals to be made.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Efficient bi-directional coupling of 3D computational fluid-particle dynamics and 1D Multiple Path Particle Dosimetry lung models for multiscale modeling of aerosol dosimetry

The development of predictive aerosol dosimetry models has been a major focus of environmental toxicology and pharmaceutical health research for decades. One-dimensional (1D) models successfully predict overall deposition averages but fail to accurately predict local deposition. Computational fluid-particle dynamics (CFPD) models provide site-specific predictions but at a computational cost that prohibits whole lung predictions. Thus, there is a need for developing multiscale strategies to provide a realistic subject-specific picture of the fate of inhaled aerosol in the lungs. CT-based 3D/CFPD models of the large airways were bidirectionally coupled with individualized 1D Navier-Stokes airflow and particle transport based upon the widely used Multiple Path Particle Dosimetry Model (MPPD). Distribution of airflows among lobes was adjusted by measured lobar volume changes observed in CT images between FRC and FRC + 1.5 L. Additionally, as a test of the effectiveness of the coupling procedures, deposition modeling of previous 1 µm aerosol exposure studies was performed. The complete coupled model was run for 3 breaths, with the computation-intense portion being the 3D CFD Lagrangian particle tracking calculation. The average deposition per breath was 11% in the combined multiscale model with site-specific doses available in the CFPD portion of the model and airway- or region-specific deposition available for the MPPD portion. In conclusion, the key methods developed in this study enable predictions of ventilation heterogeneities and aerosol deposition across the lungs that are not captured by 3D or 1D models alone. Overall, these methods can be used as the foundation for multi-scale modeling of the full respiratory system.

60 APPLIED LIFE SCIENCES↗

Transplatformer: translating toxicogenomic profiles between generations of platforms

Background Transcriptomic profiling technologies have advanced the analysis of biological and toxicological responses. However, substantial differences in probe design, dynamic range, gene coverage, and preprocessing pipelines across platforms introduce artifacts that limit cross-study integration and hinder the reuse of historical datasets. We aim to develop computational methods for accurate cross-platform translation to maximize the value of legacy resources. Results We present TransPlatformer a deep learning framework for translating gene expression profiles across heterogeneous toxicogenomics platforms. TransPlatformer employs a novel attention-based architecture to map high-dimensional fold-change vectors from legacy microarray technologies to current platforms. Models are trained and evaluated using DrugMatrix, spanning three technological generations. We investigate mixed-tissue, single-tissue, and cross-tissue training paradigms and benchmark performance against multilayer perceptron and matrix-completion baselines. In mixed-tissue training, TransPlatformer achieves a greater than 50% reduction in mean absolute error (0.043 vs. 0.09) and nearly doubles Pearson correlation ( ≈ 0.71 vs. 0.37) relative to baseline methods. Importantly, TransPlatformer preserves rare but biologically meaningful over- and under-expressed signals, with mean absolute error below 0.22. Single-tissue models yield further improvements for well-represented organs, such as a 10% reduction in liver mean absolute error, while underscoring the need for data augmentation strategies in low-sample tissues.ra Conclusions TransPlatformer provides an effective and scalable computational solution for cross-platform transcriptomic translation. By enabling biologically faithful harmonization of gene expression data, the proposed approach facilitates the reuse of legacy toxicogenomics datasets, enhances downstream biomarker discovery, and supports more reproducible predictive modeling in toxicology.

59 BASIC BIOLOGICAL SCIENCES↗

Inferring pesticide toxicity to honey bees from a field‐based feeding study using a colony model and Bayesian inference

Abstract Honey bees are crucial pollinators for agricultural crops but are threatened by a multitude of stressors including exposure to pesticides. Linking our understanding of how pesticides affect individual bees to colony‐level responses is challenging because colonies show emergent properties based on complex internal processes and interactions among individual bees. Agent‐based models that simulate honey bee colony dynamics may be a tool for scaling between individual and colony effects of a pesticide. The U.S. Environmental Protection Agency (USEPA) and U.S. Department of Agriculture (USDA) are developing the VarroaPop + Pesticide model, which simulates the dynamics of honey bee colonies and how they respond to multiple stressors, including weather, Varroa mites, and pesticides. To evaluate this model, we used Approximate Bayesian Computation to fit field data from an empirical study where honey bee colonies were fed the insecticide clothianidin. This allowed us to reproduce colony feeding study data by simulating colony demography and mortality from ingestion of contaminated food. We found that VarroaPop + Pesticide was able to fit general trends in colony population size and structure and reproduce colony declines from increasing clothianidin exposure. The model underestimated adverse effects at low exposure (36 µg/kg), however, and overestimated recovery at the highest exposure level (140 µg/kg), for the adult and pupa endpoints, suggesting that mechanisms besides oral toxicity‐induced mortality may have played a role in colony declines. The VarroaPop + Pesticide model estimates an adult oral LD 50 of 18.9 ng/bee (95% CI 10.1–32.6) based on the simulated feeding study data, which falls just above the 95% confidence intervals of values observed in laboratory toxicology studies on individual bees. Overall, our results demonstrate a novel method for analyzing colony‐level data on pesticide effects on bees and making inferences on pesticide toxicity to individual bees.

59 BASIC BIOLOGICAL SCIENCES↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Review of machine learning and deep learning models for toxicity prediction

The ever-increasing number of chemicals has raised public concerns due to their adverse effects on human health and the environment. To protect public health and the environment, it is critical to assess the toxicity of these chemicals. Traditional in vitro and in vivo toxicity assays are complicated, costly, and time-consuming and may face ethical issues. These constraints raise the need for alternative methods for assessing the toxicity of chemicals. Recently, due to the advancement of machine learning algorithms and the increase in computational power, many toxicity prediction models have been developed using various machine learning and deep learning algorithms such as support vector machine, random forest, k-nearest neighbors, ensemble learning, and deep neural network. This review summarizes the machine learning- and deep learning-based toxicity prediction models developed in recent years. Support vector machine and random forest are the most popular machine learning algorithms, and hepatotoxicity, cardiotoxicity, and carcinogenicity are the frequently modeled toxicity endpoints in predictive toxicology. It is known that datasets impact model performance. The quality of datasets used in the development of toxicity prediction models using machine learning and deep learning is vital to the performance of the developed models. The different toxicity assignments for the same chemicals among different datasets of the same type of toxicity have been observed, indicating benchmarking datasets is needed for developing reliable toxicity prediction models using machine learning and deep learning algorithms. This review provides insights into current machine learning models in predictive toxicology, which are expected to promote the development and application of toxicity prediction models in the future.

Research & Experimental Medicine↗

TransPlatformer

We propose TransPlatformer for translating toxicogenomics from one platform to another. Transcriptomic profiling has evolved through multiple generations of technology, from microarrays (e.g., Affymetrix, CodeLink) to more recent high-throughput sequencing and targeted panels such as S1500+. Microarrays, which dominated gene expression studies in the early 2000s, provided affordable and high-throughput transcript quantification but suffered from cross-hybridization issues and limited dynamic range . RNA-Seq, introduced in the late 2000s, revolutionized transcriptomics by enabling unbiased and comprehensive gene expression analysis, albeit at higher costs and computational demands . Despite advances, many studies rely on historical microarray data, necessitating the translation of legacy data into modern platforms to ensure continuity and comparability. This translation is complicated by factors such as platform-specific probe design, differences in transcript coverage, and batch effects . Existing methods for cross-platform mapping include statistical normalization, machine learning models, and biological anchoring approaches. The ability to translate transcriptomic data between platforms has broad implications, including enhanced meta-analyses, improved toxicological modeling, and better integration of historical datasets with contemporary research. TransPlatformer seeks to contribute to this effort by evaluating translation methodologies and proposing novel strategies to improve cross-platform gene expression harmonization. In this repository there are code examples for TransPlatformer implementation

Cong, Guojing↗

National User Resource for Biological Accelerator Mass Spectrometry

The National User Resource for Biological Accelerator Mass Spectrometry (User Resource) will provide isotopic analysis (primarily radiocarbon or 14C) by accelerator mass spectrometry (AMS) for NIH- funded researchers across the United States and will be the only User Resource of its type in the United States. The User Resource will provide measurement capability and expertise to a research community that requires highly sensitive, quantitative isotope analyses. Since commissioning a new accelerator mass spectrometer in June 2014, we have measured over 4000 samples a year for collaborators and service users. The User Resource will enable us to continue to meet these research needs, as well as provide for new users whose research programs would benefit from AMS as a measurement tool. The User Resource’s forte will be ultra-high sensitivity quantitation of radiocarbon and selected other radioisotopes for research studies where isotopes are required. Radioisotope labeling studies have been and will continue to be an important tool for addressing many complex biomedical science problems. AMS is a specialized and unique type of mass spectrometry that provides absolute quantitation of radiocarbon and other relevant radioisotopes with extreme sensitivity, having limits of detection in real samples on the order of a few attomol/mg of sample at measurement precisions of ~3%. It is the only instrumental method capable of quantifying radioisotope-labeled agents routinely in real-world samples with such precision and sensitivity. The sensitivity of AMS allows for the quantification of radiolabeled metabolites in extremely complex matrices of cells and organisms at very low concentrations and in small samples. AMS allows studies to be conducted without perturbing metabolism leading to more relevant quantification of metabolic rates and pathways. In addition, it enables quantification of pharmacokinetic and metabolic properties of toxicants at environmentally relevant concentrations in model systems as well as the ability to quantify pharmacokinetics and other molecular endpoints directly in humans. Such quantitative assessments can 1) improve risk assessment for toxicants, 2) address safety and efficacy considerations for therapeutic entities, 3) deepen understanding of xenobiotic and intermediary metabolism, 4) help understand the interactions between critical molecular pathways, and 5) improve efforts to model and predict various metabolic and biological states. These capabilities have been applied in a number of areas including research in carcinogenesis, toxicology, nutrition, pharmacology/drug development and basic biological science. As a NIGMS National Resource the National User Resource for Biological Accelerator Mass Spectrometry will help NIH funded scientists achieve a deeper understanding of the etiology of human health concerns by (1) enabling the quantification of pharmacokinetics and other molecular endpoints directly in humans; (2) offering the ability to conduct quantitative studies using biologics such as proteins or lipids, and thereby reducing the amount of radioisotope usage in biomedical labs; and (3) enabling more relevant studies of metabolic pathways in health and disease through the use of much lower, more biologically-relevant, concentrations of metabolic substrates in cells and intact organisms. Such studies support NIGMS’s basic biomedical research areas that contribute to the understanding of fundamental cellular and physiological principles and enable research supported by the Biophysics, Biomedical Technology, and Computational Biosciences (BBCB); Genetics and Molecular, Cellular, and Developmental Biology (GMCDB); Pharmacology, Physiology, Biological Chemistry (PPBC) and Training, Workforce Development, and Diversity (TWD) Divisions.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

National User Resource for Biological Accelerator Mass Spectrometry (Final Report)

The National User Resource for Biological Accelerator Mass Spectrometry (User Resource) will provide isotopic analysis (primarily radiocarbon or 14C) by accelerator mass spectrometry (AMS) for NIH- funded researchers across the United States and will be the only User Resource of its type in the United States. The User Resource will provide measurement capability and expertise to a research community that requires highly sensitive, quantitative isotope analyses. Since commissioning a new accelerator mass spectrometer in June 2014, we have measured over 4000 samples a year for collaborators and service users. The User Resource will enable us to continue to meet these research needs, as well as provide for new users whose research programs would benefit from AMS as a measurement tool. The User Resource’s forte will be ultra-high sensitivity quantitation of radiocarbon and selected other radioisotopes for research studies where isotopes are required. Radioisotope labeling studies have been and will continue to be an important tool for addressing many complex biomedical science problems. AMS is a specialized and unique type of mass spectrometry that provides absolute quantitation of radiocarbon and other relevant radioisotopes with extreme sensitivity, having limits of detection in real samples on the order of a few attomol/mg of sample at measurement precisions of ~3%. It is the only instrumental method capable of quantifying radioisotope-labeled agents routinely in real-world samples with such precision and sensitivity. The sensitivity of AMS allows for the quantification of radiolabeled metabolites in extremely complex matrices of cells and organisms at very low concentrations and in small samples. AMS allows studies to be conducted without perturbing metabolism leading to more relevant quantification of metabolic rates and pathways. In addition, it enables quantification of pharmacokinetic and metabolic properties of toxicants at environmentally relevant concentrations in model systems as well as the ability to quantify pharmacokinetics and other molecular endpoints directly in humans. Such quantitative assessments can 1) improve risk assessment for toxicants, 2) address safety and efficacy considerations for therapeutic entities, 3) deepen understanding of xenobiotic and intermediary metabolism, 4) help understand the interactions between critical molecular pathways, and 5) improve efforts to model and predict various metabolic and biological states. These capabilities have been applied in a number of areas including research in carcinogenesis, toxicology, nutrition, pharmacology/drug development and basic biological science. As a NIGMS National Resource the National User Resource for Biological Accelerator Mass Spectrometry will help NIH funded scientists achieve a deeper understanding of the etiology of human health concerns by (1) enabling the quantification of pharmacokinetics and other molecular endpoints directly in humans; (2) offering the ability to conduct quantitative studies using biologics such as proteins or lipids, and thereby reducing the amount of radioisotope usage in biomedical labs; and (3) enabling more relevant studies of metabolic pathways in health and disease through the use of much lower, more biologically-relevant, concentrations of metabolic substrates in cells and intact organisms. Such studies support NIGMS’s basic biomedical research areas that contribute to the understanding of fundamental cellular and physiological principles and enable research supported by the Biophysics, Biomedical Technology, and Computational Biosciences (BBCB); Genetics and Molecular, Cellular, and Developmental Biology (GMCDB); Pharmacology, Physiology, Biological Chemistry (PPBC) and Training, Workforce Development, and Diversity (TWD) Divisions. Over the next five years, our goals are to: 1. Improve the efficiency of operation for AMS measurements through installation of new interfaces to our AMS systems, technical modifications to improve gas accepting ion source efficiency and upgrading our data analysis software for improved ease of use and data reporting. 2. Increase the accessibility and visibility of ultra-sensitive 14C measurements for the biomedical research community by training of new investigators and expanding our national user base. 3. Provide high throughput, ultra-sensitive 14C analysis for the NIGMS and NIH user community.

47 OTHER INSTRUMENTATION↗

MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration

Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES↗