Search NASA⌕ Search

SEARCH · Search NASA

Results for “Pathology reports”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Path-BigBird: An AI-Driven Transformer Approach to Classification of Cancer Pathology Reports

PURPOSE Surgical pathology reports are critical for cancer diagnosis and management. To accurately extract information about tumor characteristics from pathology reports in near real time, we explore the impact of using domain-specific transformer models that understand cancer pathology reports. METHODS We built a pathology transformer model, Path-BigBird, by using 2.7 million pathology reports from six SEER cancer registries. We then compare different variations of Path-BigBird with two less computationally intensive methods: Hierarchical Self-Attention Network (HiSAN) classification model and an offthe-shelf clinical transformer model (Clinical BigBird). We use five pathology information extraction tasks for evaluation: site, subsite, laterality, histology, and behavior. Model performance is evaluated by using macro and micro F 1 scores. RESULTS We found that Path-BigBird and Clinical BigBird outperformed the HiSAN in all tasks. Clinical BigBird performed better on the site and laterality tasks. Versions of the Path-BigBird model performed best on the two most difficult tasks: subsite (micro F 1 score of 72.53, macro F 1 score of 35.76) and histology (micro F 1 score of 80.96, macro F 1 score of 37.94). The largest performance gains over the HiSAN model were for histology, for which a Path-BigBird model increased the micro F 1 score by 1.44 points and the macro F 1 score by 3.55 points. Overall, the results suggest that a Path-BigBird model with a vocabulary derived from wellcurated and deidentified data is the best-performing model. CONCLUSION The Path-BigBird pathology transformer model improves automated information extraction from pathology reports. Although Path-BigBird outperforms Clinical BigBird and HiSAN, these less computationally expensive models still have utility when resources are constrained.

60 APPLIED LIFE SCIENCES↗

Large-scale deep learning for metastasis detection in pathology reports

Objectives No existing algorithm can reliably identify metastasis from pathology reports across multiple cancer types and the entire US population. In this study, we develop a deep learning model that automatically detects patients with metastatic cancer by using pathology reports from many laboratories and of multiple cancer types. Materials and Methods We use 60 471 unstructured pathology reports from 4 Surveillance, Epidemiology, and End Results (SEER) registries. The reports were coded into 1 of 3 labels: metastasis negative, metastases positive, or metastasis undetermined. We utilize a task-specific deep neural network trained from scratch and compare its performance with a widely used large language model (LLM). Results Our deep learning architecture trained on task-specific data outperforms a general-purpose LLM, with a recall of 0.894 compared to 0.824. We quantified model uncertainty and used it to defer reports for human review. We found that retaining 72.9% of reports increased recall from 0.894 to 0.969. Discussion A smaller deep learning architecture trained on task-specific data outperforms a general LLM. Equally critical to model performance is the incorporation of uncertainty quantification, achieved here through an abstention mechanism. Conclusions This study’s finding demonstrate the feasibility of developing algorithms to automatically identify metastatic cancer cases from unstructured pathology reports.

machine learning↗

Evaluating algorithmic bias on biomarker classification of breast cancer pathology reports

Objectives: This work evaluated algorithmic bias in biomarkers classification using electronic pathology reports from female breast cancer cases. Bias was assessed across 5 subgroups: cancer registry, race, Hispanic ethnicity, age at diagnosis, and socioeconomic status. Materials and Methods: We utilized 594 875 electronic pathology reports from 178 121 tumors diagnosed in Kentucky, Louisiana, New Jersey, New Mexico, Seattle, and Utah to train 2 deep-learning algorithms to classify breast cancer patients using their biomarkers test results. We used balanced error rate (BER), demographic parity (DP), equalized odds (EOD), and equal opportunity (EOP) to assess bias. Results: We found differences in predictive accuracy between registries, with the highest accuracy in the registry that contributed the most data (Seattle Registry, BER ratios for all registries >1.25). BER showed no significant algorithmic bias in extracting biomarkers (estrogen receptor, progesterone receptor, human epidermal growth factor receptor 2) for race, Hispanic ethnicity, age at diagnosis, or socioeconomic subgroups (BER ratio <1.25). DP, EOD, and EOP all showed insignificant results. Discussion: We observed significant differences in BER by registry, but no significant bias using the DP, EOD, and EOP metrics for socio-demographic or racial categories. This highlights the importance of employing a diverse set of metrics for a comprehensive evaluation of model fairness. Conclusion: A thorough evaluation of algorithmic biases that may affect equality in clinical care is a critical step before deploying algorithms in the real world. We found little evidence of algorithmic bias in our biomarker classification tool. Artificial intelligence tools to expedite information extraction from clinical records could accelerate clinical trial matching and improve care.

60 APPLIED LIFE SCIENCES↗

Development of message passing-based graph convolutional networks for classifying cancer pathology reports

Abstract Background Applying graph convolutional networks (GCN) to the classification of free-form natural language texts leveraged by graph-of-words features (TextGCN) was studied and confirmed to be an effective means of describing complex natural language texts. However, the text classification models based on the TextGCN possess weaknesses in terms of memory consumption and model dissemination and distribution. In this paper, we present a fast message passing network (FastMPN), implementing a GCN with message passing architecture that provides versatility and flexibility by allowing trainable node embedding and edge weights, helping the GCN model find the better solution. We applied the FastMPN model to the task of clinical information extraction from cancer pathology reports, extracting the following six properties: main site, subsite, laterality, histology, behavior, and grade. Results We evaluated the clinical task performance of the FastMPN models in terms of micro- and macro-averaged F1 scores. A comparison was performed with the multi-task convolutional neural network (MT-CNN) model. Results show that the FastMPN model is equivalent to or better than the MT-CNN. Conclusions Our implementation revealed that our FastMPN model, which is based on the PyTorch platform, can train a large corpus (667,290 training samples) with 202,373 unique words in less than 3 minutes per epoch using one NVIDIA V100 hardware accelerator. Our experiments demonstrated that using this implementation, the clinical task performance scores of information extraction related to tumors from cancer pathology reports were highly competitive.

59 BASIC BIOLOGICAL SCIENCES↗

Global Explainability of A Deep Abstaining Classifier for Cancer Pathology Reports

We present a global explainability method to characterize sources of errors in a real-world multitask deep abstaining classifier (DAC), in the context of cancer histology prediction. Our multitask classifier, currently deployed for automated annotation of cancer pathology reports from NCI-SEER registries, was trained and evaluated on 1.04 million hand-annotated samples and makes simultaneous predictions of cancer site, subsite, histology, laterality, and behavior for each report. The DAC framework enables the model to abstain on ambiguous reports and confusing classes to achieve the target accuracy on the retained (non-abstained) samples, but at the cost of decreased coverage. Requiring 97% accuracy on the histology task caused our model to retain only 22% of all samples, mostly the less ambiguous and common classes. Local explainability with the GradInp technique provided a computationally efficient way of obtaining contextual reasoning for hundreds of thousands of individual predictions. Our method, involving dimensionality reduction of approximately 13000 aggregated local explanations (ALE), offers a tractable path to true global explainability. It enabled identification of sources of errors in histology classification, globally, as hierarchical complexity among classes, label noise, insufficient information, and conflicting evidence. This suggests several strategies for iterative improvement of our DAC, including well-designed exclusion criteria, focused annotation, and reduced penalties for errors involving hierarchically related classes.

59 BASIC BIOLOGICAL SCIENCES↗

Deep learning uncertainty quantification for clinical text classification

Machine learning algorithms are expected to work side-by-side with humans in decision-making pipelines. Thus, the ability of classifiers to make reliable decisions is of paramount importance. Deep neural networks (DNNs) represent the state-of-the-art models to address real-world classification. Although the strength of activation in DNNs is often correlated with the network’s confidence, in-depth analyses are needed to establish whether they are well calibrated. In this paper, we demonstrate the use of DNN-based classification tools to benefit cancer registries by automating information extraction of disease at diagnosis and at surgery from electronic text pathology reports from the US National Cancer Institute (NCI) Surveillance, Epidemiology, and End Results (SEER) population-based cancer registries. In particular, we introduce multiple methods for selective classification to achieve a target level of accuracy on multiple classification tasks while minimizing the rejection amount—that is, the number of electronic pathology reports for which the model’s predictions are unreliable. We evaluate the proposed methods by comparing our approach with the current in-house deep learning-based abstaining classifier. Overall, all the proposed selective classification methods effectively allow for achieving the targeted level of accuracy or higher in a trade-off analysis aimed to minimize the rejection rate. On in-distribution validation and holdout test data, with all the proposed methods, we achieve on all tasks the required target level of accuracy with a lower rejection rate than the deep abstaining classifier (DAC). Interpreting the results for the out-of-distribution test data is more complex; nevertheless, in this case as well, the rejection rate from the best among the proposed methods achieving 97% accuracy or higher is lower than the rejection rate based on the DAC. We show that although both approaches can flag those samples that should be manually reviewed and labeled by human annotators, the newly proposed methods retain a larger fraction and do so without retraining—thus offering a reduced computational cost compared with the in-house deep learning-based abstaining classifier.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning and deep learning tools for the automated capture of cancer surveillance data

The National Cancer Institute and the Department of Energy strategic partnership applies advanced computing and predictive machine learning and deep learning models to automate the capture of information from unstructured clinical text for inclusion in cancer registries. Applications include extraction of key data elements from pathology reports, determination of whether a pathology or radiology report is related to cancer, extraction of relevant biomarker information, and identification of recurrence. With the growing complexity of cancer diagnosis and treatment, capturing essential information with purely manual methods is increasingly difficult. These new methods for applying advanced computational capabilities to automate data extraction represent an opportunity to close critical information gaps and create a nimble, flexible platform on which new information sources, such as genomics, can be added. This will ultimately provide a deeper understanding of the drivers of cancer and outcomes in the population and increase the timeliness of reporting. These advances will enable better understanding of how real-world patients are treated and the outcomes associated with those treatments in the context of our complex medical and social environment.

60 APPLIED LIFE SCIENCES↗

Deformable phrase level attention: A flexible approach for improving AI based medical coding

Objective: Improving the AI-driven automated medical encoding of clinical text plays a vital role in gathering information on the occurrence of diseases to improve population-level health. This work presents a novel attention mechanism designed to enhance text classification models and ensure appropriate classification of medical concepts in unstructured electronic health records. Materials and Methods: We developed a deformable, phrase-level attention mechanism to identify important lexical word-level and contextual phrase-level information from clinical text documents. We evaluated conventional and transformer-based deep learning models that we extended with our attention mechanism on the extraction of critical cancer information (e.g., site, subsite, laterality, histology, behavior) from 629,908 electronic pathology reports and on the automated medical encoding of 52,722 hospital discharge summaries. Results: Transformer-based models with the deformable, phrase-level attention mechanism achieved the best performance on the extraction of critical cancer information from pathology reports. Conventional- and transformer-based models show similar or better performance than their baseline counterparts on the automated medical encoding of clinical documents. Discussion: The addition of phrase-level information allowed models extended with our proposed method to outperform standard word-level attention. Our method showed favorable properties for the real-world application in terms of model robustness and phenotyping. These results indicate that our method is promising for automated data harmonization for common data models. Conclusion: This work proposes a novel deformable, phrase-level attention mechanism that enhances text classification models in the extraction of medical concepts from clinical text documents. We demonstrate strong performances on two clinical text datasets and showcase real-world deployability of our method.

Automated medical encoding↗

Colonoscopy Screening in the US Astronaut Corps

BACKGROUND: Historically, colonoscopy screenings for astronauts have been conducted to ensure that astronauts are in good health for space missions. Recently this historical data has been identified as being useful for developing an occupational surveillance requirement. It can be used to assess overall colon health and to have a point of reference for future tests in current and former astronauts, as well as to follow-up and track rates of colorectal cancer and polyps. These rates can be compared to military and other terrestrial populations. In 2003, the active astronaut colonoscopy requirements changed to require less frequent colonoscopies. Since polyp removal during a colonoscopy is an intervention that prevents the polyp from potentially developing into cancer, the procedure decreases the individual's risk for colon cancer. The objective of this study is to evaluate the possible effect of increased follow-up times between colonoscopies on the number and severity of polyps identified during the procedures among both current and former NASA astronauts. Initial results and forward work regarding astronaut colonoscopy screenings will be presented. METHODS: A retrospective study of all colonoscopy procedures performed on NASA astronauts between 1962 and 2015 (both during active career and retirement) was conducted by review of the JSC Clinic Electronic Medical Record and Lifetime Surveillance of Astronaut Health (LSAH) database for colonoscopy screening procedures and pathology reports. The timeframe of interest was from the time of selection into the Astronaut Corps through May 2015 or death. For each colonoscopy report, the following data were captured: date of procedure, age at time of procedure, reason for procedure, quality of bowel prep, completion of procedure and/or reason for termination of procedure, findings of procedure, subsequent treatment (if any), recommended follow-up interval, actual follow up interval, family history of polyps or colon cancer, and other significant items or discrepancies. The population consisted of 338 astronauts: 52 females, 286 males. Of these, 56 were deceased, and 11 astronauts had no record of any colonoscopies. Because of a screening requirement change in 2003, analyses were conducted to determine if there were differences between the two time periods. One-sided Wilcoxon rank sum tests were used to identify statistically significant differences between the two time periods. RESULTS: There was a combined total of 1,964 colonoscopy screenings identified. The average follow-up intervals between colonoscopies were indeed longer after the screening requirement change than before the change. The mean follow-up interval pre -2003 was 3.59 years, while post-2003 it increased to 4.35 years. The statistical significance of this difference was confirmed using a one sided Wilcoxon rank sum test which yielded p<.001. Colonoscopies performed after the requirement change tended to have a higher incidence and greater severity of polyps. From pre-2003 to post-2003 the percentage of colonoscopy procedures yielding no polyps decreased from 83.77% to 74.70%. Not only did post-2003 procedures yield more polyp findings, but the polyps recorded were more often of severe pathology. Before 2003 3.62% of colonoscopy findings were polyps of the hyperplastic type (the least severe polyp type) and only 3.35% were of greater severity. Post-2003, 4.21% of findings were hyperplastic polyps while 11.44% were of greater severity. Upon the investigation of other possible contributing factors to these results, we also found that mean age post-2003 was 54.55 years which was significantly higher than during the pre-2003 timeframe (47.32 years). This was observed with a one-sided Wilcoxon rank sum test, resulting in a p<0.001. The increased average age of astronauts could also be a contributing factor to the greater number of polyps found since the risk of developing polyps increases with age. Further work is needed to better understand the increased incidence and greater severity of polyps found in astronaut colonoscopy outcomes.

Masterova, K.↗

Pitfalls in the n -mode representation of vibrational potentials

Simulations of anharmonic vibrational motion rely on computationally expedient representations of the governing potential energy surface. The n-mode representation (n-MR)—effectively a many-body expansion in the space of molecular vibrations—is a general and efficient approach that is often used for this purpose in vibrational self-consistent field (VSCF) calculations and correlated analogues thereof. In the present analysis, a lack of convergence in many VSCF calculations is shown to originate from negative and unbound potentials at truncated orders of the n-MR expansion. For cases of strong anharmonic coupling between modes, the n-MR can both dip below the true global minimum of the potential surface and lead to effective single-mode potentials in VSCF that do not correspond to bound vibrational problems, even for bound total potentials. The present analysis serves mainly as a pathology report of this issue. Furthermore, this insight into the origin of VSCF non-convergence provides a simple, albeit ad hoc, route to correct the problem by “painting in” the full representation of groups of modes that exhibit these negative potentials at little additional computational cost. Somewhat surprisingly, this approach also reasonably approximates the results of the next-higher n-MR order and identifies groups of modes with particularly strong coupling. The method is shown to identify and correct problematic triples of modes—and restore SCF convergence—in two-mode representations of challenging test systems, including the water dimer and trimer, as well as protonated tropine.

Chemistry↗

Multi-Label Classification with Constraint-Based Learning for Hierarchical Consistency

We explore the limitations of traditional crossentropy loss in a hierarchical multi-label classification setting and introduce a novel loss function. This function is designed to integrate hierarchical constraints directly into the training process. By incorporating such constraints into the loss, our approach slightly improves the logical consistency of predictions in structured domains. We demonstrate the efficacy of our approach through experiments on primary site and histology classification by using electronic pathology reports. These results show that our proposed hierarchical loss function enhances the model's ability to produce predictions that are logically consistent with the natural data hierarchies, and it slightly improves predictive accuracy. Our framework may be extended to other hierarchical domains, however the performance gains are context specific.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

Risks Associated with Sharing the MOSSAIC APIs

The MOSSAIC APIs contain two files which, in theory, could be used to discover information about the pathology report data from the SEER registries on which the AI models were trained. In this document, we explain the contents of these files and assess the associated risk. APPENDIX A contains a set of slides to aid in the dissemination of this information.

97 MATHEMATICS AND COMPUTING↗

Color-televised medical microscopy

Color television microscopy used at laboratory range magnifications, reproduces a slide image with sufficient fidelity for medical laboratory and instructional use. The system is used for instant pathological reporting between operating room and remotely located pathologist viewing a biopsy through this medium.

Heath, M. A.↗

An FDA-approved drug structurally and phenotypically corrects the K210del mutation in genetic cardiomyopathy models

Dilated cardiomyopathy (DCM) due to genetic disorders results in decreased myocardial contractility, leading to high morbidity and mortality rates. There are several therapeutic challenges in treating DCM, including poor understanding of the underlying mechanism of impaired myocardial contractility and the difficulty of developing targeted therapies to reverse mutation-specific pathologies. In this report, we focused on K210del, a DCM-causing mutation, due to 3-nucleotide deletion of sarcomeric troponin T (TnnT), resulting in loss of Lysine210. We resolved the crystal structure of the troponin complex carrying the K210del mutation. K210del induced an allosteric shift in the troponin complex resulting in distortion of activation Ca 2+ -binding domain of troponin C (TnnC) at S69, resulting in calcium discoordination. Next, we adopted a structure-based drug repurposing approach to identify bisphosphonate risedronate as a potential structural corrector for the mutant troponin complex. Cocrystallization of risedronate with the mutant troponin complex restored the normal configuration of S69 and calcium coordination. Risedronate normalized force generation in K210del patient-induced pluripotent stem cell–derived (iPSC-derived) cardiomyocytes and improved calcium sensitivity in skinned papillary muscles isolated from K210del mice. Systemic administration of risedronate to K210del mice normalized left ventricular ejection fraction. Collectively, these results identify the structural basis for decreased calcium sensitivity in K210del and highlight structural and phenotypic correction as a potential therapeutic strategy in genetic cardiomyopathies.

Research & Experimental Medicine↗

Inhibition of HSP70 and a Collagen-Specific Molecular Chaperone (HSP47) Expression in Rat Osteoblasts by Microgravity

Rat osteoblasts were cultured aboard a space shuttle for 4 or 5 days. Cells were exposed to 1alpha, 25 dihydroxyvitamin D(3) during the last 20 h and then solubilized by guanidine solution. The mRNA levels for molecular chaperones were analyzed by semi-quantitative RT-PCR. ELISA was used to quantify TGF-beta1 in the conditioned medium. The HSP70 mRNA levels in the flight cultures were almost completely suppressed, as compared to the ground (1 x g) controls. The inducible HSP70 is known as the major heat shock protein that prevents stress-induced apoptosis. The mean mRNA levels for the constitutive HSC73 in the flight cultures were reduced to 69%, approximately 60% of the ground controls. HSC73 is reported to prevent the pathological state that is induced by disruption of microtubule network. The mean HSP47 mRNA levels in the flight cultures were decreased to 50% and 19% of the ground controls on the 4th and 5th days. Concomitantly, the concentration of TGF-beta1 in the conditioned medium of the flight cultures was reduced to 37% and 19% of the ground controls on the 4th and 5th days. HSP47 is the collagen-specific molecular chaperone that controls collagen processing and quality and is regulated by TGF-beta1. Microgravity differentially modulated the expression of molecular chaperones in osteoblasts, which might be involved in induction and/or prevention of osteopenia in space.

Animals↗