Search NASASearch

SEARCH · Search NASA

Results for “NLP”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Nanolipoprotein particle (NLP) vaccine confers protection against Yersinia pestis aerosol challenge in a BALB/c mouse model

Introduction: Yersinia pestis is the etiological agent of plague, a disease that remains a concern as demonstrated by recent outbreaks in Madagascar. Infection with Y. pestis results in a rapidly progressing illness that can only be successfully treated with antibiotics given shortly after symptom onset. Live attenuated or whole cell inactivated vaccines confer protection against bubonic plague, but pneumonic plague has been more difficult to prevent. Novel effective subunit vaccine formulations may circumvent some of these shortfalls. Here, we compare the immunogenicity generated by an advanced subunit vaccine (F1V fusion protein) and a nanolipoprotein particle (NLP)-based vaccine. Methods: The NLP, a high-density lipoprotein mimetic, provides a nanoscale delivery platform for recombinant Y. pestis antigens LcrV (V) and F1. BALB/c mice were immunized via subcutaneous injection twice, three or four weeks apart. Four weeks later, splenocytes and sera were collected for immune profiling, and mice were challenged with aerosolized Y. pestis CO92. Results: Both formulations induced a strong IgG response against the F1 and V proteins, along with a robust memory B cell response and a balanced cell-mediated immune response as evidenced by both Th1- and Th2-related cytokines. The NLP-based vaccine induced a stronger cytokine response against F1, V, and F1V proteins relative to the F1V vaccine. As with F1V, the inclusion of Alhydrogel (Alu) in NLP vaccine formulations was critical for enhanced immunogenicity and protective efficacy. Mice that received two doses of F1:V:NLP + Alu and CpG were completely protected from a challenge with approximately eight median lethal doses of aerosolized Y. pestis CO92 and this protection confirmed the well-documented synergy between the F1 and V antigens in context of pneumonic plague. The NLPs have defined regions of polarity that facilitates the incorporation of a wide range of adjuvants and antigens with distinct physicochemical properties and are an excellent candidate platform for the development of multi-antigen vaccines.

F1

Informing Plant Asset Reliability and Availability Through AI-Driven Analysis of Operator Logs

The availability and reliability of nuclear power plant (NPP) structures, systems, and components (SSCs) are critical parameters for NPP safety. Tracking these parameters is necessary but costly and labor-intensive, requiring the collection and evaluation of SSC event data such as shutdowns, startups, and failures. To show how these events are needed for the parameters an example is given: one measure of reliability is based on the number of equipment failure events and the number of run hours (i.e., the time from a startup event to a shutdown event). Here, this work investigates using artificial intelligence (AI) to mine NPP operator log entry texts for SSC event data. Four AI approaches were explored for identifying these events, including natural language processing (NLP) methods, generative AI, generative AI combined with NLP, and topic modeling. A key challenge addressed with all four approaches is the brevity of operator log entries. Among these four a neural network–based NLP method was shown to be the most promising for this application, achieving F1 scores of 86.0% for shutdowns, 92.2% for startups, and 80.4% for failures on a subject-matter-expert-curated dataset from NPP operator logs, compared to a baseline of 66.6% for a random classifier. This shows that NLP methods can perform better than generative AI. Additionally, the NLP methods combined with generative AI were shown to perform better than generative AI alone. Generative AI was most successful at providing the background information for the NLP methods to use. This work demonstrates the potential to use AI to automate parameter collection from NPP operator log entries and other records.

97 - MATHEMATICS AND COMPUTING

Excipient screening by lyophilization provides insights into spray drying formulations for nanoparticle vaccines

Nanoparticles have shown great promise as delivery platforms in the development of tunable and safe vaccines. Nanolipoprotein particles (NLPs), also known as nanodiscs, are discoidal nanoparticles composed of a lipid bilayer stabilized at their periphery by apolipoproteins. Under the right conditions, the NLP self-assembly process is highly customizable in terms of lipids and apolipoprotein constituents, allowing for tunable physical and chemical characteristics. This flexibility allows a wide range of vaccine antigens and adjuvants to be incorporated onto the NLP platform for tailored vaccine design. The stability of NLPs during long term storage is a very important factor in developing a vaccine delivery platform suitable for widespread global use. When stored in a solution for extended periods of time, NLPs dissociate into their corresponding lipids and protein constituents, leading to particle degradation. Proper stabilization of NLPs can often be achieved by lyophilization (i.e. freeze-drying), a method widely used for various applications including pharmaceuticals. This process, however, can be damaging to particles without the presence of lyoprotectants or excipients that help maintain particle stability during lyophilization. Another method used to stabilize vaccines and pharmaceuticals is spray drying, a process that converts liquid formulations into dry powders through controlled heating and airflow. While spray drying is rapid, scalable, and cost-effective, lyophilization is typically a gentler process that better retains biomolecule structure and function. Both processes use excipients for particle stabilization, so lyophilization can be used as a surrogate to test stability of NLPs, to down-select formulations that may withstand the harsher conditions of spray drying. To screen different formulations, NLPs were synthesized and purified to homogeneity by size exclusion chromatography (SEC) and samples were prepared with a wide range of lyoprotectants and/or excipients. To assess the protective effects of excipients on NLPs upon spray drying, both pre- and post-lyophilized samples were analyzed by SEC. To assess the protective effects upon heating (encountered during the spray drying process), NLP samples were incubated at elevated temperatures prior to SEC analysis. The lyoprotectants and excipients evaluated in this study had different efficiencies in protecting NLPs during lyophilization and heating tests. Trehalose, for example, exhibits stabilization on NLPs both upon lyophilization and heating whereas leucine accelerated NLP dissociation. Although some lyoprotectants are effective by themselves, different combinations can decrease the stabilization of NLPs. Shelf-stable vaccines that do not require cold-chain storage are essential for global accessibility and our findings provide fundamental insight into how to advance NLP-based vaccines for these applications.

Serrano, Litzay J [Lawrence Livermore National Lab

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records

Language models for materials discovery and sustainability: Progress, challenges, and opportunities

Significant advancements have been made in one of the most critical branches of artificial intelligence: natural language processing (NLP). These advancements are exemplified by the remarkable success of OpenAI’s GPT-3.5/4 and the recent release of GPT-4.5, which have sparked a global surge of interest akin to an NLP gold rush. Here, in this article, we offer our perspective on the development and application of NLP and large language models (LLMs) in materials science. We begin by presenting an overview of recent advancements in NLP within the broader scientific landscape, with a particular focus on their relevance to materials science. Next, we examine how NLP can facilitate the understanding and design of novel materials and its potential integration with other methodologies. To highlight key challenges and opportunities, we delve into three specific topics: (i) the limitations of LLMs and their implications for materials science applications, (ii) the creation of a fully automated materials discovery pipeline, and (iii) the potential of GPT-like tools to synthesize existing knowledge and aid in the design of sustainable materials.

36 MATERIALS SCIENCE

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),

Developing a Vaccine Platform for a Balanced Mucosal Immune Response (CB11446), Project Report: Year 1

Mucosal vaccines can elicit protective immune responses at the infection site of respiratory or aerosolized pathogens. Achieving balanced mucosal and systemic responses is a key challenge in the design of mucosal vaccines, especially in terms of developing a broadly applicable vaccine delivery platform amenable to a wide range of subunit vaccine antigens. The overarching goal of the project entitled “Developing a Vaccine Platform for a Balanced Mucosal Immune Response” is to develop a pathogen-agnostic vaccine platform that elicits robust and balanced mucosal immune responses upon intranasal vaccination, in the context of a nanolipoprotein particle (NLP) platform. In year one, we investigated NLP-based vaccine formulations incorporating select immune adjuvants (Monophosphoryl Lipid A (MPLA), FSL-1, L18-muramyl dipeptide (MDP), and cholesterol-tagged ODN2006 (cCpG)) along with antigens relevant to plague (LcrV), tularemia (IglC), and Ebola virus (GP). We demonstrated the ability to prepare, characterize, and quantify these formulations. We then carried out studies in vivo using a BALB/c mouse model, comparing intranasal and intramuscular administration and systematically evaluated the resulting mucosal and systemic immune responses using ELISpot and ELISAs. All proposed tasks (Table 1) were completed within the twelve-month base period, and the established go/no go metric was successfully achieved. Numerous adjuvant:LcrV:NLP formulations were found to have statistically significant increases in IgA titers in serum and/or lung upon IN administration (compared to IM administration), while also producing robust IgG responses in serum. Some formulations also showed significant IFNγ responses using restimulated splenocytes. The data collected during the base period provide valuable insights for the future design of safe and effective mucosal subunit vaccines while also highlighting areas of research that merit further investigation.

59 BASIC BIOLOGICAL SCIENCES

Hybrid Quantum–Classical Graph Transformers for Efficient Sentiment Analysis

Quantum Machine Learning (QML) offers a promising paradigm that leverages quantum computing principles to develop efficient and expressive models for learning from complex and structured data. Recent advances in natural language processing (NLP) and artificial intelligence (AI) have demonstrated capabilities in understanding, generating, and reasoning over linguistic and multimodal information. In this work, we present the Quantum Graph Transformer (QGT), a hybrid quantum–classical architecture that extends graph transformer capabilities through quantum self-attention. The QGT models variable-length sentences as token graphs, where both the embedding encoding and the self-attention mechanisms are implemented using parameterized quantum circuits (PQCs), enabling efficient contextual learning with significantly fewer trainable parameters. We train QGT using both fully connected and 𝑘 -nearest-neighbor graph structures and evaluate it on five benchmark sentiment-classification datasets. Experimental results show that QGT consistently achieves higher or comparable accuracy to existing quantum NLP models and outperforms a Classical Graph Transformer (CGT) baseline with identical architecture, achieving 29.4 × fewer parameters while requiring 3–5 × fewer samples to reach comparable performance. These findings highlight the potential of graph-based quantum models as scalable and data-efficient architectures for natural language understanding.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Spatio-temporal multivariate cluster evolution analysis for detecting and tracking climate impacts

Recent years have seen a growing concern about climate change and its impacts. While Earth System Models (ESMs) can be invaluable tools for studying the impacts of climate change, the complex coupling processes encoded in ESMs and the large amounts of data produced by these models, together with the high internal variability of the Earth system, can obscure important source-to-impact relationships. Here, this paper presents a novel and efficient unsupervised data-driven approach for detecting statistically-significant impacts and tracing spatio-temporal source-impact pathways in the climate through a unique combination of ideas from anomaly detection, clustering and Natural Language Processing (NLP). Using as an exemplar the 1991 eruption of Mount Pinatubo in the Philippines, we demonstrate that the proposed approach is capable of detecting known post-eruption impacts/events. We additionally describe a methodology for extracting meaningful sequences of post-eruption impacts/events by using NLP to efficiently mine frequent multivariate cluster evolutions, which can be used to confirm or discover the chain of physical processes between a climate source and its impact(s).

Anomaly detection

Ice storage model-predictive control in an office building with PV: scenario, error and sensitivity analysis

Thermal energy storage (TES) can enable more building-sited renewable electricity generation and lower utility bill costs for buildings owners and occupants, especially when there are high demand and variable time-of-use (TOU) charges. A model predictive control (MPC) strategy can offer additional savings over a schedule-based control with added complexity and reliance on forecasts. Here, this study examines savings for medium office buildings with chiller plants in three locations with building-installed solar photovoltaics (PV) to understand the impact of MPC. Control setpoints are fixed by a schedule-based control or optimized by nonlinear MPC. These control setpoints are actuated within EnergyPlus building models to simulate the utility cost of the chiller plant. NLP solutions can be unstable or unrealistic, but our results show that by regularizing the NLP, the solutions can be reasonably followed by the building model. MPC models make simplifications that lead to errors once the controller is participating in and changing the operation of the building. These errors average 9 % across the cases, showing that the most important parts of the system are represented. The no-thermal load costs are computed to show that the optimization can in some cases achieve both the minimum TOU and minimum monthly demand costs by demand management while reducing TOU energy costs by energy arbitrage. The MPC saves 35–66 % in the annual chiller plant operating costs, which is an additional savings above the schedule by 1–33 %. PV and TES are complementary and mostly independent, but a load with PV often results in better performance for the schedule. Our case study and sensitivity analysis show the importance of modeling and optimization for complex rates, but also the circumstances wherein a simpler strategy achieves the same performance with less potential for error.

14 SOLAR ENERGY

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa

Domain-specific text embedding model for accelerator physics

Accelerator physics presents unique challenges for natural language processing (NLP) due to its specialized terminology and complex concepts. A key component in overcoming these challenges is the development of robust text embedding models that transform textual data into dense vector representations, facilitating efficient information retrieval and semantic understanding. In this work, we introduce AccPhysBERT, a sentence embedding model fine-tuned specifically for accelerator physics. Our model demonstrates superior performance across a range of downstream NLP tasks, surpassing existing models in capturing the domain-specific nuances of the field. We further showcase its practical applications, including semantic paper-reviewer matching and integration into retrieval-augmented generation systems, highlighting its potential to enhance information retrieval and knowledge discovery in accelerator physics. Published by the American Physical Society 2025

Hellert, Thorsten (ORCID:0000000227970926)

Data for KETCHUP: Parameterizing of Large-Scale Kinetic Models Using Multiple Datasets with Different Reference States

Repository for Kinetic Estimation Tool Capturing Heterogeneous Datasets Using Pyomo (KETCHUP), a flexible parameter estimation tool that leverages a primal-dual interior-point algorithm to solve a nonlinear programming (NLP) problem that identifies a set of parameters capable of recapitulating the steady-state fluxes and concentrations in wild-type and perturbed metabolic networks. KETCHUP can use K-FIT [2] input files. Example K-FIT input files are located in the K-FIT repository at https://github.com/maranasgroup/K-FIT.

Metabolomics

Comprehensive Database of Environmental Mitigations Extracted from FERC-Licensed Hydropower Projects Using Artificial Intelligence Techniques, 1998-2023

This dataset provides a comprehensive inventory of environmental mitigation measures required by Federal Energy Regulatory Commission (FERC) licensed hydropower facilities from 461 licenses that were issued from 1998 to 2023. These licenses constitute 446 of the 1015 FERC projects that were active at the end of 2023. 17,612 mentions of environmental mitigations were identified and categorized in 128 unique categories. Mitigations were identified using a Natural Language Processing (NLP) approach, specifically with a Bidirectional Encoder Representations from Transformer (BERT) model. Model-derived results were then reviewed and updated by a subject matter expert as needed. This dataset introduces important enhancements to previous efforts to inventory environmental mitigations, such as including associated license text for each mitigation, tracking the number of instances a mitigation was identified within a license, and providing improved location information. These enhancements significantly expand the dataset's utility, offering greater analytical capabilities and ensuring reproducibility. The dataset is downloadable as a zip file containing the metadata and dataset files.

Ruggles, Thomas [Oak Ridge National Laboratory (OR

From Machine Learning to Machine Reasoning: A Model-based Approach to Analyze Equipment Reliability Data

In current nuclear power plants (NPPs) a large amount of condition-based data which can be used to assess and monitor component health and performance. Assessing component health from such data can be performed with a large variety of methods. While the analysis of numeric data can be performed with several methods, the extraction of information from textual data remains a challenge. Currently employed natural language processing (NLP) methods do not really provide quantitative information that might be contained in IRs. In addition, the integration of numeric and textual data to identify possible causal relationships between data elements is still an unresolved challenge. This paper presents an approach to extract information from textual (e.g., incident or maintenance reports) and numeric data that relies on model based system engineer (MBSE) models. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence while semantic analysis is designed to analyze the logic structure of a sentence. An innovative element of our approach is that semantic analysis uses MBSE models to identify links between textual elements. Similarly, numeric data is directly linked to elements of the MBSE models in order to map which functions are being monitored.

97 - MATHEMATICS AND COMPUTING

Next-to-next-to-leading power corrections to unpolarized Semi-Inclusive Deep Inelastic Scattering

Semi-Inclusive Deep Inelastic Scattering (SIDIS) is a key tool for exploring the three-dimensional structure of the nucleon through Transverse Momentum Dependent parton distributions and fragmentation functions. While leading-power contributions to the SIDIS cross-section are well established, next-to-leading power (NLP) corrections of order 1/Q and next-to-next-to-leading power (NNLP) corrections of order 1/Q 2 to the hadronic tensor have only recently begun to be systematically investigated. These corrections are essential for reliable phenomenology and interpretation of modern high-precision data. In recent papers by one of the authors, NNLP corrections to the Drell-Yan process were derived using the rapidity factorization formalism. In the present work, we extend this approach to SIDIS and obtain analytic expressions for the unpolarized structure functions. We derive NNLP corrections that include convolutions of unpolarized distributions, f 1 , with unpolarized fragmentation functions, D 1 , and Boer-Mulders functions, ${h}_1^{\perp }$, with Collins fragmentation functions, ${H}_1^{\perp }$. We compare our results with previous formulations, provide numerical studies, confront our predictions with HERMES and COMPASS measurements, and present predictions for future experiments at Jefferson Lab and the Electron-Ion Collider.

deep inelastic scattering

Domain Shift Analysis in Chest Radiographs Classification in a Veterans Healthcare Administration Population

This study aims to assess the impact of domain shift on chest X-ray classification accuracy and to analyze the influence of ground truth label quality and demographic factors such as age group, sex, and study year. We used a DenseNet121 model pre-trained MIMIC-CXR dataset for deep learning-based multi-label classification using ground truth labels from radiology reports extracted using the CheXpert and CheXbert Labeler. We compared the performance of the 14 chest X-ray labels on the MIMIC-CXR and Veterans Healthcare Administration chest X-ray dataset (VA-CXR). The validation of ground truth and the assessment of multi-label classification performance across various NLP extraction tools revealed that the VA-CXR dataset exhibited lower disagreement rates than the MIMIC-CXR datasets. Additionally, there were notable differences in AUC scores between models utilizing CheXpert and CheXbert. When evaluating multi-label classification performance across different datasets, minimal domain shift was observed in the unseen VA dataset, except for the label “Enlarged Cardiomediastinum.” The subgroup with the most significant variations in multi-label classification performance was study year. These findings underscore the importance of considering domain shift in chest X-ray classification tasks, paying particular attention to the temporality of the exam. Our study reveals the significant impact of domain shift and demographic factors on chest X-ray classification, emphasizing the need for improved transfer learning and robust model development. Addressing these challenges is crucial for advancing medical imaging research and improving patient care.

chest X-ray image classification