Search NASASearch

SEARCH · Search NASA

Results for “BERT”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Traveling Wave Tube (TVT) RF Power Combining Demonstration for use in the Jupiter Icy Moons Orbiter (JIMO)

The Jupiter Icy Moons Orbiter (JIMO) is set to launch between the years 2012 and 2015. It will possibly utilize a nuclear reactor power source and ion engines as it travels to the moons of Jupiter. The nuclear reactor will produce hundreds of kilowatts of power for propulsion, communication and various scientific instruments. Hence, the RF amplification devices aboard will be able to operate at a higher power level and data rate. The initial plan for the communications system is for an output of 1000 watts of RF power, a data rate of at least 10 megabits a second, and a frequency of 32 GHz. A higher data rate would be ideal to fully utilize the instruments aboard JIMO. At NASA Glenn, one of our roles in the JIMO project is to demonstrate RF power combining using multiple traveling wave tubes (TWT). In order for the power of separate TWT s to be combined, the RF output waves from each must be in-phase and have the same amplitude. Since different tubes act differently, we had to characterize each tube using a Network Analyzer. We took frequency sweeps and power sweeps to characterize each tube to ensure that they will behave similarly under the same conditions. The 200 watt Dornier tubes had been optimized to run at a lower power level (120 watts) for their extensive use in the ACTS program, so we also had to experiment with adjusting the voltage settings on several internal components (helix, anode, collector) of the tubes to reach the full 200 watt potential. from the ACTS program. Phase shifters and power attenuators were placed in the waveguide circuit at the inputs to the tubes so that adjustments could be made individually to match them exactly. A magic tee was used to route and combine the amplified electromagnetic RF waves on the tube output side. The demonstration of 200 watts of combined power was successful with efficiencies greater than 90% over a 500 MHz bandwidth. The next step will be to demonstrate the use of three amplifiers using two magic tees by adding a 200 watt Dornier tube to the Varian and Logimetrics combined setup for a total of 400 watts. After that we will use two 200 watt Dorniers for 400 watts and eventually four 200 watt Dornier tubes to demonstrate 800 watts. After demonstrating the success of power combining, we will need to verify the integrity of a modulated signal sent through the combined tubes. The purpose will be to see what effects separating and recombining will have on the modulated signal and also what effect it will have on combining efficiency. A Bit Error Rate (BER) will be determined by a Bit Error Rate Tester (BERT) by comparing the random information it transmits to what it receives back. The process began with two 100 watt tubes, a Varian and a Logimetrics, salvaged

Downey, Joseph A.

Natural Language Processing Analysis of Notices to Airmen for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized.

Natural Language Processing

Natural Language Processing (NLP) Analysis of NOTAMs for Air Traffic Management Optimization

With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized. Video is an mp4 download, with a play time of 9 min 35 secs.

Natural Language Processing

A Brief Introduction to AI/ML Applications of Air Traffic Management Data at NASA Ames

This presentation will give a brief overview of several AI/ML This presentation will give a brief overview of several AI/ML projects that NASA Ames interns are exploring in partnership with NASA Aeronautic Research Institute (NARI) and the FAA. NASA is interested in Natural Language Processing (NLP) of various legacy text and speech data within air traffic management e.g., Notices To Airmen (NOTAMs), Letters of Agreement (LoAs), Standard Operating Procedures (SOPs), and Air Traffic Control Center audio briefings. Since our focus is on applying state of the art AI/ML tools to legacy air traffic management data, we first showcase the different data sources of interest followed by a brief introduction to the techniques and language models used. We present some exciting preliminary results on each topic including both unsupervised learning techniques (e.g., clustering) and other modern language models (e.g., BERT) that help extract useful information from these data sources that are interpretable by both man and machine.

Air Traffic Management

TopiQAL: Topic-aware Question Answering using Scalable Domain-specific Supercomputers

We all have questions. About today's temperature, scores of our favorite baseball team, the Universe, and about vaccine for COVID-19. Life, physical, and natural scientists have been trying to find answers to various topics using scientific methods and experiments, while computer scientists have built language models as a tiny step towards automatically answering all of these questions across domains given a little bit of context. In this paper, we propose an architecture using state-of-the-art Natural Language Processing language models namely Topic Models and Bidirectional Encoder Representations from Transformers (BERT) that can transparently and automatically retrieve articles of relevance to questions across domains, and fetch answers to topical questions related to COVID-19 current and historical medical research literature. We demonstrate the benefits of using domain-specific supercomputers like Tensor Processing Units (TPUs), residing on cloud-based infrastructure, using which we could achieve significant gains in training and inference times, also with very minimal cost.

Penberthy, Scott

What Went Wrong: A Survey of Wildfire UAS Mishaps through Named Entity Recognition

Increasingly, unmanned aircraft systems (UAS) are being applied to wildfire incidents for tasks such as mapping, aerial ignition, and delivery. As a result, incident reporting systems for wildfires are beginning to accumulate data related to UAS mishaps in wildfire response. In this research, we apply state-of-the-art natural language processing (NLP) techniques to develop a custom Named Entity Recognition (NER) model which extracts a Failure Modes and Effects Analysis (FMEA)-style survey of wildfire UAS mishaps reported in SAFECOM. The custom NER model is built by fine-tuning an existing (BERT) model, resulting in a generalizable NER model that can extract engineering relevant entities including failure modes, causes, effects, control processes, and recommendations from any failure-relevant text. Similar mishaps are clustered and reported as single rows within the FMEA. For each cluster, frequency, severity, and overall risk are computed. The methodology can be applied as part of a broader safety management system to track trends in mishaps and discover knowledge that can be utilized to improve safety outcomes and system performance.

Machine Learning

What Went Wrong: A Survey of Wildfire UAS Mishaps through Named Entity Recognition

Increasingly, unmanned aircraft systems (UAS) are being applied to wildfire incidents for tasks such as mapping, aerial ignition, and delivery. As a result, aviation incident reporting systems for wildfires are beginning to accumulate data related to UAS mishaps in wildfire response. In this research, we apply state-of-the-art natural language processing (NLP) techniques to develop a custom Named Entity Recognition (NER) model which extracts entities relevant to safety analysts. The custom NER model is built by fine-tuning an existing Bidirectional Encoder Representations from Transformers (BERT) model, resulting in a generalizable NER model that can extract engineering relevant entities including failure modes, causes, effects, control processes, and recommendations from failure-relevant text. This model performs passably, with a weighted average f1 score of 0.33 across entity types, indicating more labeled training data is needed. Extracted entities are used to form a Failure Modes and Effects Analysis (FMEA)-style survey of wildfire UAS mishaps reported using the SAFECOM system. Similar mishaps are manually clustered and reported as single rows within an FMEA. Foreach cluster, we compute frequency, severity, and overall riskin accordance with FAA standards. This methodology can beapplied as part of a broader safety management system totrack trends in mishaps (e.g., likelihood, severity) and discoverknowledge (e.g., causes, effects) that can be utilized to improvesafety outcomes and system performance.

Machine Learning

What Went Wrong: A Survey of Wildfire UAS Mishaps through Named Entity Recognition

Increasingly, unmanned aircraft systems (UAS) are being applied to wildfire incidents for tasks such as mapping, aerial ignition, and delivery. As a result, incident reporting systems for wildfires are beginning to accumulate data related to UAS mishaps in wildfire response. In this research, we apply state-of-the-art natural language processing (NLP) techniques to develop a custom Named Entity Recognition (NER) model which extracts a Failure Modes and Effects Analysis (FMEA)-style survey of wildfire UAS mishaps reported in SAFECOM. The custom NER model is built by fine-tuning an existing (BERT) model, resulting in a generalizable NER model that can extract engineering relevant entities including failure modes, causes, effects, control processes, and recommendations from any failure-relevant text. Similar mishaps are clustered and reported as single rows within the FMEA. For each cluster, frequency, severity, and overall risk are computed. The methodology can be applied as part of a broader safety management system to track trends in mishaps and discover knowledge that can be utilized to improve safety outcomes and system performance.

Machine Learning

Evaluating the Use of Foundational Chemical Language Models in Multimodal Graph Fusion

Rapid and accurate prediction of the physicochemical properties of molecules given their structures remains a key challenge in cheminformatics. Machine learning approaches offer high-throughput options, but the optimality of inductive biases and data representations are up for debate. For example, BERT-based masked language models (MLMs) can be trained in a self-supervised way on hundreds of millions to billions of readily available SMILES strings. Another option is graph neural networks (GNNs), which can operate directly on molecular structures. Yet, generating accurate molecular geometry is computationally expensive, leading to a relative scarcity in data compared to SMILES strings. It is attractive to combine these two paradigms by pre-training an LM on a large corpus of SMILES strings and embedding these representation into a geometric graph neural network. Despite the promise of such an approach, and contrary to previous studies, we find mixed results with the combination of the LMs and GNNs on several molecule datasets. In particular, we found evidence for improvement on the FreeSolv and QM7 benchmarks, but degraded performance on the ESOL, LIPO and QM9 datasets compared to a GNN baseline.

Francel, Collin [University of Alabama]

>24% screen printed Cu contacted n-TOPCon solar cells with successful implementation of LECO process

In this paper, we report the successful fabrication of >24.0 % efficiency n-TOPCon Si solar cells with screen-printed, fire-through Cu contact to n-TOPCon on the rear side and Ag contacted boron emitter on the front side by implementing optimized firing and LECO conditions. The highest efficiency (24.3%) Cu contacted n-TOPCon cell in this study showed excellent cell performance parameters with V oc >730 mV, J sc of 41.1 mA/cm 2 and FF of 80.8%, resulting in an absolute efficiency gap of 0.2% between Cu-contacted and fully Ag contacted n-TOPCon cells (24.5%). The mini-module fabricated with the Cu contacted n-TOPCon cell showed excellent reliability and durability of open-circuit voltage (V oc ), pseudo fill factor (pFF) and efficiency after prolonged damp-heat tests. Such high efficiency screen printed Cu contacted n-TOPCon cells provide unique opportunity to replace very expensive Ag contact on n-TOPCon with cheaper screen printable Cu metal pastes.

14 SOLAR ENERGY

Review of Technical Photovoltaic Key Performance Indicators and the Importance of Data Quality Routines

Technical key performance indicators (KPIs) are important metrics used to assess and quantitatively summarize various aspects of photovoltaic (PV) systems, including long-term performance, economic viability, and carbon footprint. Herein, a group of experts of the International Energy Agency's Photovoltaic Power Systems Programme Task 13 collect and describ the most important technical KPIs used in the industry. Thereby, a set of best practices for reliably handling PV system data is presented and the impact of data quality and climatic variability on KPI calculation is investigated. Further, the effective use of technical KPIs allows triggering data-driven and informed decisions to optimize PV systems and providing a comprehensive overview of how PV systems operate across different conditions and climates. With the worldwide growth of the PV industry, more companies operate/own PV systems in different regions, where the climatic and seasonal profiles differ. This requires context-aware evaluation of KPIs, or the judicious application of multiple KPIs, to ensure that each asset is evaluated correctly. Beyond that, there is untapped potential in the utilization of KPIs through geospatial mapping and extrapolation of fleet KPIs. This study demonstrates that the uncertainty in KPI estimation is not well understood and depends on data quality, climatic variability, and system configuration.

14 SOLAR ENERGY

How should reproducibility be approached in plastic recycling?

With the growing importance of developing new and improved methodologies for plastic recycling, conducting reproducible research and ensuring that results are transferable across labs are increasingly important. This Voices article reflects on how academia and industry view the path forward for strengthening reproducibility to advance science and enable a circular plastics economy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Consistent performance of large language models in rare disease diagnosis across ten languages and 4917 cases

Background Large language models (LLMs) are increasingly used medicine for diverse applications including differential diagnostic support. The training data used to create LLMs such as the Generative Pretrained Transformer (GPT) predominantly consist of English-language texts, but LLMs could be used across the globe to support diagnostics if language barriers could be overcome. Initial pilot studies on the utility of LLMs for differential diagnosis in languages other than English have shown promise, but a large-scale assessment on the relative performance of these models in a variety of European and non-European languages on a comprehensive corpus of challenging rare-disease cases is lacking. Methods We created 4917 clinical vignettes using structured data captured with Human Phenotype Ontology (HPO) terms with the Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema. These clinical vignettes span a total of 360 distinct genetic diseases with 2525 associated phenotypic features. We used translations of the Human Phenotype Ontology together with language-specific templates to generate prompts in English, Chinese, Czech, Dutch, French, German, Italian, Japanese, Spanish, and Turkish. We applied GPT-4o, version gpt-4o-2024-08-06, and the medically fine-tuned Meditron3-70B to the task of delivering a ranked differential diagnosis using a zero-shot prompt. An ontology-based approach with the Mondo disease ontology was used to map synonyms and to map disease subtypes to clinical diagnoses in order to automate evaluation of LLM responses. Findings For English, GPT-4o placed the correct diagnosis at the first rank 19.9% and within the top-3 ranks 27.0% of the time. In comparison, for the nine non-English languages tested here the correct diagnosis was placed at rank 1 between 16.9% and 20.6%, within top-3 between 25.4% and 28.6% of cases. The Meditron3 model placed the correct diagnosis within the first 3 ranks for 20.9% of cases in English and between 19.9% and 24.0% for the other nine languages. Interpretation The differential diagnostic performance of LLMs across a comprehensive corpus of rare-disease cases was largely consistent across the ten languages tested. This suggests that the utility of LLMs in clinical settings may extend to non-English clinical settings.

Artificial intelligence

Forced flow transient safety analysis of irradiation device with adjustable orifice for research reactor fuel assemblies

The Belgium Reactor 2 (BR2) of the Belgian Nuclear Research Centre (SCK CEN) has several irradiation devices or rigs that are dedicated to the fuel performance and qualification demonstration testing of research reactor fuels. In support of the U.S. High Performance Research Reactor (USHPRR) LEU conversion project, a new flexible irradiation apparatus, MUSTANG-R, has been constructed. SCK CEN has completed the design and safety study, in cooperation with Idaho National Laboratory (INL) and Argonne National Laboratory (ANL), to allow for the irradiation testing of a full-size fuel assembly in a 200 mm diameter channel in the BR2 reactor. The moveable valve is a key design feature of the device and acts like an adjustable orifice enhancing or restricting the flow through a coolant channel inlet located in the BR2 upper plenum. This moveable valve allows the flow through the device to be adjusted prior to each BR2 cycle to obtain the necessary conditions for the fuel qualification test. This ensures accurate and representative thermal-hydraulic conditions of the fuel design are achieved. The device was designed and qualified as passively safe, implying verification by a combination of mechanical and thermal-hydraulic analysis and testing. This includes characterization of the safety margin required for a scenario where the moveable valve is assumed to be erroneously closed during irradiation. A simplified and conservative method is proposed for analyzing the corresponding forced flow transient using a critical heat flux criterion. In conclusion, this allows the required minimum valve opening to be determined for the experiments' design and safety studies.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

CRISPR/Cas9 editing of p-COUMAROYL-CoA:MONOLIGNOL TRANSFERASE 1 in maize alters phenolic metabolism, lignin structure, and lignin-first biomass processing

Valorization of lignocellulosic biomass for sustainable production of high-value chemicals is challenged by the complexity of lignin, a phenolic biopolymer. Beyond the classical lignin monomers derived from p-coumaryl, coniferyl, and sinapyl alcohol, grass lignins incorporate substantial amounts of monolignol p-coumarates that are produced by p-COUMAROYL-CoA:MONOLIGNOL TRANSFERASE (PMT). Here, the CRISPR/Cas9-mediated mutation of ZmPMT1 in maize enabled the design of biomass depleted in p-coumaroylated lignin and enriched in guaiacyl lignin. Lignin-first biorefining of stem biomass from zmpmt1 mutants by reductive catalytic fractionation (RCF) generated a lignin oil depleted in carboxylates and enriched in guaiacyl-derived alcohols, which are desirable substrates for bio-based polyurethane synthesis. Furthermore, the reported lignin engineering in maize is a promising strategy for designing a dual-purpose crop, providing both food and feed, along with a renewable feedstock for the production of plant-based chemicals.

59 BASIC BIOLOGICAL SCIENCES

Protein–Protein Interaction Networks Derived from Classical and Machine Learning-Based Natural Language Processing Tools

The study of protein-protein interactions (PPIs) provides insight into various biological mechanisms, including the binding of antibodies to antigens, enzymes to inhibitors or promoters, and receptors to ligands. Recent studies of PPIs have led to significant biological breakthroughs. For example, the study of PPIs involved in the human:SARS-CoV-2 viral infection mechanism aided in the development of the SARS-CoV-2 vaccines. Though several databases exist for the manual curation of PPI networks, text mining methods have been routinely demonstrated as useful alternatives for newly studied or understudied species where databases are incomplete. Here, the relationship extraction (RE) performance of several open-source classical text processing, machine learning (ML)-based natural language processing (NLP), and large language model (LLM)-based NLP tools were compared. Overall, our results indicated that networks derived from classical methods tend to have high true positive rates at the expense of having overconnected-networks, ML-based NLP methods have lower true positive rates but networks with the closest structures to the target network, and LLM-based NLP methods tend to exist in-between the two other approaches, with variable performances. Finally, the selection of a specific NLP approach should be tied to the needs of a study and text availability, as models varied in performance due to the amount of text provided.

59 BASIC BIOLOGICAL SCIENCES

DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases

A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.

BERT

Molar-Mass-Dependent Partitioning of Polyethylene in Nanopores of Model Catalyst Supports from Small-Angle Neutron Scattering

Heterogeneous catalysis offers opportunities to enhance valorization of plastic waste via chemical recycling through control of the upcycled product distributions. Minimizing low-value light hydrocarbons is desired; however, fundamental insights into how to control selectivity are lacking. Here we use contrast variation with small-angle neutron scattering (SANS), model perdeuterated polyethylenes (dPEs), and a model liquid hydrocracking product (tetradecane) to quantify polymer partitioning within mesoporous silica (SBA-15). Polyethylene concentration within the mesopores is increased relative to the bulk solution, and this partitioning increases as the temperature increases. However, this polyethylene partitioning is maximized when the radius of gyration of the polymer chains is comparable to the SBA-15 pore size (10 nm). An increased partitioning at higher temperatures is attributed to entropically driven adsorption of PE within the mesopores. There is no observed preferential partitioning of hexatriacontane (a model oligomer) within the mesopores at the temperatures examined. Furthermore, these results suggest that pore size could promote the selective partitioning of polymer species into the mesopores by size. For plastic upcycling, pore-size-dependent partitioning should increase the probability for the reaction of long polymers over oligomeric and small-molecule polyolefin depolymerization products.

Adsorption