Search NASASearch

SEARCH · Search NASA

Results for “text extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

BOREAS TE-2 NSA Soil Lab Data

This data set contains the major soil properties of soil samples collected in 1994 at the tower flux sites in the Northern Study Area (NSA). The soil samples were collected by Hugo Veldhuis and his staff from the University of Manitoba. The mineral soil samples were largely analyzed by Barry Goetz, under the supervision of Dr. Harold Rostad at the University of Saskatchewan. The organic soil samples were largely analyzed by Peter Haluschak, under the supervision of Hugo Veldhuis at the Centre for Land and Biological Resources Research in Winnipeg, Manitoba. During the course of field investigation and mapping, selected surface and subsurface soil samples were collected for laboratory analysis. These samples were used as benchmark references for specific soil attributes in general soil characterization. Detailed soil sampling, description, and laboratory analysis were performed on selected modal soils to provide examples of common soil physical and chemical characteristics in the study area. The soil properties that were determined include soil horizon; dry soil color; pH; bulk density; total, organic, and inorganic carbon; electric conductivity; cation exchange capacity; exchangeable sodium, potassium, calcium, magnesium, and hydrogen; water content at 0.01, 0.033, and 1.5 MPascals; nitrogen; phosphorus: particle size distribution; texture; pH of the mineral soil and of the organic soil; extractable acid; and sulfur. These data are stored in ASCII text files. The data files are available on a CD-ROM (see document number 20010000884), or from the Oak Ridge National Laboratory (ORNL) Distributed Active Archive Center (DAAC).

Veldhuis, Hugo

Strangeness enhancement at its extremes: multiple (multi-)strange hadron production in pp collisions at \(\sqrt{s}=5.02\) TeV

The probability to observe a specific number of strange and multi-strange hadrons (nS), denoted as P(nS), is measured by ALICE at midrapidity (|y| < 0.5) in $$\sqrt{s}=5.02$$ TeV proton-proton (pp) collisions, dividing events into several multiplicity-density classes. Exploiting, for the first time, a technique based on counting the number of strange-particle candidates event-by-event, this measurement allows one to extend the study of strangeness production beyond the mean of the distribution. This constitutes a new test bench for production mechanisms, probing events with a large imbalance between strange and non-strange content. The analysis of a large-statistics data sample makes it possible to extract P(nS) up to a maximum nS of 7 for $${\text{K}}_{\text{S}}^{0}$$, 5 for Λ and $$\overline{\Lambda }$$, 4 for Ξ− and $${\overline{\Xi } }^{+}$$, and 2 for Ω− and $${\overline{\Omega } }^{+}$$. From this, the probability of producing strange hadron multiplets per event is calculated, thereby enabling the extension of the study of strangeness enhancement to extreme situations where several strange quarks hadronize in a single event at midrapidity. Moreover, comparing hadron combinations with different u and d quark compositions and equal overall s quark content, the contribution to the enhancement pattern coming from non-strangeness related mechanisms is isolated. The results are compared with state-of-the-art phenomenological models implemented in commonly used Monte Carlo event generators, including PYTHIA 8 Monash 2013, PYTHIA 8 with QCD-based Color Reconnection and Rope Hadronization (QCD-CR + Ropes), and EPOS LHC, which incorporates both partonic interactions and hydrodynamic evolution. These comparisons show that the new approach dramatically enhances the sensitivity to the different underlying physics mechanisms modeled by each generator.

Abualrob, I J

BOREAS TE-1 SSA Soil Lab Data

This data set was collected by TE-1 to provide a set of soil properties for BOREAS investigators in the SSA. The soil samples were collected at sets of soil pits in 1993 and 1994. Each set of soil pits was in the vicinity of one of the five flux towers in the BOREAS SSA. The collected soil samples were sent to a lab, where the major soil properties were determined. These properties include, but are not limited to, soil horizon; dry soil color; pH; bulk density; total, organic, and inorganic carbon; electric conductivity; cation exchange capacity; exchangeable sodium, potassium, calcium, magnesium, and hydrogen; water content at 0.01, 0.033, and 1.5 MPascals; nitrogen; phosphorus; particle size distribution; texture; pH of the mineral soil and of the organic soil; extractable acid; and sulfur. The data are stored in tabular ASCII text files. The data files are available on a CD-ROM (see document number 20010000884), or from the Oak Ridge National Laboratory (ORNL) Distributed Active Archive Center (DAAC).

Hall, Forrest G.

Towards Automated Analytics of Research Publications

For readers of scientific publications it remains a big challenge to unambiguously relate the published research with the data used. To a substantial degree it is attributed to authors, journals, editors, and reviewers not prioritizing correct data citation, which impacts traceability, repeatability, and giving credits to published authors and their funding sources. Furthermore, uniform classification of the content of the published research is hampered by journals using journal specific topics and letting authors to assign free text keywords to their papers. We demonstrate automated analytics methods for extracting and relating datasets used and the research application areas by processing 1,300 research papers that referenced the NASA Giovanni service (but probably not the datasets in particular) as supporting their publication process. This presentation was given during the 2022 ESIP January meeting held virtually in January 2022.

Irina Gerasimov

Semantic Theme Analysis of Pilot Incident Reports

Pilots report accidents or incidents during take-off, on flight and landing to airline authorities and Federal aviation authority as well. The description of pilot reports for an incident contains technical terms related to Flight instruments and operations. Normal text mining approaches collect keywords from text documents and relate them among documents that are stored in database. Present approach will extract specific theme analysis of incident reports and semantically relate hierarchy of terms assigning weights of themes. Once the theme extraction has been performed for a given document, a unique key can be assigned to that document to cross linking the documents. Semantic linking will be used to categorize the documents based on specific rules that can help an end-user to analyze certain types of accidents. This presentation outlines the architecture of text mining for pilot incident reports for autonomous categorization of pilot incident reports using semantic theme analysis.

Thirumalainambi, Rajkumar

Optimizing Geospatial Assessments for Nuclear Safeguards Applications with Large Language Models

A multidisciplinary team at Argonne National Laboratory evaluated the ability of large language models (LLMs) to identify geographic locations from open-source text and assessed post-processing measures to strengthen the reliability of those extractions in support of international nuclear safeguards. The study focused on addressing challenges such as toponym ambiguity, imprecise descriptions, and misinformation, which often undermine the accuracy of LLM-derived geospatial assessments. By integrating authoritative geospatial datasets, employing rigorous validation techniques, and leveraging human-in-the-loop processes, the project aimed to enhance the precision, transparency, and reproducibility of geospatial localization workflows. The findings demonstrate that while LLMs exhibit significant potential for accelerating geospatial analysis, their outputs require systematic grounding and verification to ensure reliability in high-stakes applications. This work contributes to the broader field of geospatial intelligence and supports strategic objectives of international organizations such as the International Atomic Energy Agency (IAEA) and the U.S. Department of Energy (DOE).

97 MATHEMATICS AND COMPUTING

Topic Modeling of NASA Space System Problem Reports: Research in Practice

Problem reports at NASA are similar to bug reports: they capture defects found during test, post-launch operational anomalies, and document the investigation and corrective action of the issue. These artifacts are a rich source of lessons learned for NASA, but are expensive to analyze since problem reports are comprised primarily of natural language text. We apply topic modeling to a corpus of NASA problem reports to extract trends in testing and operational failures. We collected 16,669 problem reports from six NASA space flight missions and applied Latent Dirichlet Allocation topic modeling to the document corpus. We analyze the most popular topics within and across missions, and how popular topics changed over the lifetime of a mission. We find that hardware material and flight software issues are common during the integration and testing phase, while ground station software and equipment issues are more common during the operations phase. We identify a number of challenges in topic modeling for trend analysis: 1) that the process of selecting the topic modeling parameters lacks definitive guidance, 2) defining semantically-meaningful topic labels requires nontrivial effort and domain expertise, 3) topic models derived from the combined corpus of the six missions were biased toward the larger missions, and 4) topics must be semantically distinct as well as cohesive to be useful. Nonetheless,topic modeling can identify problem themes within missions and across mission lifetimes, providing useful feedback to engineers and project managers.

Data Mining

Production and Improved Separation of Therapeutic Radionuclides Tb-161, Er-165, and Lu-177

This project resulted in the synthesis and utility of new solid-phase extractants using diglycolamide (DGA) extractants grafted onto mesoporous silica. We used DGA ligands that offer multidentate coordination sites for lanthanides and offer pre-arranged binding sites that may facilitate radiolanthanide metal binding. Also, the Hunter graduate student worked at University of Utah on production and separation of 161 Tb with high specific activity resulting in a publication. These are reported in the full text upload.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records

The flat fielding and achievable signal-to-noise of the MAMA detectors

The Space Telescope Imaging Spectrograph (STIS) was designed to achieve a signal-to-noise (S/N) of at least 100:1 per resolution element. Multi-Anode Microchannel Arrays (MAMA) observations during Servicing Mission Orbital Verification (SMOV) confirm that this specification can be met. From analysis of a single spectrum of GD153, with counting statistics of approximately 165 a S/N of approximately 125 is achieved per spectral resolution element in the far ultraviolet (FUV) over the spectral range of 1280A to 1455A. Co-adding spectra of GRW+7OD5824 to increase the counting statistics to approximately 300 yields a S/N of approximately 190 per spectral resolution element over the region extending from 1347A to 1480A in the FUV. In the near ultraviolet (NUV), a single spectrum of GRW+7OD5824 with counting statistics of approximately 200 yields a S/N of approximately 150 per spectral resolution element over the spectral region extending from 2167 to 2520A. Details of the flat field construction, the spectral extraction, and the definition of a spectral resolution element will be described in the text.

Kaiser, Mary Elizabeth

Creative Analytics of Mission Ops Event Messages

Historically, tremendous effort has been put into processing and displaying mission health and safety telemetry data; and relatively little attention has been paid to extracting information from missions time-tagged event log messages. Todays missions may log tens of thousands of messages per day and the numbers are expected to dramatically increase as satellite fleets and constellations are launched, as security monitoring continues to evolve, and as the overall complexity of ground system operations increases. The logs may contain information about orbital events, scheduled and actual observations, device status and anomalies, when operators were logged on, when commands were resent, when there were data drop outs or system failures, and much much more. When dealing with distributed space missions or operational fleets, it becomes even more important to systematically analyze this data. Several advanced information systems technologies make it appropriate to now develop analytic capabilities which can increase mission situational awareness, reduce mission risk, enable better event-driven automation and cross-mission collaborations, and lead to improved operations strategies: Industry Standard for Log Messages. The Object Management Group (OMG) Space Domain Task Force (SDTF) standards organization is in the process of creating a formal standard for industry for event log messages. The format is based on work at NASA GSFC. Open System Architectures. The DoD, NASA, and others are moving towards common open system architectures for mission ground data systems based on work at NASA GSFC with the full support of the commercial product industry and major integration contractors. Text Analytics. A specific area of data analytics which applies statistical, linguistic, and structural techniques to extract and classify information from textual sources. This presentation describes work now underway at NASA to increase situational awareness through the collection of non-telemetry mission operations information into a common log format and then providing display and analytics tools to provide in-depth assessment of the log contents. The work includes: Common interface formats for acquiring time-tagged text messages Conversion of common files for schedules, orbital events, and stored commands to the common log format Innovative displays to depict thousands of messages on a single display Structured English text queries against the log message data store, extensible to a more mature natural language query capability Goal of speech-to-text and text-to-speech additions to create a personal mission operations assistant to aid on-console operations. A wide variety of planned uses identified by the mission operations teams will be discussed.

events

Computer-Aided System Engineering and Analysis (CASE/A) Programmer's Manual, Version 5.0

The Computer Aided System Engineering and Analysis (CASE/A) Version 5.0 Programmer's Manual provides the programmer and user with information regarding the internal structure of the CASE/A 5.0 software system. CASE/A 5.0 is a trade study tool that provides modeling/simulation capabilities for analyzing environmental control and life support systems and active thermal control systems. CASE/A has been successfully used in studies such as the evaluation of carbon dioxide removal in the space station. CASE/A modeling provides a graphical and command-driven interface for the user. This interface allows the user to construct a model by placing equipment components in a graphical layout of the system hardware, then connect the components via flow streams and define their operating parameters. Once the equipment is placed, the simulation time and other control parameters can be set to run the simulation based on the model constructed. After completion of the simulation, graphical plots or text files can be obtained for evaluation of the simulation results over time. Additionally, users have the capability to control the simulation and extract information at various times in the simulation (e.g., control equipment operating parameters over the simulation time or extract plot data) by using "User Operations (OPS) Code." This OPS code is written in FORTRAN with a canned set of utility subroutines for performing common tasks. CASE/A version 5.0 software runs under the VAX VMS(Trademark) environment. It utilizes the Tektronics 4014(Trademark) graphics display system and the VTIOO(Trademark) text manipulation/display system.

Knox, J. C.

A Solid State Zwitterionic Plastic Crystal with High Static Dielectric Constant

The dielectric data in Figure 3, Figure 4, Figure S6 of the published paper was extracted from 2EOIMTSA-BDS-DATA .txt file. This file can be directly opened using a text file editor. It can also be imported to Excel/ Origin for further plotting and analysis. The G' and G'' in Figure 3 of the publihsed paper was plotted from data in file 2EOImTSA-temperature-sweep.xlsx. This file can be directly opend using Excel. The details of DFT simulations mentioned in Figure 2, Figure 7, and Figure S9 of the published paper are included in the DFT.zip file.

Huang, Zitan [Pennsylvania State University]

Fe(III) reducing bacterial activities in Old Woman Creek wetland sediments, June 2023

To evaluate the Fe(III) reducing microbiological activities in Old Woman Creek Nature Preserve (OWC) wetland sediments, we incubated OWC sediments under anoxic and oxic conditions and with or without Fe(III) amendment [as hydrous ferric oxide (HFO)]. No Fe(III) reduction was observed in heat-deactivated incubations. In non-sterile anoxic incubations, measurement of 0.5 M HCl-extractable Fe(II) indicated that Fe(III) reduction occurred in both Fe(III)-amended and -unamended incubations, indicating that abundant Fe(III) is associated with the OWC sediments. Little Fe(II) accumulated in solution, indicating that the most biogenic Fe(II) adsorbs to the sediments. When air was added to the headspace of non-sterile incubations, Fe(III) reduction was halted and any biogenic Fe(II) that accumulated was oxidized. These experiments were used to guide preparation and analyses of incubations to determine if electrochemical measuements can be used to detect microbiological activities in contrasting terminal electron accepting regimes (i.e., aerobic and Fe(III) reducing conditions). Data package includes methods and data from experiments, including dissolved anion concentrations, dissolved Fe(II) concentrations, and 0.5 M HCl-extractable Fe(II) concentrations. All files are either .txt or .csv and can be opened by any plain text editor application.

EARTH SCIENCE

What Went Wrong: A Survey of Wildfire UAS Mishaps through Named Entity Recognition

Increasingly, unmanned aircraft systems (UAS) are being applied to wildfire incidents for tasks such as mapping, aerial ignition, and delivery. As a result, incident reporting systems for wildfires are beginning to accumulate data related to UAS mishaps in wildfire response. In this research, we apply state-of-the-art natural language processing (NLP) techniques to develop a custom Named Entity Recognition (NER) model which extracts a Failure Modes and Effects Analysis (FMEA)-style survey of wildfire UAS mishaps reported in SAFECOM. The custom NER model is built by fine-tuning an existing (BERT) model, resulting in a generalizable NER model that can extract engineering relevant entities including failure modes, causes, effects, control processes, and recommendations from any failure-relevant text. Similar mishaps are clustered and reported as single rows within the FMEA. For each cluster, frequency, severity, and overall risk are computed. The methodology can be applied as part of a broader safety management system to track trends in mishaps and discover knowledge that can be utilized to improve safety outcomes and system performance.

Machine Learning

What Went Wrong: A Survey of Wildfire UAS Mishaps through Named Entity Recognition

Increasingly, unmanned aircraft systems (UAS) are being applied to wildfire incidents for tasks such as mapping, aerial ignition, and delivery. As a result, incident reporting systems for wildfires are beginning to accumulate data related to UAS mishaps in wildfire response. In this research, we apply state-of-the-art natural language processing (NLP) techniques to develop a custom Named Entity Recognition (NER) model which extracts a Failure Modes and Effects Analysis (FMEA)-style survey of wildfire UAS mishaps reported in SAFECOM. The custom NER model is built by fine-tuning an existing (BERT) model, resulting in a generalizable NER model that can extract engineering relevant entities including failure modes, causes, effects, control processes, and recommendations from any failure-relevant text. Similar mishaps are clustered and reported as single rows within the FMEA. For each cluster, frequency, severity, and overall risk are computed. The methodology can be applied as part of a broader safety management system to track trends in mishaps and discover knowledge that can be utilized to improve safety outcomes and system performance.

Machine Learning

Collecting and Processing Earth Science Data Metrics at NASA ESDIS

Since the launch of Terra satellite in 1999, the number of Earth Science remote sensing data products created and distributed by NASA's Earth Observing System (EOS) Data and Information System (EOSDIS) has increased from a few hundred to nearly ten thousand. NASA's Earth Science Data and Information System (ESDIS) Metrics System (EMS) collects metrics on data ingest, archive, and distribution by its Distributed Active Archive Centers (DAACs) and the Science Investigator-led Systems (SIPS), known as Data Providers. These metrics are critical in helping NASA management as well as data producers in resource planning and gaining a wide range of knowledge of data users and data usage.EMS receives flat files, or log files of data archive, ingest, and distribution either in their raw format, such as Apache web logs, or text files of log records formatted by the Data Providers. Tens of millions of records are processed each day to extract metrics on data products, user information, distribution protocols and services, and so on. The metrics are then made available to designated parties.This presentation provides an overview of the EMS processing workflow and improvement efforts made in recent years to handle ever-increasing number of data records and new metrics requirements, discusses several key steps including mapping log records to data products and identifying user communities along with geo-distribution, and demonstrates typical metrics capabilities produced by the EMS system. Challenges and potential approaches to improve the system are also discussed.

Pan, Jianfu

Transcribing Air Traffic Control System Command Center Planning Telecons Using Cloud-Based Automatic Speech Recognition

This paper addresses the challenge of using Automatic Speech Recognition (ASR) technology to transcribe regular teleconferences that happen between FAA Air Traffic Control System Command Center (ATCSCC) planners, stakeholders and air users. These planning teleconferences (aka telecons or planning webinars) are an integral part of managing air traffic in the U.S. National Airspace System (NAS). In particular, the meetings facilitate the creation and modification of various traffic management initiatives (TMIs), that are used to regulate the flow of air traffic. This is typically a human intensive process, requiring specialists to listen to the entire meeting audio (10-20 minutes duration) and inferring the state of the NAS (e.g., weather phenomenon) that was discussed. It would be advantageous to have digital transcripts of the audio and have useful information (e.g., related to TMIs) automatically extracted from the transcripts. In this regard, we are exploring the adoption of state-of-the-art speech to text and Natural Language Processing (NLP) tools that will achieve our objective of digitizing the webinar audio. Unfortunately, the highly technical phraseology present in the audio and limited data availability for model building make ASR difficult. To overcome this challenge, we have taken the critical first step in creating a human transcription dataset from ~20 hours of speech in the ATCSCC audio with the help of subject matter experts. A novelty of our work is the creation of a ground truth transcription dataset for ATCSCC teleconference webinars, which is particularly important for Aviation domain-specific NLP tasks. Using Microsoft Speech Studio, a cloud-based ASR platform, we have fine-tuned the English pre-trained ASR models (available in speech studio) and achieved an average word error rate (WER) of 6.81%. The baseline ASR also provides a digital version of each planning webinar, making it accessible and text-searchable for future references. Additionally, the transcriptions can serve as a bridge between raw audio data and a range of text-based NLP tasks, such as named entity recognition (NER) and intent classification, potentially enhancing the digital footprint of the webinars and other connected data sources. Our work has several potential applications. Firstly, the transcriptions can be analyzed to understand the complex decision process of creating, implementing and modifying TMIs and may also contribute to TMI prediction services. Secondly, our dataset and model can be used to develop more accurate ASR systems for aviation-specific language, which can bring about digital communication in the aviation industry (and aid current “voice only” communications, which are inherently error-prone). Lastly, the transcriptions themselves can be used as a valuable resource for training other NLP models.

Stephen S. B. Clarke