Search NASASearch

SEARCH · Search NASA

Results for “text extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

TropiRoot 1.0: Database of tropical root characteristics across environments

Tropical ecosystems contain the world's largest biodiversity of vascular plants. Yet, our understanding of tropical functional diversity and its contribution to global diversity patterns is constrained by data availability. This discrepancy underscores an urgent need to bridge data gaps by incorporating comprehensive tropical root data into global datasets. Here, we provide a database of tropical root characteristics. This new database, TropiRoot 1.0, will be instrumental in evaluating an array of hypotheses pertaining to root functional ecology and plant biogeography, both within the tropics and relative to other global biomes. The data compilation was conducted by the TropiRoot Initiative, in partnership with the Fine-Root Ecology Database (FRED) and the Global Root Trait (GRooT) database, Colorado State University (CSU) and the Smithsonian Tropical Research Institute (STRI). Literature search and data extraction were conducted between 2020 and 2024. Literature was identified using Web of Science, Scopus, and complemented using the expert knowledge of members of TropiRoot. To provide broad environmental and geographical distributions, literature searches included root characteristics (traits) across global change drivers, natural gradients, and from different continents. We adopted FRED standardized data columns and streamlined the format to enhance accessibility for data extraction across various user groups. This optimized framework resulted in a smaller, yet comprehensive datasheet. To make the database compatible with other global root trait initiatives, column identification was standardized following the codes provided by FRED. These efforts culminated in data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 include root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology, and root chemistry. This initiative represents a 30% increase in the currently available data for tropical roots in FRED. TropiRoot 1.0 contains root characteristics from 25 different countries, where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data were available, including soil data, these data were either extracted and included in the database or its availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match those reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models. The data are freely available and should be cited when used.

FRED

Data from TropiRoot 1.0 database: tropical root characteristics across environments

TropiRoot 1.0 is a new tropical root database with root characteristics across environment gradients. It has data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 includes root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology and root chemistry. This initiative represents an approximately 30% increase in the currently available data for tropical roots in the Fine Root Ecology Database (FRED). TropiRoot 1.0, contains root characteristics from 25 different countries where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data was available, including soil data, these data was either extracted and included in the database or their availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match the ones reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions, and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models.

54 ENVIRONMENTAL SCIENCES

Information Extraction on an Earth Science Knowledge Graphs with Semantic Parsing

Knowledge graphs are an important tool, both for representing knowledge and for retrieving information. Fundamentally, they are semantic networks that represent entities and relationships in the form of nodes and edges. A large corpus of natural language text can bebroken down into discrete entities and relationships to form a useful knowledge graph. Existing research breaks down text into a subject, object, and verb relationship triple. Although this is a useful first step, it loses much of the original contextual information encoded within the text. Our process uses a novel 7-tuple approach, in which elements of sentences are programmatically parsed into seven categories: initiator, impacted, receiver, beneficiary, result, and context. In this presentation, we show a knowledge graph built using this 7-tupleprocessing of an Earth science corpus. We explain the techniques used to create the graph and analyze its information retrieval capability while assessing the accuracy and limitations of the results.

Carson Davis

Isolating contour information from arbitrary images

Aspects of natural vision (physiological and perceptual) serve as a basis for attempting the development of a general processing scheme for contour extraction. Contour information is assumed to be central to visual recognition skills. While the scheme must be regarded as highly preliminary, initial results do compare favorably with the visual perception of structure. The scheme pays special attention to the construction of a smallest scale circular difference-of-Gaussian (DOG) convolution, calibration of multiscale edge detection thresholds with the visual perception of grayscale boundaries, and contour/texture discrimination methods derived from fundamental assumptions of connectivity and the characteristics of printed text. Contour information is required to fall between a minimum connectivity limit and maximum regional spatial density limit at each scale. Results support the idea that contour information, in images possessing good image quality, is (centered at about 10 cyc/deg and 30 cyc/deg). Further, lower spatial frequency channels appear to play a major role only in contour extraction from images with serious global image defects.

Jobson, Daniel J.

Old Woman Creek Wetland Sediment and Electrochemical Sensor Microbial Community, 2023

We are developing a technique to monitor microbiological activities referred to as zero resistance ammetry, which entails the deployment of graphite electrodes in sediments. Measurement of current between electrodes of contrasting redox regimes and/or predominant terminal electron accepting processes can be used as an indicator of the extents of microbiological activity. We deployed an electrode array at depths of 2 mm, 4 mm, 76 mm, 78 mm, 152 mm, 154 mm, 227 mm, and 229 mm below the wetland sediment water interface in the Old Woman Creek National Estuarine Research Center, Huron, OH, USA (Lat. = 41.380833, Long. = -82.508889). A core was collected from adjacent sediment and subsamples were collected from depth intervals of 0 – 25 mm, 25 – 127 mm, 127 – 128 mm, and below 178 mm. To determine if the microbial communities attached to the electrodes were reflective of the adjacent sediment-associated microbial community, we conducted a 16S rRNA gene-based (V4 region) survey of these respective materials. This data package contains the results of these surveys, including metadata on the depths from which samples were collected (samples.csv), DNA extraction and sequencing information (OWC_DEPTH_AMPLICON_SEQUENCING_METADATA), sequence processing information (OWC_DEPTH_BIOINFORMATIC_METADATA.csv), an operational taxonomic unit (OTU) table (OWC_DEPTH_97OTUS_TABLE.csv), and nucleotide sequences of OTUs (OWC_DEPTH_97OTUS_SEQS.fasta). All files can be opened using a text-editing application. The fasta file is compatible with bioinformatics applications.

54 ENVIRONMENTAL SCIENCES

Blast from the Past: ASDC Curation for NASA Suborbital Legacy Missions to Promote Data Discovery and Accessibility

NASA has an extensive history of conducting suborbital field campaigns to further advances in atmospheric sciences. Beginning with the Chemical Instrument Test and Evaluation (CITE) conducted in 1983-1984, NASA has completed many suborbital campaigns over the past three decades. Since the early 2010s, suborbital missions are typically assigned to a NASA Distributed Active Archive Center (DAAC) prior to the mission for long-term archival and distribution. Efforts are being made by NASA’s Earth Science Data and Information System (ESDIS) Project and the Airborne Data Management Group (ADMG) to assign legacy missions to DAACs for permanent archival and distribution, so that these valuable datasets remain to be available to the scientific community. NASA’s Atmospheric Science Data Center (ASDC) has been named the assigned DAAC for nearly 20 atmospheric composition legacy missions, including missions conducted as part of the Global Tropospheric Experiment (GTE) and expects to be named the assigned DAAC for more of these missions over the next few years. The primary goal of the ASDC is to provide access to the datasets as they are currently formatted to the broad user community and enhance their findability and accessibility. However, data reporting standards have evolved significantly since 1983 and the datasets span a wide variety of file formats, including text, Ames, GTE, and ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), and the amount of metadata and relevant information included in the files also varies greatly and can not be readily extracted without subject matter knowledge. This has caused challenges for the ASDC’s suborbital metadata extraction pipeline in ensuring that accurate and necessary metadata is being provided for the missions by all the ASDC’s existing search mechanisms. To make the data more findable and accessible, the ASDC has begun researching ways to further enhance the datasets, including distributing value-added products (i.e. consistent file format such as ICARTT or netCDF), adding standard names from the ESDIS Standards Coordination Office (ESCO)-approved Atmospheric Composition Variable Standard Names Convention (ACVSNC), and creating outreach materials such as ArcGIS StoryMaps, User Guides, and Micro Articles, providing overviews of the missions and what type of data was collected during the missions. These efforts also help support NASA’s Open-Source Science by enhancing the FAIRness of the legacy data products. This presentation will review the ASDC’s ongoing efforts, progress made, and future plans for legacy missions.

Megan Buzanowicz

Multicriteria Measures to Assess the Sustainability of Diets: A Systematic Review

Abstract Context Assessing the overall sustainability of a diet is a challenging undertaking requiring a holistic approach capable of addressing the multicriteria nature of this concept. Objective The aim was to identify and summarize the multicriteria measures used to assess the sustainability characteristics of diets reported at the individual level by healthy adults. Data Sources Articles were identified via PubMed, Scopus, and Web of Science. The search strategy consisted of key words and MeSH terms, and was concluded in September 2022, covering references in English, Spanish, and Portuguese. Data Extraction This systematic review followed the PRISMA guidelines. The search identified 5663 references, from which 1794 were duplicates. Two reviewers independently screened the titles and abstracts of each of the 3869 records and the full-text of the 144 references selected. Of these, 7 studies met the inclusion criteria. Data Analysis A total of 6 multicriteria measures were identified: 3 different Sustainable Diet Indices, the Quality Environmental Costs of Diet, the Quality Financial Costs of Diet, and the Environmental Impact of Diet. All of these incorporated a health/nutrition dimension, while the environmental and economic dimensions were the second and the third most integrated, respectively. A sociocultural sustainability dimension was included in only 1 of the measures. Conclusion Despite some methodological concerns in the development and validation process of the identified measures, their inclusion is considered indispensable in assessing the transition towards sustainable diets in future studies. Systematic Review Registration PROSPERO registration no. CRD42022358824.

Rei, Mariana (ORCID:0000000189453708)

A Natural Language Understanding Approach for Digitizing Aircraft Ground Taxi Instructions

Advancements in natural language processing (NLP) technologies offer a unique opportunity to furnish aircraft crews, primarily pilots, with digital instructions for taxiing operations. Digital taxi instructions, delivered either as text or graphics, can streamline taxiing procedures, thereby reducing radio congestion, minimizing communication errors, and enhancing aircraft monitoring. Techniques used for natural language understanding (NLU), a subset of NLP focused on machine comprehension of natural language, can extract taxi instructions directly from verbal radio communications. This capability paves the way for implementing a digital taxi communication framework with minimal adjustments to the existing air traffic controller operations. This paper delves into a novel application of NLU: the automated generation of digital taxi instructions from air traffic controller speech. We detail the development of an annotation scheme to represent aircraft ground traffic communications within the US National Airspace System (NAS), employing intent classification (IC) and slot filling (SF) to extract taxi instructions using NLU models. Several neural network models were trained on a dataset annotated with our scheme, achieving notable accuracy and F1 scores. Our research demonstrates the feasibility of using NLU to automatically generate digital taxi instructions, showcasing its potential to streamline the implementation of digital taxi communications.

LSTM

A Natural Language Understanding Approach for Digitizing Aircraft Ground Taxi Instructions

Advancements in natural language processing (NLP) technologies offer a unique opportunity to furnish aircraft crews, primarily pilots, with digital instructions for taxiing operations. Digital taxi instructions, delivered either as text or graphics, can streamline taxiing procedures, thereby reducing radio congestion, minimizing communication errors, and enhancing aircraft monitoring. Techniques used for natural language understanding (NLU), a subset of NLP focused on machine comprehension of natural language, can extract taxi instructions directly from verbal radio communications. This capability paves the way for implementing a digital taxi communication framework with minimal adjustments to the existing air traffic controller operations. This paper delves into a novel application of NLU: the automated generation of digital taxi instructions from air traffic controller speech. We detail the development of an annotation scheme to represent aircraft ground traffic communications within the US National Airspace System (NAS), employing intent classification (IC) and slot filling (SF) to extract taxi instructions using NLU models. Several neural network models were trained on a dataset annotated with our scheme, achieving notable accuracy and 𝐹1 scores. Our research demonstrates the feasibility of using NLU to automatically generate digital taxi instructions, showcasing its potential to streamline the implementation of digital taxi communications.

ATC

Open Source Subtitle Editor Software Study for Section 508 Close Caption Applications

This paper will focus on a specific item within the NASA Electronic Information Accessibility Policy - Multimedia Presentation shall have synchronized caption; thus making information accessible to a person with hearing impairment. This synchronized caption will assist a person with hearing or cognitive disability to access the same information as everyone else. This paper focuses on the research and implementation for CC (subtitle option) support to video multimedia. The goal of this research is identify the best available open-source (free) software to achieve synchronized captions requirement and achieve savings, while meeting the security requirement for Government information integrity and assurance. CC and subtitling are processes that display text within a video to provide additional or interpretive information for those whom may need it or those whom chose it. Closed captions typically show the transcription of the audio portion of a program (video) as it occurs (either verbatim or in its edited form), sometimes including non-speech elements (such as sound effects). The transcript can be provided by a third party source or can be extracted word for word from the video. This feature can be made available for videos in two forms: either Soft-Coded or Hard-Coded. Soft-Coded is the more optional version of CC, where you can chose to turn them on if you want, or you can turn them off. Most of the time, when using the Soft-Coded option, the transcript is also provided to the view along-side the video. This option is subject to compromise, whereas the transcript is merely a text file that can be changed by anyone who has access to it. With this option the integrity of the CC is at the mercy of the user. Hard-Coded CC is a more permanent form of CC. A Hard-Coded CC transcript is embedded within a video, without the option of removal.

Murphy, F. Brandon

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam

Electromagnetic Pain Relief/Blocking: Feasibility Assessment

Context/Background: Astronauts use pharmaceuticals during spaceflight to manage acute and chronic pain, but use of analgesics will have drawbacks for exploration-class missions because the shelf life of these medications is limited, resupply will be curtailed, astronauts may develop tolerance and/or addiction to these medications, and side effects can include impairment of cognitive abilities. Electromagnetic devices have been developed that treat pain terrestrially by affecting neuromodulation–dubbed “electroceuticals”, these devices have varied mechanisms of action that either stimulate or suppress neural activity in the central nervous system or peripheral nerves. Objective/Purpose: The available literature was reviewed and FDA-approved pain treatments (both pharmacological and non-pharmacological), as well as those currently under development, were assessed for their suitability for use in exploration class spaceflight missions. Data Sources: Due to the COVID-19 pandemic and the resulting closure of libraries, data sources were restricted to those available digitally. Online database searches included PubMed, U.S. Patent and Trademark Office, federal grant award databases (National Aeronautics and Space Administration (NASA), Department of Defense (DoD), National Institutes of Health (NIH)), and general internet searches. More than 1,600 records were reviewed in this effort. Study Selection/Eligibility Criteria: Targeted searches included different aspects of pain management. Priority was given to review studies, to cover as much of the available literature as possible in this limited effort. Study Appraisal and Synthesis Methods/Data Extraction and Data Synthesis: The titles of the studies and the awards that were obtained by searching online databases were reviewed and further information was sought for the relevant titles. Abstracts or award summaries were generally available online; for journal abstracts, full text articles were either available online or were requested via interlibrary loan. Results: An overwhelming majority of the literature focuses on the treatment of chronic rather than acute pain because it is assumed that acute pain only rarely fails to resolve and instead transitions into chronic pain when the central nervous system becomes hypersensitized. The available electromagnetic devices marketed for pain treatment have varying levels of invasiveness, use different mechanisms of action, and have demonstrated varying efficacy when evaluated scientifically. A truly noninvasive, highly efficient device is desired for use during spaceflight. One portable, self-contained, FDA-approved device was identified that, from preliminarily assessment, best met these criteria; the device noninvasively applies pulsed shortwave therapy (PSWT) to modify pain signals from peripheral nerves, however, the device has limited battery life and the effects are relatively non-selective in type of neural signal modified. Limitations: This current effort, although extensive, did not identify a comprehensive list of all alternatives for pain treatment. Once the pandemic limitations are lifted, a longer, more thorough effort may find additional options. Conclusions/Implications: The ideal electromagnetic pain treatment device for use on exploration-class spaceflight missions does not yet exist, but it may be available soon. It is not feasible for NASA to develop medical devices due to the schedule constraints for pending exploration-class missions, but adapting a promising device that is already FDA-approved might be an option. Monitoring research that is ongoing at other federal agencies is recommended, and further review of the candidate PSWT device identified in this current effort may be warranted.

Carol Mullenax

Electromagnetic Pain Relief/Blocking: Feasibility Assessment

CONTEXT/BACKGROUND Astronauts use pharmaceuticals during spaceflight to manage acute and chronic pain, but use of analgesics will have drawbacks for exploration-class missions because the shelf life of these medications is limited, resupply will be curtailed, astronauts may develop tolerance and/or addiction to these medications, and side effects can include impairment of cognitive abilities. Electromagnetic devices have been developed that treat pain terrestrially by affecting neuromodulation–dubbed “electroceuticals”, these devices have varied mechanisms of action that either stimulate or suppress neural activity in the central nervous system or peripheral nerves. OBJECTIVE/PURPOSE The available literature was reviewed and FDA-approved pain treatments (both pharmacological and non-pharmacological), as well as those currently under development, were assessed for their suitability for use in exploration class spaceflight missions. DATA SOURCES Due to the COVID-19 pandemic and the resulting closure of libraries, data sources were restricted to those available digitally. Online database searches included PubMed, U.S. Patent and Trademark Office, federal grant award databases (National Aeronautics and Space Administration (NASA), Department of Defense (DoD), National Institutes of Health (NIH)), and general internet searches. More than 1,600 records were reviewed in this effort. STUDY SELECTION/ELIGIBILITY CRITERIA Targeted searches included different aspects of pain management. Priority was given to review studies, to cover as much of the available literature as possible in this limited effort. STUDY APPRAISAL AND SYNTHESIS METHODS/DATA EXTRACTION AND DATA SYNTHESIS The titles of the studies and the awards that were obtained by searching online databases were reviewed and further information was sought for the relevant titles. Abstracts or award summaries were generally available online; for journal abstracts, full text articles were either available online or were requested via interlibrary loan. RESULTS An overwhelming majority of the literature focuses on the treatment of chronic rather than acute pain because it is assumed that acute pain only rarely fails to resolve and instead transitions into chronic pain when the central nervous system becomes hypersensitized. The available electromagnetic devices marketed for pain treatment have varying levels of invasiveness, use different mechanisms of action, and have demonstrated varying efficacy when evaluated scientifically. A truly noninvasive, highly efficient device is desired for use during spaceflight. One portable, self-contained, FDA-approved device was identified that, from preliminarily assessment, best met these criteria; the device noninvasively applies pulsed shortwave therapy (PSWT) to modify pain signals from peripheral nerves, however, the device has limited battery life and the effects are relatively non-selective in type of neural signal modified. LIMITATIONS This current effort, although extensive, did not identify a comprehensive list of all alternatives for pain treatment. Once the pandemic limitations are lifted, a longer, more thorough effort may find additional options. CONCLUSIONS/IMPLICATIONS The ideal electromagnetic pain treatment device for use on exploration-class spaceflight missions does not yet exist, but it may be available soon. It is not feasible for NASA to develop medical devices due to the schedule constraints for pending exploration-class missions, but adapting a promising device that is already FDA-approved might be an option. Monitoring research that is ongoing at other federal agencies is recommended, and further review of the candidate PSWT device identified in this current effort may be warranted.

C A Mullenax

Text-mined dataset of solid-state syntheses with impurity phases using Large Language Model

Solid-state synthesis is widely used to obtain various inorganic materials, such as battery materials and bulk thermoelectrics. Despite its prevalence, the process remains challenging due to the lack of a general theory and well-understood underlying reaction mechanisms. While prior works have successfully extracted structured datasets from literature, they often neglect product phase purity or yield. In this work, we construct a solid-state synthesis dataset consisting of 80,806 syntheses extracted with a large language model (LLM), including 18,869 reactions with impurity phase(s). Our dataset not only validates expected thermodynamic trends for impurity phase formation but also identifies challenging cases where impurity phases emerge even when the target phase is significantly more stable.

Lee, Sanghoon

Decoding substance use disorder severity from clinical notes using a large language model

Substance use disorder (SUD) poses a major concern due to its detrimental effects on health and society. SUD identification and treatment depend on a variety of factors such as severity, co-determinants (e.g., withdrawal symptoms), and social determinants of health. Existing diagnostic coding systems used by insurance providers, like the International Classification of Diseases (ICD-10), lack granularity for certain diagnoses, but American clinicians will add this granularity (as that found within the Diagnostic and Statistical Manual of Mental Disorders classification or DSM-5) as supplemental unstructured text in clinical notes. Traditional natural language processing (NLP) methods face limitations in accurately parsing such diverse clinical language. Large language models (LLMs) offer promise in overcoming these challenges by adapting to diverse language patterns. This study investigates the application of LLMs for extracting severity-related information for various SUD diagnoses from clinical notes. We propose a workflow employing zero-shot learning of LLMs with carefully crafted prompts and post-processing techniques. Through experimentation with Flan-T5, an open-source LLM, we demonstrate its superior recall compared to the rule-based approach. Focusing on 11 categories of SUD diagnoses, we show the effectiveness of LLMs in extracting severity information, contributing to improved risk assessment and treatment planning for SUD patients.

60 APPLIED LIFE SCIENCES

Asteroid Exploration and Exploitation

John S. Lewis is Professor of Planetary Sciences and Co-Director of the Space Engineering Research Center at the University of Arizona. He was previously a Professor of Planetary Sciences at MIT and Visiting Professor at the California Institute of Technology. Most recently, he was a Visiting Professor at Tsinghua University in Beijing for the 2005-2006 academic year. His research interests are related to the application of chemistry to astronomical problems, including the origin of the Solar System, the evolution of planetary atmospheres, the origin of organic matter in planetary environments, the chemical structure and history of icy satellites, the hazards of comet and asteroid bombardment of Earth, and the extraction, processing, and use of the energy and material resources of nearby space. He has served as member or Chairman of a wide variety of NASA and NAS advisory committees and review panels. He has written 17 books, including undergraduate and graduate level texts and popular science books, and has authored over 150 scientific publications.

Lewis, John S.

Simplifying NASA Earth Science Data and Information Access Through Natural Language Processing Based Data Analysis and Visualization

NASA Earth science data collected from satellites, model assimilation, airborne missions, and field campaigns, are large, complex and evolving. Such characteristics pose great challenges for end users (e.g., Earth science and applied science users, students, citizen scientists), particularly for those who are unfamiliar with NASA's EOSDIS and thus unable to access and utilize datasets effectively. For example, a novice user may simply ask: what is the total rainfall for a flooding event in my county yesterday? For an experienced user (e.g., algorithm developer), a question can be: how did my rainfall product perform, compared to ground observations, during a flooding event? Nonetheless, with rapid information technology development such as natural language processing, it is possible to develop simplified Web interfaces and back-end processing components to handle such questions and deliver answers in terms of text, data, or graphic results directly to users.In this presentation, we describe the main challenges for end users with different levels of expertise in accessing and utilizing NASA Earth science data. Surveys reveal that most non-professional users normally do not want to download and handle raw data as well as conduct heavy-duty data processing tasks. Often they just want some simple graphics or data for various purposes. To them, simple and intuitive user interfaces are sufficient because complicated ones can be difficult and time-consuming to learn. Professionals also want such interfaces to answer many questions from datasets. One solution is to develop a natural language based search box like Google and the search results can be text, data, graphics and more. Now the challenge is, with natural language processing, can we design a system to process a scientific question typed in by a user? In this presentation, we describe our plan for such a prototype. The workflow is: 1) extract needed information (e.g., variables, spatial and temporal information, processing methods, etc.) from the input, 2) process the data in the backend, and 3) deliver the results (data or graphics) to the user.

Liu, Zhong

Cosmological constraints from the Planck cluster catalogue with DES shear profiles and Chandra observations

We present cosmological constraints from the Planck PSZ2 cosmological cluster sample, using weak-lensing shear profiles from Dark Energy Survey (DES) data and X-ray observations from the Chandra telescope for the mass calibration. We compute hydrostatic mass estimates for all clusters in the PSZ2 sample with a scaling relation between their Sunyaev-Zeldovich signal and X-ray derived hydrostatic mass, calibrated with the Chandra data. We introduce a method to correct these masses with a hydrostatic mass bias using shear profiles from wide-field galaxy surveys. We simultaneously fit the number counts of the PSZ2 sample and the mass calibration with the DES data, finding $Ω_\text{m}=0.312^{+0.018}_{-0.024}$, $σ_8=0.777\pm 0.024$, $S_8\equiv σ_8 \sqrt{Ω_\text{m} / 0.3}=0.791^{+0.023}_{-0.021}$, and $(1-b)=0.844^{+0.055}_{-0.062}$ for our baseline analysis when combined with BAO data. When considering a hydrostatic mass bias evolving with mass, we find $Ω_\text{m}=0.353^{+0.025}_{-0.031}$, $σ_8=0.751\pm 0.023$, and $S_8=0.814^{+0.019}_{-0.020}$. We verify the robustness of our results by exploring a variety of analysis settings, with a particular focus on the definition of the halo centre used for the extraction of shear profiles. We compare our results with a number of other analyses, in particular two recent analyses of cluster samples obtained from SPT and eROSITA data that share the same mass calibration data set. We find that our results are in overall agreement with most late-time probes, in very mild tension with CMB results (1.6$σ$), and in significant tension with results from eROSITA clusters (2.9$σ$). We confirm that our mass calibration is consistent with the eROSITA analysis by comparing masses for clusters present in both Planck and eROSITA samples, eliminating it as a potential cause of tension.

Aymerich, G. [Orsay, IAS; AIM, Saclay] (ORCID:0009