Search NASASearch

SEARCH · Search NASA

Results for “text extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Space Communications and Navigation Validation: Extracting Data for the Strategic Center for Networking, Integration, and Communications Scheduling Algorithms

Efficiency in communication system architecture performance between Space Communications and Navigation (SCaN) assets and missions is crucial, as space communication is varied, complex, and often not utilized to its full potential. The SCaN Strategic Center for Networking, Integration, and Communications (SCENIC) new scheduling algorithms, which are designed to simulate the allocation of resources between SCaN assets and missions, have the potential to simulate an increase of this efficiency; however, they require real-world data to be validated against. The purpose of this project was to extract said validation data, which details the frequency and duration of utilized contact windows between missions and assets in the Near Earth Network (NEN), Space Network (SN), and Deep Space Network (DSN). Stored as images in daily operations summaries (DOSs), the tabular data existed in a variety of file formats such as.pdf, .docx, and .doc. Since the tables were stored as images, ABBYY® FineReader® (ABBYY Software Ltd.) optical character recognition (OCR) was implemented, which is a proprietary software that reads images from text. The comma separated value (CSV) output was utilized as input to a series of MATLAB® (The MathWorks, Inc.) methods for reformatting, at which point it was ready to be machine-read. Finally, the results were converted to a Microsoft Excel format for human readability. Along with being used for validation purposes, the data will also be used to map equipment degradation as a function of time to analyze the reliability of network assets.

Kontur, Noah P.

Space-time evolution of particle emission in p–Pb collisions at $\sqrt{{s}_{\text{NN}}} = 5.02$ TeV with 3D kaon femtoscopy

The measurement of three-dimensional femtoscopic correlations between identical charged kaons (K ± K ± ) produced in p–Pb collisions at center-of-mass energy per nucleon pair $\sqrt{{s}_{\text{NN}}}=5.02$ TeV with ALICE at the LHC is presented for the first time. This measurement, supplementary to those in pp and Pb–Pb collisions, allows understanding the particle-production mechanisms at different charged-particle multiplicities and provides information on the dynamics of the source of particles created in p–Pb collisions, for which a general consensus does not yet exist. It is shown that the measured source sizes increase with charged-particle multiplicity and decrease with increasing pair transverse momentum. These trends for K ± K ± are similar to the ones observed earlier in identical charged-pion and ${\text{K}}_{\text{s}}^{0}{\text{K}}_{\text{s}}^{0}$ correlations in Pb–Pb collisions at various energies and in π ± π ± correlations in p–Pb collisions at $\sqrt{{s}_{\text{NN}}}=5.02$ TeV. At comparable multiplicity, the source sizes measured in p–Pb collisions agree within uncertainties with those observed in pp collisions, and there is an indication that they are smaller than those observed in Pb–Pb collisions. The obtained results are also compared with predictions from the hadronic interaction model EPOS 3, which tends to underestimate the source size for the most central collisions and agrees with the data for semicentral and peripheral events. Furthermore, the time of maximal emission for kaons is extracted. It turns out to be comparable with the value obtained in highly peripheral Pb–Pb collisions at the same energy, indicating that the kaon emission evolution is similar to that in p–Pb collisions.

Heavy Ion Experiments

Comprehensive Database of Environmental Mitigations Extracted from FERC-Licensed Hydropower Projects Using Artificial Intelligence Techniques, 1998-2023

This dataset provides a comprehensive inventory of environmental mitigation measures required by Federal Energy Regulatory Commission (FERC) licensed hydropower facilities from 461 licenses that were issued from 1998 to 2023. These licenses constitute 446 of the 1015 FERC projects that were active at the end of 2023. 17,612 mentions of environmental mitigations were identified and categorized in 128 unique categories. Mitigations were identified using a Natural Language Processing (NLP) approach, specifically with a Bidirectional Encoder Representations from Transformer (BERT) model. Model-derived results were then reviewed and updated by a subject matter expert as needed. This dataset introduces important enhancements to previous efforts to inventory environmental mitigations, such as including associated license text for each mitigation, tracking the number of instances a mitigation was identified within a license, and providing improved location information. These enhancements significantly expand the dataset's utility, offering greater analytical capabilities and ensuring reproducibility. The dataset is downloadable as a zip file containing the metadata and dataset files.

Ruggles, Thomas [Oak Ridge National Laboratory (OR

A Brief Introduction to AI/ML Applications of Air Traffic Management Data at NASA Ames

This presentation will give a brief overview of several AI/ML This presentation will give a brief overview of several AI/ML projects that NASA Ames interns are exploring in partnership with NASA Aeronautic Research Institute (NARI) and the FAA. NASA is interested in Natural Language Processing (NLP) of various legacy text and speech data within air traffic management e.g., Notices To Airmen (NOTAMs), Letters of Agreement (LoAs), Standard Operating Procedures (SOPs), and Air Traffic Control Center audio briefings. Since our focus is on applying state of the art AI/ML tools to legacy air traffic management data, we first showcase the different data sources of interest followed by a brief introduction to the techniques and language models used. We present some exciting preliminary results on each topic including both unsupervised learning techniques (e.g., clustering) and other modern language models (e.g., BERT) that help extract useful information from these data sources that are interpretable by both man and machine.

Air Traffic Management

Image Analysis via Fuzzy-Reasoning Approach: Prototype Applications at NASA

A set of imaging techniques based on Fuzzy Reasoning (FR) approach was built for NASA at Kennedy Space Center (KSC) to perform complex real-time visual-related safety prototype tasks, such as detection and tracking of moving Foreign Objects Debris (FOD) during the NASA Space Shuttle liftoff and visual anomaly detection on slidewires used in the emergency egress system for Space Shuttle at the launch pad. The system has also proved its prospective in enhancing X-ray images used to screen hard-covered items leading to a better visualization. The system capability was used as well during the imaging analysis of the Space Shuttle Columbia accident. These FR-based imaging techniques include novel proprietary adaptive image segmentation, image edge extraction, and image enhancement. Probabilistic Neural Network (PNN) scheme available from NeuroShell(TM) Classifier and optimized via Genetic Algorithm (GA) was also used along with this set of novel imaging techniques to add powerful learning and image classification capabilities. Prototype applications built using these techniques have received NASA Space Awards, including a Board Action Award, and are currently being filed for patents by NASA; they are being offered for commercialization through the Research Triangle Institute (RTI), an internationally recognized corporation in scientific research and technology development. Companies from different fields, including security, medical, text digitalization, and aerospace, are currently in the process of licensing these technologies from NASA.

Dominguez, Jesus A.

Geographic Information Systems and Web Page Development

The Facilities Engineering and Architectural Branch is responsible for the design and maintenance of buildings, laboratories, and civil structures. In order to improve efficiency and quality, the FEAB has dedicated itself to establishing a data infrastructure based on Geographic Information Systems, GIs. The value of GIS was explained in an article dating back to 1980 entitled "Need for a Multipurpose Cadastre which stated, "There is a critical need for a better land-information system in the United States to improve land-conveyance procedures, furnish a basis for equitable taxation, and provide much-needed information for resource management and environmental planning." Scientists and engineers both point to GIS as the solution. What is GIS? According to most text books, Geographic Information Systems is a class of software that stores, manages, and analyzes mapable features on, above, or below the surface of the earth. GIS software is basically database management software to the management of spatial data and information. Simply put, Geographic Information Systems manage, analyze, chart, graph, and map spatial information. At the outset, I was given goals and expectations from my branch and from my mentor with regards to the further implementation of GIs. Those goals are as follows: (1) Continue the development of GIS for the underground structures. (2) Extract and export annotated data from AutoCAD drawing files and construct a database (to serve as a prototype for future work). (3) Examine existing underground record drawings to determine existing and non-existing underground tanks. Once this data was collected and analyzed, I set out on the task of creating a user-friendly database that could be assessed by all members of the branch. It was important that the database be built using programs that most employees already possess, ruling out most AutoCAD-based viewers. Therefore, I set out to create an Access database that translated onto the web using Internet Explorer as the foundation. After some programming, it was possible to view AutoCAD files and other GIS-related applications on Internet Explorer, while providing the user with a variety of editing commands and setting options. I was also given the task of launching a divisional website using Macromedia Flash and other web- development programs.

Reynolds, Justin

Search for a new scalar resonance decaying to a Higgs boson and another new scalar particle in the final state with two bottom quarks and two photons in proton-proton collisions at $\sqrt{s} = 13$ TeV

A search is presented for a new scalar resonance, X, decaying to a standard model Higgs boson and another new scalar particle, Y, in the final state where the Higgs boson decays to a $\text{b}\overline{\text{b} }$ pair, while the Y particle decays to a pair of photons. The search is performed in the mass range 240–1000 GeV for the resonance X, and in the mass range 70–800 GeV for the particle Y, using proton-proton collision data collected by the CMS experiment at $\sqrt{s}=13$ TeV, corresponding to an integrated luminosity of 132 fb −1 . In general, the data are found to be compatible with the standard model expectation. Observed (expected) upper limits at 95% confidence level on the product of the production cross section and the relevant branching fraction are extracted for the X → YH process, and are found to be within the range of 0.05–2.69 (0.08–1.94) fb, depending on m X and m Y . The most significant deviation from the background-only hypothesis is observed for X and Y masses of 300 and 77 GeV, respectively, with a local (global) significance of 3.33 (0.65) standard deviations.

Beyond Standard Model

Derivation of physical equations for high-speed laser welding using large language models

It is challenging to formulate complex physical phenomena that occur in a manufacturing process, particularly when the available data are limited, rendering conventional data-driven approaches ineffective. This study aims to predict humping onset in high-speed laser welding by introducing a novel framework, namely text-to-equations generative pre-trained transformer (T2EGPT). This method leverages the capabilities of large language models (LLMs), in combination with sparse experimental data and enriched literature data, to derive an interpretable and generalizable equation for predicting humping initiation. By capturing key correlations among physical parameters, T2EGPT generates a compact and dimensionless expression that accurately predicts hump formation. The equation reveals that humping arises from the interplay between inertia-driven backward melt flow and capillary-driven surface stabilization, where inertial forces drive molten metal backward and capillary forces resist surface deformation. Furthermore, compared to traditional data-driven models, T2EGPT demonstrates enhanced predictive accuracy and cross-material transferability. More broadly, this study highlights the potential of LLMs to integrate textual information with data-driven discovery, enabling the extraction of physical laws in data-scarce scientific domains.

36 MATERIALS SCIENCE

A petabyte size electronic library using the N-Gram memory engine

A model library containing petabytes of data is proposed by Triada, Ltd., Ann Arbor, Michigan. The library uses the newly patented N-Gram Memory Engine (Neurex), for storage, compression, and retrieval. Neurex splits data into two parts: a hierarchical network of associative memories that store 'information' from data and a permutation operator that preserves sequence. Neurex is expected to offer four advantages in mass storage systems. Neurex representations are dense, fully reversible, hence less expensive to store. Neurex becomes exponentially more stable with increasing data flow; thus its contents and the inverting algorithm may be mass produced for low cost distribution. Only a small permutation operator would be recalled from the library to recover data. Neurex may be enhanced to recall patterns using a partial pattern. Neurex nodes are measures of their pattern. Researchers might use nodes in statistical models to avoid costly sorting and counting procedures. Neurex subsumes a theory of learning and memory that the author believes extends information theory. Its first axiom is a symmetry principle: learning creates memory and memory evidences learning. The theory treats an information store that evolves from a null state to stationarity. A Neurex extracts information data without a priori knowledge; i.e., unlike neural networks, neither feedback nor training is required. The model consists of an energetically conservative field of uniformly distributed events with variable spatial and temporal scale, and an observer walking randomly through this field. A bank of band limited transducers (an 'eye'), each transducer in a bank being tuned to a sub-band, outputs signals upon registering events. Output signals are 'observed' by another transducer bank (a mid-brain), except the band limit of the second bank is narrower than the band limit of the first bank. The banks are arrayed as n 'levels' or 'time domains, td.' The banks are the hierarchical network (a cortex) and transducers are (associative) memories. A model Neurex was built and studied. Data were 50 MB to 10 GB samples of text, data base, and images: black/white, grey scale, and high resolution in several spectral bands. Memories at td, S(m(sub td)), were plotted against outputs of memories at td-1. S(m(sub td)) was Boltzman distributed, and memory frequencies exhibited self-organized criticality (SOC); i.e., 'l/f(sup beta)' after long exposures to data. Whereas output signals from level n may be encoded with B(sub output) = O(-log(2)f(sup beta)) bits, and input data encoded with B(sub input) = O((S(td)/S(td-1))(sup n)), B(sup output)/B(sub input) is much less than 1 always, the Neurex determines a canonical code for data and it is a lossless data compressor. Further tests are underway to confirm these results with more data types and larger samples.

Bugajski, Joseph M.

NASA Pilot-Engaged Expert Response Using IBM Watson Technology: Prototype Evaluation of Knowledge Retrieval System

NASA Langley Research Center and IBM have been investigating the use of IBM Watson technology in aerospace research and development. One application of Watson technology is the Pilot-Engaged Expert Response (PEER) use case. The PEER system is envisioned as an in-cockpit advisor that will act as a source of situationally-relevant information for pilots and other flight crew members to assist in decision making about real-time events and situations that arise in the course of aircraft operations. PEER will make available vast stores of knowledge and information quickly and directly, putting important informational resources where they are needed most. IBM has worked with NASA to develop an architecture and articulate a roadmap for the development of the PEER system. That vision is built around Watson Discovery Advisor (WDA) software solution, derived from IBM's Jeopardy!-winning automatic question answering system. PEER makes use of WDA's sophisticated question-answering capabilities as its core, adding important User Interface components and other customizations for the cockpit environment, including communication with flight systems and other external data sources. The development plan for PEER includes four development stages, with the current project constituting the first phase. In this project, a prototype instance of PEER was successfully adapted to the aviation domain, enabling users to ask questions about aviation topics and receive useful and accurate answers to these questions. Major tasks accomplished include the development of procedures for domain adaptation through automatic lexicon extraction from domain glossaries; generation of question-answer training data which was used to train the system; and assessment of the effectiveness of domain adaptation, which showed a dramatic improvement in the ability of the PEER system to answer domain-relevant questions. In addition, the vision for the PEER system was pushed forward by the articulation of a plan for the automatic enhancement of question-answering with contextual information. This initial phase focused on two main goals: 1) the targeted domain adaptation of the underlying WDA system to the aviation domain; and, 2) the design of the software systems needed to leverage flight-contextual data. Domain adaptation of the WDA system proceeds via three main activities: Domain data ingestion, lexical customization and model training. A textual corpus consisting of 1,147 individual documents with more than 7.5 million words of text was ingested into the system and this served as the basis of all further development. A domain lexicon of over 3,500 aviation-domain terms was semi-automatically generated from domain documents and used to train the system. In addition, a set of over 500 question-answer (QA) pairs relevant to the PEER use case was developed; these were used to train and assess the system. These important first steps established the basis for the PEER system. In addition, steps were taken towards the integration of the PEER system into the cockpit environment with the development of a functional design for the Contextual Data Augmentation (CDA) subsystem. This subsystem brings to bear contextual data to improve system responses. It has three main submodules: the Contextual Data Collection module, the Contextual Data Selection module, and the Contextual QA Augmentation module. These modules form a processing pipeline that addresses the problems associated with automatically integrating information from external resources into the knowledge-retrieval mechanism.

Machine learning

A Hybrid Approach to Labeling Datasets in Earth Science Publications

NASA Data Centers provide the public with thousands of datasets that result in published papers, reports, and conference proceedings. Collecting accurate metrics on usage of these datasets is key to connecting different areas of knowledge and evaluating the datasets’ impact. While most of the datasets have Digital Object Identifiers (DOIs) assigned, most publications do not cite them hampering the automated search of these publications. Instead, articles mention attributes like organization, instrument, mission, variable, or a publication describing the dataset. Often only domain experts can deduce the dataset that was used in the publication text. The lack of a citation slows the spread of information and reduces the research’s impact. With thousands of papers produced each year, an automated means of labeling datasets is critical. This paper explores a hybrid approach of heuristics and a Natural Language Processing (NLP) Named Entity Recognition (NER) model to find and label the datasets used within Earth Science papers. Heuristics are used to produce the labelled sentences and any potential dataset candidates that can be derived from a sentence. The heuristic labels the sentences with the names of mission, instrument, re-analysis models, and science keywords taken from the Global Change Master Directory (GCMD) ontology. Additionally, it uses those labels to generate the dataset citation candidates. If the mission, instrument, and variable are sufficient to create the citation for the dataset the citation and the label the domain expert reviews the output without going through the NLP model. If the extracted label is not sufficient to label the dataset on its own, the sentence and its associated dataset labels will be inputted into the NER model. The model outputs the labeled sentence and the potential dataset candidates with their associated probabilities. The domain expert then reviews the NER model’s output and the correct labels are determined. The newly labelled papers can then be used as additional training data. This creates an iterative process for the approach to continuously improve. Because all the possible mentions are gathered by the model, the domain expert can quickly and easily label the papers resulting in large time savings.

Jacob Atkins

NEPATEC2.0: NEPA Text Corpus v2.0

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

environmental review

NEPATEC v2.0: Standardized Metadata and Text Corpus of National Environmental Policy Act Documents

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

54 ENVIRONMENTAL SCIENCES

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics

Towards an Aviation Large Language Model by Fine-tuning and Evaluating Transformers

In the aviation domain, there are many applications for machine learning and artificial intelligence tools that utilize natural language. For example, there is a desire to know the commonalities in written safety reports such as voluntary post incidents reports or aerial wildfire operations reports to better understand the risks present. Another use-case is the possibility of extracting airspace procedures and constraints currently written in documents such as Letters of Agreement. These applications can benefit from the use of state-of-the-art natural language processing techniques when adapted to the language/phraseology specific to the aviation domain. This paper evaluates the viability of adaptation of NLP tools to the aviation domain by fine-tuning transformer based models using aviation data sets. In 2018, a novel language model based on neural units (also called transformers) was created and became known as “Bidirectional Encoder Representations from Transformers” or BERT. This architecture combined with large amounts of English training data and innovative semi-supervised training tasks set the standard for what would later emerge as Large Language Models. The performance of these models was further improved by hyperparameter tuning and refinement of the semi-supervised training task and resulted in “Robustly Optimized BERT Pre-training Approach through hyperparameter tuning” or RoBERTa models. These pre-trained Large Language Models proved to be useful for a wide variety of natural language processing tasks such as text classification and question answering through a process called fine-tuning. The transformer architecture with pre-trained weights served as the basis with the last few layers replaced with layers fine-tuned to perform a new task e.g., a layer that provides a label for the entire input text. This process of fine-tuning can also be used to adapt the models to new domains; e.g., BioBERT started with the pre-trained BERT model and was completed by additional fine-tuning and training on biomedical documents. Transformer-based architectures can also be used to create rich representations of text called embeddings which can serve as the input to other machine learning models. This allows simpler algorithms such as logistic regression to use context-rich representations of the text while still remaining quick to train and evaluate. In the world of aviation, there is a growing demand for natural language processing and understanding but the domain presents unique challenges. Due to the technical content (and specialized language) of most aviation documents, fine-tuning pre-trained Large Language Models to specific tasks has not met the benchmark on natural language processing tasks set by simpler models trained from scratch on the data. To address this deficiency, this paper evaluates the improvements from fine-tuning a Large Language Model on a large set of aviation documents using the original semi-supervised training tasks before performing specific natural language tasks. In fine-tuning, a domain-specific dataset is used on the original training task but with the pre-trained Large Language Model instead of starting from a random initialization. This approach allows the model to be adapted to the specific domain language without discarding the information gained from training on general English data. This paper utilized two major dataset types to train and assess the RoBERTa fine-tuning performance. The first are 7,057 Letters of Agreement which are Federal Aviation Administration (FAA) documents that formalize airspace operations across the national airspace system. They contain many examples of ‘aviation English’ using domain specific terminology and phrasing which serves as a representative basis to perform the semi-supervised fine-tuning. The second type is the 494 document classification labels to be used for evaluation. This down-stream evaluation aims to show the performance of the fine-tuned model, better understand how much data is needed for an effective fine-tuning, and how fine-tuning can be adapted for different applications in-the domain. After semi-supervised training, evaluation begins by encoding the documents for classification using the fine-tuned RoBERTa model. Then a logistic regression classifier is trained to label the document type and compared against our ground truth labels. This currently leads to a 82.8% accuracy on 10-fold cross validation showing improvement over baseline RoBERTa which achieved 81.0%. We plan to measure the improvements on additional tasks and it is expected that these improvements will lead to more robust models that can tackle the natural language processing challenges present in aviation datasets.

ATM

Intra-EVA Space-to-Ground Interactions when Conducting Scientific Fieldwork Under Simulated Mars Mission Constraints

The Biologic Analog Science Associated with Lava Terrains (BASALT) project is a four-year program dedicated to iteratively designing, implementing, and evaluating concepts of operations (ConOps) and supporting capabilities to enable and enhance scientific exploration for future human Mars missions. The BASALT project has incorporated three field deployments during which real (non-simulated) biological and geochemical field science have been conducted at two high-fidelity Mars analog locations under simulated Mars mission conditions, including communication delays and data transmission limitations. BASALT's primary Science objective has been to extract basaltic samples for the purpose of investigating how microbial communities and habitability correlate with the physical and geochemical characteristics of chemically altered basalt environments. Field sites include the active East Rift Zone on the Big Island of Hawai'i, reminiscent of early Mars when basaltic volcanism and interaction with water were widespread, and the dormant eastern Snake River Plain in Idaho, similar to present-day Mars where basaltic volcanism is rare and most evidence for volcano-driven hydrothermal activity is relict. BASALT's primary Science Operations objective has been to investigate exploration ConOps and capabilities that facilitate scientific return during human-robotic exploration under Mars mission constraints. Each field deployment has consisted of ten extravehicular activities (EVAs) on the volcanic flows in which crews of two extravehicular and two intravehicular crewmembers conducted the field science while communicating across time delay and under bandwidth constraints with an Earth-based Mission Support Center (MSC) comprised of expert scientists and operators. Communication latencies of 5 and 15 min one-way light time and low (0.512 Mb/s uplink, 1.54 Mb/s downlink) and high (5.0 Mb/s uplink, 10.0 Mb/s downlink) bandwidth conditions were evaluated. EVA crewmembers communicated with the MSC via voice and text messaging. They also provided scientific instrument data, still imagery, video streams from chest-mounted cameras, GPS location tracking information. The MSC monitored and reviewed incoming data from the field across delay and provided recommendations for pre-sampling and sampling tasks based on their collective expertise. The scientists used dynamic priority ranking lists, referred to as dynamic leaderboards, to track and rank candidate samples relative to one another and against the science objectives for the current EVA and the overall mission. Updates to the dynamic leaderboards throughout the EVA were relayed regularly to the IV crewmembers. The use of these leaderboards enabled the crew to track the dynamic nature of the MSC recommendations and helped minimize crew idle time (defined as time spent waiting for input from Earth during which no other productive tasks are being performed). EVA timelines were strategically designed to enable continuous (delayed) feedback from an Earth-based Science Team while simultaneously minimizing crew idle time. Such timelines are operationally advantageous, reducing transport costs by eliminating the need for crews to return to the same locations on multiple EVAs while still providing opportunities for recommendations from science experts on Earth, and scientifically advantageous by minimizing the potential for cross-contamination across sites. This paper will highlight the space-to-ground interaction results from the three BASALT field deployments, including planned versus actual EVA timeline data, ground assimilation times (defined as the amount of time available to the MSC to provide input to the crew), and idle time. Furthermore, we describe how these results vary under the different communication latency and bandwidth conditions. Together, these data will provide a basis for guiding and prioritizing capability development for future human exploration missions.

Beaton, Kara H.

An Approach to Identifying Aspects of Positive Pilot Behavior within the Aviation Safety Reporting System

The National Airspace System (NAS) is constantly evolving as air traffic continues to ramp up to pre-pandemic numbers and projected to grow to unprecedented levels in the coming years. As well as increasing demand to the current system, emerging operations such as Unmanned Autonomous Systems are also expected to add to complexity in the airspace. To address these issues, the industry and government agencies supporting the NAS will need to rely upon additional automation and new technologies to address future operational requirements, while continuing to be a world-leading safe transportation system. As these new technologies are implemented, the system continues to rely on human pilots and controllers in the loop to monitor the system and intervene in situations the automation cannot handle. The goal of proactively addressing safety is of foremost concern to ensure passenger confidence. The industry has implemented various Safety Monitoring Systems to identify safety risks and proactively address them before they result in a serious incident or accident. One such program is the Aviation Safety Reporting System (ASRS). ASRS is a long-established system where pilots and controllers voluntarily and anonymously report safety incidents they experienced and observed during line operations by providing rich text narratives describing the events, the environment, and conditions leading to the safety event of concern. These narratives provide insight and context around events of interest and can be used to identify emerging problems. They can trigger investigations within Flight Operational Quality Assurance or Flight Data Monitoring programs. However, this process typically focuses on the adverse events and the unsafe aspects of the operations surrounding the reported or detected events. This perspective of investigating factors that went wrong around an adverse event is commonly referred to as Safety I. Alternatively, characterizing successful actions that operators perform every day under varying conditions that keep the system within safe operating bounds is a concept referred to as Safety II. The benefit of the Safety II view is that the scope is much larger than that of Safety I since a vast majority of the operations result in successful flights. Many of the successful techniques used to manage operational threats are not documented in standard operating procedures or taught during training. They are typically acquired over time by working with experienced pilots during line operations or in many cases after experiencing a problem for the first time and reacting to it in situ, drawing from years of experience to manage the threat. In an attempt to quantify these positive actions, we are proposing an approach to extracting key behaviors within ASRS reports that can support the Safety II concept. Our analysis assumes that ASRS reports contain some descriptions of corrective actions that operators performed to prevent a situation from leading to an accident. Leveraging recent advances in Natural Language Process modeling, we have developed an approach to extract positive sentiment from reports, embed these positive statements in a vector space where they can be numerically analyzed, and clustering these statements into similar contextual categories. From these contextualized categories we can attempt to summarized and distilled aspects of the positive behavior. The goal is to identify categories of behavior that describe consistent operator techniques that supports the Safety II concept. With this information, airlines may enable learning from these positive actions, or address procedures that need to be changed to avoid having pilots implement a workaround. These insights can provide a lens into what is “going right” in the operations that may otherwise not be known widely within the community. It is envisioned that this approach can be extended to other narrative programs such as Line Operation Safety Audit or Learning Improvement Team reports where similar observed behavior can be analyzed to extract positive actions and inform the overall operations.

NLP

An Approach to Identifying Aspects of Positive Pilot Behavior within the Aviation Safety Reporting System

The National Airspace System (NAS) is constantly evolving as air traffic continues to ramp up to pre-pandemic numbers and projected to grow to unprecedented levels in the coming years. As well as increasing demand to the current system, emerging operations such as Unmanned Autonomous Systems are also expected to add to complexity in the airspace. To address these issues, the industry and government agencies supporting the NAS will need to rely upon additional automation and new technologies to address future operational requirements, while continuing to be a world-leading safe transportation system. As these new technologies are implemented, the system continues to rely on human pilots and controllers in the loop to monitor the system and intervene in situations the automation cannot handle. The goal of proactively addressing safety is of foremost concern to ensure passenger confidence. The industry has implemented various Safety Monitoring Systems to identify safety risks and proactively address them before they result in a serious incident or accident. One such program is the Aviation Safety Reporting System (ASRS). ASRS is a long-established system where pilots and controllers voluntarily and anonymously report safety incidents they experienced and observed during line operations by providing rich text narratives describing the events, the environment, and conditions leading to the safety event of concern. These narratives provide insight and context around events of interest and can be used to identify emerging problems. They can trigger investigations within Flight Operational Quality Assurance or Flight Data Monitoring programs. However, this process typically focuses on the adverse events and the unsafe aspects of the operations surrounding the reported or detected events. This perspective of investigating factors that went wrong around an adverse event is commonly referred to as Safety I. Alternatively, characterizing successful actions that operators perform every day under varying conditions that keep the system within safe operating bounds is a concept referred to as Safety II. The benefit of the Safety II view is that the scope is much larger than that of Safety I since a vast majority of the operations result in successful flights. Many of the successful techniques used to manage operational threats are not documented in standard operating procedures or taught during training. They are typically acquired over time by working with experienced pilots during line operations or in many cases after experiencing a problem for the first time and reacting to it in situ, drawing from years of experience to manage the threat. In an attempt to quantify these positive actions, we are proposing an approach to extracting key behaviors within ASRS reports that can support the Safety II concept. Our analysis assumes that ASRS reports contain some descriptions of corrective actions that operators performed to prevent a situation from leading to an accident. Leveraging recent advances in Natural Language Process modeling, we have developed an approach to extract positive sentiment from reports, embed these positive statements in a vector space where they can be numerically analyzed, and clustering these statements into similar contextual categories. From these contextualized categories we can attempt to summarized and distilled aspects of the positive behavior. The goal is to identify categories of behavior that describe consistent operator techniques that supports the Safety II concept. With this information, airlines may enable learning from these positive actions, or address procedures that need to be changed to avoid having pilots implement a workaround. These insights can provide a lens into what is “going right” in the operations that may otherwise not be known widely within the community. It is envisioned that this approach can be extended to other narrative programs such as Line Operation Safety Audit or Learning Improvement Team reports where similar observed behavior can be analyzed to extract positive actions and inform the overall operations.

NLP