Search NASA⌕ Search

SEARCH · Search NASA

Results for “text analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Worldwide Spacecraft Crew Hatch History

The JSC Flight Safety Office has developed this compilation of historical information on spacecraft crew hatches to assist the Safety Tech Authority in the evaluation and analysis of worldwide spacecraft crew hatch design and performance. The document is prepared by SAIC s Gary Johnson, former NASA JSC S&MA Associate Director for Technical. Mr. Johnson s previous experience brings expert knowledge to assess the relevancy of data presented. He has experience with six (6) of the NASA spacecraft programs that are covered in this document: Apollo; Skylab; Apollo Soyuz Test Project (ASTP), Space Shuttle, ISS and the Shuttle/Mir Program. Mr. Johnson is also intimately familiar with the JSC Design and Procedures Standard, JPR 8080.5, having been one of its original developers. The observations and findings are presented first by country and organized within each country section by program in chronological order of emergence. A host of reference sources used to augment the personal observations and comments of the author are named within the text and/or listed in the reference section of this document. Careful attention to the selection and inclusion of photos, drawings and diagrams is used to give visual association and clarity to the topic areas examined.

Johnson, Gary↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗

JPSS-1 ATMS Post-launch Active Geolocation Analysis

A NOAA-20 (N20) ATMS active geolocation test was performed around Jan. 2018 with 24 pre-selected coastline crossing scenes. After comprehensive analysis of ATMS stare data and the corresponding VIIRS data, the ATMS pitch, roll and yaw pointing angle errors are found from the nadir perpendicular, the nadir oblique shallow angle, and the off-nadir perpendicular coastline crossing data, respectively. In this study, we first determine the ATMS radiometric coastline crossing time by using the ATMS radiometric count data. Since the coastline can be located anywhere within one ATMS FOV from the passive (regular scanning) geolocation data, depending on the scan starting time, it is not a valid assumption for the inflection point being the same as the coastline location. Consequently, using the passive geolocation data to validate the sensor’s on-orbit pointing angle performance is limited. After finding the ATMS radiometric coastline crossing time from the ATMS data, we compare it with the VIIRS effective time stamp (see main text). Specifically the VIIRS M3, M4 & M5 (true color) & M15 and M16 bands (thermal) data have a much smaller footprint size. Using the differences of the ATMS and VIIRS (effective) coastlines crossing times, the N20 ATMS pitch, roll and yaw pointing angle errors are found to be -0.09o, -0.24o and 0.28o, respectively. To determine ATMS geolocation properly, these on-orbit pointing errors need to be corrected, adding to the ATMS SDR Processing Coefficient Table, and passed on to the operational geolocation processing code.

Active Geolocation↗

Displaying Composite and Archived Soundings in the Advanced Weather Interactive Processing System

In a previous task, the Applied Meteorology Unit (AMU) developed spatial and temporal climatologies of lightning occurrence based on eight atmospheric flow regimes. The AMU created climatological, or composite, soundings of wind speed and direction, temperature, and dew point temperature at four rawinsonde observation stations at Jacksonville, Tampa, Miami, and Cape Canaveral Air Force Station, for each of the eight flow regimes. The composite soundings were delivered to the National Weather Service (NWS) Melbourne (MLB) office for display using the National version of the Skew-T Hodograph analysis and Research Program (NSHARP) software program. The NWS MLB requested the AMU make the composite soundings available for display in the Advanced Weather Interactive Processing System (AWIPS), so they could be overlaid on current observed soundings. This will allow the forecasters to compare the current state of the atmosphere with climatology. This presentation describes how the AMU converted the composite soundings from NSHARP Archive format to Network Common Data Form (NetCDF) format, so that the soundings could be displayed in AWl PS. The NetCDF is a set of data formats, programming interfaces, and software libraries used to read and write scientific data files. In AWIPS, each meteorological data type, such as soundings or surface observations, has a unique NetCDF format. Each format is described by a NetCDF template file. Although NetCDF files are in binary format, they can be converted to a text format called network Common data form Description Language (CDL). A software utility called ncgen is used to create a NetCDF file from a CDL file, while the ncdump utility is used to create a CDL file from a NetCDF file. An AWIPS receives soundings in Binary Universal Form for the Representation of Meteorological data (BUFR) format (http://dss.ucar.edu/docs/formats/bufr/), and then decodes them into NetCDF format. Only two sounding files are generated in AWIPS per day. One file contains all of the soundings received worldwide between 0000 UTC and 1200 UTC, and the other includes all soundings between 1200 UTC and 0000 UTC. In order to add the composite soundings into AWIPS, a procedure was created to configure, or localize, AWIPS. This involved modifying and creating several configuration text files. A unique fourcharacter site identifier was created for each of the 32 soundings so each could be viewed separately. The first three characters were based on the site identifier of the observed sounding, while the last character was based on the flow regime. While researching the localization process for soundings, the AMU discovered a method of archiving soundings so old soundings would not get purged automatically by AWl PS. This method could provide an alternative way of localizing AWl PS for composite soundings. In addition, this would allow forecasters to use archived soundings in AWIPS for case studies. A test sounding file in NetCDF format was written in order to verify the correct format for soundings in AWIPS. After the file was viewed successfully in AWIPS, the AMU wrote a software program in the Tool Command Language/Tool Kit (Tcl/Tk) language to convert the 32 composite soundings from NSHARP Archive to CDL format. The ncgen utility was then used to convert the CDL file to a NetCDF file. The NetCDF file could then be read and displayed in AWIPS.

Barrett, Joe H., III↗

PFLOTRAN modeling data and scripts associated with “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics” submitted to Water Resources Research (Terry et al. 2025). The data package contains the groundwater modeling dataset from PFLOTRAN software. It includes the python script for mesh generation, boundary condition setting, PFLOTRAN input deck formation and postprocessing. It couples groundwater flow and species transport for Hanford Reach river corridor and pipelines the model generation and processing. This model can be used to easily generate the model and analysis for Hanford site. It can also be adjusted to other hydrologic area with ease. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package consists of 6 folders: (1) “data” contains all necessary data as input and intermediate data for processing; (2) “mesh” contains all mesh related files to generate mesh in Hanford Reach river corridor; (3) “model_run” contains the generated script for PFLOTRAN modeling; (4) “notebooks” contains all the Python script to generate the model; (5) “output” contains all the output from the computation; (6) “postprocessing” contains the Python script to generate scientific figure for manuscript. All files are .csv (comma-separated values), .h5 (HDF5 format), .in (input files), .ipynb (Jupyter notebooks), .p (Python pickle), .png (images), .PNG (images), .py (Python scripts), .pyc (Python bytecode), .r (R scripts), .sh (shell scripts), .txt (text files), .vtu (3D mesh/visualization format), .xz (compressed archive), or .zip (compressed archive).

54 ENVIRONMENTAL SCIENCES↗

Computer-Aided System Engineering and Analysis (CASE/A) Programmer's Manual, Version 5.0

The Computer Aided System Engineering and Analysis (CASE/A) Version 5.0 Programmer's Manual provides the programmer and user with information regarding the internal structure of the CASE/A 5.0 software system. CASE/A 5.0 is a trade study tool that provides modeling/simulation capabilities for analyzing environmental control and life support systems and active thermal control systems. CASE/A has been successfully used in studies such as the evaluation of carbon dioxide removal in the space station. CASE/A modeling provides a graphical and command-driven interface for the user. This interface allows the user to construct a model by placing equipment components in a graphical layout of the system hardware, then connect the components via flow streams and define their operating parameters. Once the equipment is placed, the simulation time and other control parameters can be set to run the simulation based on the model constructed. After completion of the simulation, graphical plots or text files can be obtained for evaluation of the simulation results over time. Additionally, users have the capability to control the simulation and extract information at various times in the simulation (e.g., control equipment operating parameters over the simulation time or extract plot data) by using "User Operations (OPS) Code." This OPS code is written in FORTRAN with a canned set of utility subroutines for performing common tasks. CASE/A version 5.0 software runs under the VAX VMS(Trademark) environment. It utilizes the Tektronics 4014(Trademark) graphics display system and the VTIOO(Trademark) text manipulation/display system.

Knox, J. C.↗

First-results from the Perseverance SHERLOC Investigation: Aqueous Alteration Processes and Implications for Organic Geochemistry in Jezero Crater, Mars

The Perseverance rover landed in Jezero crater, a site selected to fulfill the Mars-2020 mission goals of characterizing the geology of habitable environments and searching for signs of life while collecting samples for return to Earth [1]. Jezero hosted an open-basin lake during the late Noachian/early Hesperian (~3.7 Ga) [1-2], has units associated with the largest carbonate deposit identified on Mars [3-4], and has a well-preserved delta with clay and carbonate-bearing sediments, well-suited to preservation of organics [1,3-4]. Investigating the nature of organics and aqueous environments within their geologic con-text allows us to understand important aqueous processes and determine habitability within Jezero crater. Previous in situ landed measurements of organics could not resolve their spatial and mineralogical con-text [5-6]. Although Martian meteorites lack geological context, the spatial distribution of organic compounds in Martian meteorites have allowed recognition of an association between aqueous processes and organics [7-8]. Here, we show for the first time in-situ associations between carbonate-forming ultramafic alteration process, later stage aqueous sulfate and perchlorate formation, and organics on the Martian surface. Methodology and geological context: We use the Perseverance rover’s SHERLOC instrument (Scanning Habitable Environments with Raman and Lumines-ence of Organics and Chemicals), a deep-ultraviolet fluorescence and Raman scattering spectrometer capable of mapping the organic and mineral composition with a spatial resolution of 100 μm resolution to report the presence of organics and aqueously formed minerals at Jezero crater [9]. These spectral detections were compared with co-located images obtained with the autofocus context imager (ACI) and the WATSON camera for textural analysis [9]. As of writing, the Per-severance rover has abraded five targets that were measured with the SHERLOC instrument. The five targets are located in two different orbitally-identified geological units within the floor of Jezero crater; the Crater Floor Fractured Rough unit (CF-Fr) and the Séítah region within the Crater Floor Fractured 1 unit (CF-F1) [10]. In orbital infrared spectroscopic data, the CF-Fr unit is associated with pyroxene spectral signatures and minor alteration, while the Séítah region is associated with olivine and minor Mg-rich carbonates and clays [3-4,10]. Carbonation of ultramafic protolith recorded within Jezero crater: All scans of abraded targets within the Séítah region reveal strong peaks at 1080–1090 cm−1 consistent with carbonate and peak singlets or doublets at 820–840 cm−1 attributed to olivine (Fig. 1), consistent with orbital infrared observations. Our detailed micron-scale petrographic and spectroscopic evidence shows that these carbonates formed through carbonation of an ultramafic protolith. The supporting observations include: (1) Carbonate cation compositions match those of olivine, suggesting mixed Fe- and Mg-olivine gave rise to mixed Fe- and Mg-carbonates, similar to observations of ultramafic systems on Earth and within Martian meteorites [3-4,7-8]. (2) The ob-served carbonates co-occur with hydrated materials, gypsum, and potentially aqueously-formed phases, amorphous silicates and phosphate. (3) The spectral and textural variation of olivine and carbonate dominated zones and olivine-carbonate mixtures within both primary grains and interstitial zones are expected for carbonated ultramafic protoliths. (4) These mineral associations and textures closely resemble those observed within the ALH84001 and Nakhlite meteorites attributed to olivine carbonation on Mars [6-7]. Taken together, micron-scale SHERLOC documentation of these phenomena bridge previous orbital and meteorite observations and demonstrate in-situ regionally extensive (~106 km2) ultramafic alteration resulting in geo-logical deposition of carbonates. Furthermore, we observe that olivine carbonation was involved in preserving and possibly synthesizing organics, which makes this environment potentially habitable, as previously suggested in [1,3-4] (Fig. 1). Late-stage aqueous perchlorate and sulfate in Jezero crater: An abrasion target within the CF-Fr unit contains combinations of high intensity 950-955 cm−1 peaks and minor 1090-1095 cm−1 and 1150-1155 cm−1 peaks that are spectral fits to anhydrous perchlorate (Fig. 1). Some spectra show a combination of 950-955 cm−1 peaks with equally strong 1010-1020 cm−1 peaks, low intensity broad features at 1120 cm−1, and occasional broad 3450 cm−1 hydration (-OH) features, indicating a mixture of Ca-sulfate and perchlorate that is minimally hydrated (Fig. 1). The detections of per-chlorates within Jezero crater differ from previous measurements (e.g. Phoenix lander, Curiosity rover, Tissint meteorite [7,11]) because they are observed to be intimately related to aqueous processes including sulfate formation, they present as a secondary white void-fill occurring within the interior of the rock, and they are found to likely be Na-perchlorate. implications for their formation: Three different types of organics embedded within three different lithologies were observed within the abraded targets. Organics associated with low intensity ~340 nm fluorescence were widespread within targets with no apparent association to particular minerals (Fig. 1). Organics associated with ~305 nm and ~275 nm fluorescence correlated with sulfates within the Bellegarde target in the CF-Fr unit, while organics associated with high intensity ~340 nm fluorescence correlated with carbonate, phosphate, and amorphous silicate mixtures within the Garde target in the Séítah region (Fig. 1). Although assignment of fluorescence signatures to specific organic compounds is not conclusive, ~340 nm fluorescence is generally more consistent with 2-ring aromatic organics, ~275 nm fluorescence is more consistent with 1-ring aromatic organics, and ~305 nm fluorescence can be created by either 2- or 1-ring aromatics [12]. These observations indicate that the strongest fluorescence signatures interpreted as organics were found in materials associated with aqueous processes, i.e. sulfate- and carbonate-bearing materials, suggesting both brines and ultramafic carbonation aqueous environments were capable of preserving organics on ancient Mars. In Martian meteorites, simple aromatic organics proposed to have been synthesized through aqueous processes can be found within minerals associated with olivine carbonation and in spatial association with perchlorate and sulfate materials [7-8], similar to SHERLOC observations. Hence, we advance an abiotic aqueous synthesis origin for the organics although we cannot rule out the presence of organics from meteoritic in-fall or putative organic biosignatures. Detailed analyses will be required upon return of these materials to Earth.

E L Scheller↗

Exploration Clinical Decision Support System: Medical Data Architecture

The Exploration Clinical Decision Support (ECDS) System project is intended to enhance the Exploration Medical Capability (ExMC) Element for extended duration, deep-space mission planning in HRP. A major development guideline is the Risk of "Adverse Health Outcomes & Decrements in Performance due to Limitations of In-flight Medical Conditions". ECDS attempts to mitigate that Risk by providing crew-specific health information, actionable insight, crew guidance and advice based on computational algorithmic analysis. The availability of inflight health diagnostic computational methods has been identified as an essential capability for human exploration missions. Inflight electronic health data sources are often heterogeneous, and thus may be isolated or not examined as an aggregate whole. The ECDS System objective provides both a data architecture that collects and manages disparate health data, and an active knowledge system that analyzes health evidence to deliver case-specific advice. A single, cohesive space-ready decision support capability that considers all exploration clinical measurements is not commercially available at present. Hence, this Task is a newly coordinated development effort by which ECDS and its supporting data infrastructure will demonstrate the feasibility of intelligent data mining and predictive modeling as a biomedical diagnostic support mechanism on manned exploration missions. The initial step towards ground and flight demonstrations has been the research and development of both image and clinical text-based computer-aided patient diagnosis. Human anatomical images displaying abnormal/pathological features have been annotated using controlled terminology templates, marked-up, and then stored in compliance with the AIM standard. These images have been filtered and disease characterized based on machine learning of semantic and quantitative feature vectors. The next phase will evaluate disease treatment response via quantitative linear dimension biomarkers that enable image content-based retrieval and criteria assessment. In addition, a data mining engine (DME) is applied to cross-sectional adult surveys for predicting occurrence of renal calculi, ranked by statistical significance of demographics and specific food ingestion. In addition to this precursor space flight algorithm training, the DME will utilize a feature-engineering capability for unstructured clinical text classification health discovery. The ECDS backbone is a proposed multi-tier modular architecture providing data messaging protocols, storage, management and real-time patient data access. Technology demonstrations and success metrics will be finalized in FY16.

Biomedical support↗

Control And Optimization Modular Modeling Application For Nuclear Deployment

The purpose of the COMMAND code is to provide a flexible, scalable tool for use in developing, integrating, and testing the technologies necessary for achieving autonomous operations of advanced nuclear reactors. The code enables users to efficiently implement custom simulations and experiments by combining key methods from different software modules. These modules are focused on: modeling and simulation tools, such as nuclear simulation tools used for high-fidelity modeling (e.g., Reactor Excursion and Leak Analysis Program [RELAP5-3D] and Monte Carlo N-Particle [MCNP]); machine learning and optimization tools (e.g., anomaly detection and data-driven modeling techniques); advanced control in its digital, high-performance, and supervisory control forms (e.g., proportional integral derivative (PID) control and model predictive control (MPC); and integration with hardware through industrial communication protocols. To ensure flexibility and scalability, COMMAND was designed to be both modular—the software “pieces” all inherit from generic building blocks and can be combined and connected to create complicated simulations—and high performing—designed for parallel processing, enabling simulations and experiments to take advantage of multi-core computers, servers, and nodes. The code is written in the Python programming language due to the language's popularity, active community, and open-source and cross-platform nature. Maintaining consistency with other simulation tools used within the nuclear energy community, users implement simulations and experiments through text input files, which define components, parameters, connections, etc., through lines of text. Given that COMMAND is written in Python, these input files are native Python scripts, and so use the standard Python structure and formatting. This also enables users to take advantage of Python's extensive package library to develop custom capabilities for their specific use cases.

Faber, Jacob [Idaho National Laboratory (INL), Ida↗

An MBSE Framework to Identify Regulatory Gaps for Electrified Transport Aircraft

A typical aircraft certification process consists of obtaining a type, production, airworthiness, and continued airworthiness certificates. During this process, a type certification plan is created that includes the intended regulatory operating environment, the proposed certification basis, means of compliance, and a list of documentation to show compliance. Earlier work by the authors demonstrated a model-based framework for the management of these certification artifacts for normal category airplanes. Presently, it is expanded and adapted to consider certification for transport category airplanes regulated under 14 CFR Part 25, providing clear transparency and traceability between the text of the regulations and imposed requirements, contextual information, and specified test activities. In particular, a capability to identify potential gaps in the applicability of regulations for novel architectures such as electrified aircraft is proposed. This capability, based on mismatches between the functional intent and the corresponding prescribed physical implementation, is developed. A sample implementation of the proposed capability is presented for a notional electrified powertrain aircraft architecture.

electrified transport aircraft↗

Generative AI for Power Grid Operations

Generative artificial intelligence (AI) has captured into the mainstream, demonstrating capabilities that once belonged solely to the realm of human cognition. From defeating world champions in complex games to generating human-quality text and images, Generative AI has proven its potential to revolutionize countless industries. The electric power grid is no exception. Generative AI's ability to process vast amounts of data rapidly, assist decision support and identify patterns could significantly enhance power grid operations. For example, Generative AI could improve state estimation where measurements are not available or integrate renewable energy sources more efficiently with probabilistic forecasting. The key contributions of this whitepaper are outlined below: (1) Comprehensive overview of Generative AI's applications in power grid operations: It highlights the opportunities in areas such as forecasting, state estimation, and demonstrating the potential for enhancing efficiency, reliability, and resilience. (2) Expanding Generative AI's impact through synergies with emerging technologies: The paper introduce NREL developed eGridGPT and explores how AI orchestration, multi-agent systems, and Digital Twins can collaborate to optimize grid operations, addressing the complexities of a decarbonized and electrified future. (3) In-depth analysis of challenges in implementing Generative AI: This includes considerations like data availability and quality, model validation, certification, and ethical concerns, ensuring responsible AI deployment. (4) Emphasizing human-AI collaboration: The whitepaper underscores the importance of trustworthy, transparency, and explainability in AI systems to promote seamless interaction between human operators and AI, ultimately improving decision-making. (5) Exploring future research and development: It identifies critical areas for further advancement to fully realize Generative AI's potential in power grid operations. This whitepaper serves as a valuable resource for researchers, practitioners, and policymakers looking to harness Generative AI for a more reliable, stable, and cost-effective power grid.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhancing NASA Earth Science Data Discovery from Scientific Publications

Earth observations from space borne instruments have evolved explosively in the past decades. Following closely are reanalysis systems assimilating model and observational data, yielding even longer records and larger number of variables. Thanks to advances in internet technology, it is now easier than ever to visualize and analyze these data using web interfaces. On the other hand, it also becomes an increasingly daunting task to build upon the existing knowledge published in various peer reviewed sources, and navigate toward the most relevant data, analysis, and visualization. We present an analysis of a subset of publications that utilized a popular visualization web interface at the NASA Goddard Earth Science Data and Information Services Center. Known as "Giovanni", it allows researchers from wide backgrounds to work with hundreds of variables from space observations and assimilation systems. Since coming online more than a decade ago, Giovanni has been credited in more than 100 papers per year, and the total count now is estimated to be nearly 1,500. Many of these papers contain valuable information about when, where and how Giovanni has been used, and hence forge an opportunity to learn and share the knowledge of which variables were used for what research projects. The purpose of our work is to retrieve the information from the papers and organize it as a knowledge repository which links together datasets, variables, places, dates and phenomena all of which reflect the essence of the published research. Since the publications are unstructured texts, we use natural language processing along with machine learning methods in the retrieval process. One of the challenges is deciphering the dataset names, because in many cases researchers refer to variables, rather than the datasets containing them. To constrain the number of terms, we deploy Earth Science ontologies as dictionaries for the term extraction. We demonstrate that storing these terms and underlying ontologies, along with datasets, variables and papers in the knowledge graph database, enables various linkages between all these entities facilitating the data discovery. Thus, we are setting a qualitatively new stage in improvements of web data interfaces, where machine learning techniques are used to establish and optimize usage-based discovery of data.

Irina V Gerasimov↗

Storage of Physical Sample Metadata in the Astrobiology Habitable Environments Database (AHED)

The National Aeronautics and Space Administration has begun an effort to store, curate, and publish information about physical samples collected and analyzed in conjunction with NASA-funded astrobiology research. Astrobiology is a multidisciplinary area of scientific research being conducted by collaborating teams of biologists, chemists, geologists, atmospheric scientists, oceanographers, astrophysicists, astronomers, and other specialists. Astrobiology studies the origin, evolution, and distribution of life in the Universe. NASA uses the results of astrobiology research to focus its future missions on targets of opportunity for the discovery of life off Earth. Astrobiology researchers conduct both field-based and laboratory-based research, during which physical samples are collected, processed, and catalogued. The cataloguing practices employed by different teams of astrobiologists vary widely, and there are no specific standards available to guide the collection and recording of astrobiology sample data. The disparity in data collection approaches and the lack of a centralized sample repository makes it difficult for astrobiology teams to share data and benefit from resultant synergies.To facilitate data sharing within the astrobiology community, NASA is developing a prototype database the Astrobiology Habitable Environments Database (AHED) and an associated set of data collection templates. The database will store information about samples, along with associated measurements and analyses, including information about biological cultures enriched or isolated from samples, and the results of analyses performed on the samples (e.g., via spectrography, microscopy, etc.). In addition, the system will store contextual information about field sites where samples were collected, the instruments or equipment used for analysis, and people and institutions involved in their collection. AHED is being implemented on top of Open Data Repository's Data Publisher [1], an open source software platform for the publication of scientific datasets. The data collection templates under development represent an initial attempt to propose a set of metadata for capture and storage within AHED. The design of these templates is being conducted by a consolidated group of astrobiologists from active research teams at NASA Ames Research Center, assisted by data science and software engineering specialists. These initial templates must be vetted with the broader astrobiology community through a defined process to ensure that they meet community needs. Each template captures a different type of data collection record. For each template, we are developing a list of fields to be captured, including a set of required entry fields, a set of recommended but optional fields, and a set of discretionary fields. A datatype selected from a variety of text and numeric types is specified for each field. Included is a 'choice' type that restricts user input to an enumerated list of values. Many of the fields and field values capture information of particular interest to the astrobiology community, and are intended to facilitate search and retrieval of relevant data across multiple datasets.

Keller, Rich↗

TRMM Version 7 Level 3 Gridded Monthly Accumulations of GPROF Precipitation Retrievals

In July 2011, improved versions of the retrieval algorithms were approved for TRMM. All data starting with June 2011 are produced only with the version 7 code. At the same time, version 7 reprocessing of all TRMM mission data was started. By the end of August 2011, the 14+ years of the reprocessed mission data became available online to users. This reprocessing provided the opportunity to redo and enhance upon an analysis of V7 impacts on L3 data accumulations that was presented at the 2010 EGU General Assembly. This paper will discuss the impact of algorithm changes made in th GPROF retrieval on the Level 2 swath products. Perhaps the most important change in that retrieval was to replacement of a model based a priori database with one created from Precipitation Radar (PR) and TMI brightness temperature (Tb) data. The radar pays a major role in the V7 GPROF (GPROF2010) in determining existence of rain. The level 2 retrieval algorithm also introduced a field providing the probability of rain. This combined use of the PR has some impact on the retrievals and created areas, particularly over ocean, where many areas of low-probability precipitation are retrieved whereas in version 6, these areas contained zero rain rates. This paper will discuss how these impacts get translated to the space/time averaged monthly products that use the GPROF retrievals. The level 3 products discussed are the gridded text product 3G68 and the standard 3A12 and 3B31 products. The paper provides an overview of the changes and explanation of how the level 3 products dealt with the change in the retrieval approach. Using the .25 deg x .25 degree grid, the paper will show that agreement between the swath product and the level 3 remains very high. It will also present comparisons of V6 and V7 GPROF retrievals as seen both at the swath level and the level 3 time/space gridded accumulations. It will show that the various L3 products based on GPROF level 2 retrievals are in close agreement. The paper concludes by outlining some of the challenges of the TRMM version 7 level 3 products.

Stocker, E. F.↗

Inelastic response of metal matrix composites under biaxial loading

Elements of the analytical/experimental program to characterize the response of silicon carbide titanium (SCS-6/Ti-15-3) composite tubes under biaxial loading are outlined. The analytical program comprises prediction of initial yielding and subsequent inelastic response of unidirectional and angle-ply silicon carbide titanium tubes using a combined micromechanics approach and laminate analysis. The micromechanics approach is based on the method of cells model and has the capability of generating the effective thermomechanical response of metal matrix composites in the linear and inelastic region in the presence of temperature and time-dependent properties of the individual constituents and imperfect bonding on the initial yield surfaces and inelastic response of (0) and (+ or - 45)sub s SCS-6/Ti-15-3 laminates loaded by different combinations of stresses. The generated analytical predictions will be compared with the experimental results. The experimental program comprises generation of initial yield surfaces, subsequent stress-strain curves and determination of failure loads of the SCS-6/Ti-15-3 tubes under selected loading conditions. The results of the analytical investigation are employed to define the actual loading paths for the experimental program. A brief overview of the experimental methodology is given. This includes the test capabilities of the Composite Mechanics Laboratory at the University of Virginia, the SCS-6/Ti-15-3 composite tubes secured from McDonnell Douglas Corporation, a text fixture specifically developed for combined axial-torsional loading, and the MTS combined axial-torsion loader that will be employed in the actual testing.

Mirzadeh, F.↗

Reaching For New Physics With MeV-scale Reconstruction In The MicroBooNE LArTPC Neutrino Detector

Large neutrino liquid argon time projection chamber (LArTPC) experiments can broaden their physics reach by reconstructing MeV-Scale energy depositions, or blips, in their data. We demonstrate new calorimetric and particle discrimination capabilities at the MeV scale using reconstructed blips in MicroBooNE LArTPC data at Fermilab. A concentration of low-energy ($<$3 MeV) blips is observed around fiberglass mechanical support struts along the TPC edges, with spectral features consistent with the Compton edge of the 2.614 MeV $^{208}$Tl decay $\gamma$ ray. With these features we perform the electron energy scale calibration to few-percent precision and yield the specific activity of $^{208}$Tl in the struts, $(11.7 \pm 0.2 \text{(stat)} \pm 2.8 \text{(syst)})$ Bq/kg. Using cosmogenic blips above 3 MeV, we demonstrate the ability of large LArTPCs to discriminate low-energy proton and electron depositions. An enriched low-energy proton sample selected with this technique is smaller in data than in dedicated CORSIKA simulations, pointing to possible mismodeling in CORSIKA incident cosmic fluxes or Geant4 particle transport. These methods are applied to MicroBooNE's inclusive single-photon search, which reported a 2.2$\sigma$ excess below 600 MeV in shower energy for events with no reconstructed protons. By identifying and classifying blips near single-photon events selected by the WireCell reconstruction framework, a more comprehensive labeling of nearby hadronic activity is established: blips upstream of the shower axis indicate previously unidentified final-state protons, while elevated blip counts at wide angles signal final-state neutrons. Taken together with MiniBooNE's long-standing low-energy excess (LEE) and MicroBooNE electron-like and sterile neutrino searches disfavored as possible explanations of the MiniBooNE anomaly, this analysis motivates an expanded exploration of the single-photon channel in Fermilab's short-baseline LArTPC program. This thesis documents the current status of this enhanced analysis, which will form a key part of MicroBooNE's final low-energy-excess results.

Andrade Aldana, Diego Armando [IIT, Chicago (main)↗

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics↗

FAST User Guide

The Flow Analysis Software Toolkit, FAST, is a software environment for visualizing data. FAST is a collection of separate programs (modules) that run simultaneously and allow the user to examine the results of numerical and experimental simulations. The user can load data files, perform calculations on the data, visualize the results of these calculations, construct scenes of 3D graphical objects, and plot, animate and record the scenes. Computational Fluid Dynamics (CFD) visualization is the primary intended use of FAST, but FAST can also assist in the analysis of other types of data. FAST combines the capabilities of such programs as PLOT3D, RIP, SURF, and GAS into one environment with modules that share data. Sharing data between modules eliminates the drudgery of transferring data between programs. All the modules in the FAST environment have a consistent, highly interactive graphical user interface. Most commands are entered by pointing and'clicking. The modular construction of FAST makes it flexible and extensible. The environment can be custom configured and new modules can be developed and added as needed. The following modules have been developed for FAST: VIEWER, FILE IO, CALCULATOR, SURFER, TOPOLOGY, PLOTTER, TITLER, TRACER, ARCGRAPH, GQ, SURFERU, SHOTET, and ISOLEVU. A utility is also included to make the inclusion of user defined modules in the FAST environment easy. The VIEWER module is the central control for the FAST environment. From VIEWER, the user can-change object attributes, interactively position objects in three-dimensional space, define and save scenes, create animations, spawn new FAST modules, add additional view windows, and save and execute command scripts. The FAST User Guide uses text and FAST MAPS (graphical representations of the entire user interface) to guide the user through the use of FAST. Chapters include: Maps, Overview, Tips, Getting Started Tutorial, a separate chapter for each module, file formats, and system administration.

Walatka, Pamela P.↗