Search NASASearch

SEARCH · Search NASA

Results for “text analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

CATLAC - Calibration and Validation Analysis Tool of Local Area Coverage for the SeaWiFS Mission

Calibration and validation Analysis Tool of Local Area Coverage (CATLAC) is an analysis package for selecting and graphically displaying Earth and space targets for calibration and validation activities on a polar orbiting satellite. The package is written in the Interactive Data Language (IDL) and includes a graphical user interface. Although it is designed specifically for the Sea-viewing Wide Field-of-view Sensor (SeaWiFS) mission, the package can be used for analysis on other Earth-viewing missions. An individual can use text or graphical methods in CATLAC to select Earth targets to be scanned by a satellite. Additional onboard calibration activities (such as observations of the moon, or solar irradiance from a solar diffuser), which use data recorder time, can also be specified. All information pertinent to the creation of a command schedule can be written to a file which is read by a command scheduler. The scheduler can be invoked and the Local Area Coverage (LAC) recording periods can be visually verified using CATLAC. The schedule can also be verified by examining record and error files written by the scheduler.

Woodward, Robert H.

Data Science and the Knowledge Discovery Adventure

This talk will cover the important steps involved in the data science and knowledge discovery process: • Initial fact gathering (interview domain experts, review reports, articles, state-of-the-art) • Identify the problem (prediction, classification, statistical analysis, etc.) • Survey supporting data sources • Understand the data (numerical, categorical, text, sampling rate, data quality issues, etc.) • Selecting relevant features and sources • Acquire the data (set up agreements with the data stewards, APIs to download, etc.) • Merge data sources (temporal, spatial, common key, other ontologies...) • Feature Engineering (non linear domain knowledge or physics-based relationships) • Build data processing pipeline (may need to tap into data stream, develop parallel processing algorithm, federated learning etc.) • Build model and test (tune hyper-parameters, cross validation.) • Analyze/Validate results (do the results make sense. Does it answer the original question). • Deploy/Publish (Monitor and assess benefits)

Data science

Eucalyptus – An Analysis Suite for Fault Trees with Uncertainty Quantification

Eucalyptus is a novel code developed at Lawrence Livermore National Laboratory to incorporate uncertainty quantification into Fault Tree Analysis (FTA). This tool addresses the challenge of imperfect knowledge in “grey-box” systems by allowing analysts to incorporate and propagate uncertainty from component-level assessments to system-level effects. Eucalyptus facilitates a consistent evaluation of the impact of subject matter expert judgment and knowledge gaps on overall system response by Monte Carlo generation of possible system fault trees, sampling probabilities of the existence of subsystems and components. Here, the code supports the specification of fault trees through text and allows export to various formats, including auto-generated images, easing analysis and reducing errors. It has undergone extensive verification testing, demonstrating its reliability and readiness for deployment, and leverages on-node parallelism for rapid analysis. Example analyses are shown that include the identification of system failure paths and quantification of the value of further information about system components.

Fault Tree Analysis

Information Extraction for System-Software Safety Analysis: Calendar Year 2007 Year-End Report

This annual report describes work to integrate a set of tools to support early model-based analysis of failures and hazards due to system-software interactions. The tools perform and assist analysts in the following tasks: 1) extract model parts from text for architecture and safety/hazard models; 2) combine the parts with library information to develop the models for visualization and analysis; 3) perform graph analysis on the models to identify possible paths from hazard sources to vulnerable entities and functions, in nominal and anomalous system-software configurations; 4) perform discrete-time-based simulation on the models to investigate scenarios where these paths may play a role in failures and mishaps; and 5) identify resulting candidate scenarios for software integration testing. This paper describes new challenges in a NASA abort system case, and enhancements made to develop the integrated tool set.

Malin, Jane T.

From Text to Maps: LLM-Driven Extraction and Geotagging of Epidemiological Data

Epidemiological datasets are essential for public health analysis and decision-making, yet they remain scarce and often difficult to compile due to inconsistent data formats, language barriers, and evolving political boundaries. Traditional methods of creating such datasets involve extensive manual effort and are prone to errors in accurate location extraction. To address these challenges, we propose utilizing large language models (LLMs) to automate the extraction and geotagging of epidemiological data from textual documents. Our approach significantly reduces the manual effort required, limiting human intervention to validating a subset of records against text snippets and verifying the geotagging reasoning, as opposed to reviewing multiple entire documents manually to extract, clean, and geotag. Additionally, the LLMs identify information often overlooked by human annotators, further enhancing the dataset’s completeness. Our findings demonstrate that LLMs can be effectively used to semi-automate the extraction and geotagging of epidemiological data, offering several key advantages: (1) comprehensive information extraction with minimal risk of missing critical details; (2) minimal human intervention; (3) higher-resolution data with more precise geotagging; and (4) significantly reduced resource demands compared to traditional methods.

Harrod, Karly

Integrated Software for Analyzing Designs of Launch Vehicles

Launch Vehicle Analysis Tool (LVA) is a computer program for preliminary design structural analysis of launch vehicles. Before LVA was developed, in order to analyze the structure of a launch vehicle, it was necessary to estimate its weight, feed this estimate into a program to obtain pre-launch and flight loads, then feed these loads into structural and thermal analysis programs to obtain a second weight estimate. If the first and second weight estimates differed, it was necessary to reiterate these analyses until the solution converged. This process generally took six to twelve person-months of effort. LVA incorporates text to structural layout converter, configuration drawing, mass properties generation, pre-launch and flight loads analysis, loads output plotting, direct solution structural analysis, and thermal analysis subprograms. These subprograms are integrated in LVA so that solutions can be iterated automatically. LVA incorporates expert-system software that makes fundamental design decisions without intervention by the user. It also includes unique algorithms based on extensive research. The total integration of analysis modules drastically reduces the need for interaction with the user. A typical solution can be obtained in 30 to 60 minutes. Subsequent runs can be done in less than two minutes.

Philips, Alan D.

A critical study of the role of the surface oxide layer in titanium bonding

Scanning electron microscope/X-ray photoelectron spectroscopy (SEM/XPS) analysis of fractured adhesively bonded Ti 6-4 samples is discussed. The text adhesives incuded NR 056X polyimide, polypheylquinoxaline (PPQ), and LARC-13 polyimide. Differentiation between cohesive and interfacial failure was based on the absence of presence of a Ti 2p XPS photopeak. In addition, the surface oxide layer on Ti-(6A1-4V) adherends is characterized and bond strength and durability are addressed. Bond durability in various environmental conditions is discussed.

Dias, S.

DESI DR1 Ly$α$ forest: 3D full-shape analysis and cosmological constraints

We perform an analysis of the full shapes of Lyman-$α$ (Ly$α$) forest correlation functions measured from the first data release (DR1) of the Dark Energy Spectroscopic Instrument (DESI). Our analysis focuses on measuring the Alcock-Paczynski (AP) effect and the cosmic growth rate times the amplitude of matter fluctuations in spheres of $8$$h^{-1}\text{Mpc}$, $fσ_8$. We validate our measurements using two different sets of mocks, a series of data splits, and a large set of analysis variations, which were first performed blinded. Our analysis constrains the ratio $D_M/D_H(z_\mathrm{eff})=4.525\pm0.071$, where $D_H=c/H(z)$ is the Hubble distance, $D_M$ is the transverse comoving distance, and the effective redshift is $z_\mathrm{eff}=2.33$. This is a factor of $2.4$ tighter than the Baryon Acoustic Oscillation (BAO) constraint from the same data. When combining with Ly$α$ BAO constraints from DESI DR2, we obtain the ratios $D_H(z_\mathrm{eff})/r_d=8.646\pm0.077$ and $D_M(z_\mathrm{eff})/r_d=38.90\pm0.38$, where $r_d$ is the sound horizon at the drag epoch. We also measure $fσ_8(z_\mathrm{eff}) = 0.37\; ^{+0.055}_{-0.065} \,(\mathrm{stat})\, \pm 0.033 \,(\mathrm{sys})$, but we do not use it for cosmological inference due to difficulties in its validation with mocks. In $Λ$CDM, our measurements are consistent with both cosmic microwave background (CMB) and galaxy clustering constraints. Using a nucleosynthesis prior but no CMB anisotropy information, we measure the Hubble constant to be $H_0 = 68.3\pm 1.6\;\,{\rm km\,s^{-1}\,Mpc^{-1}}$ within $Λ$CDM. Finally, we show that Ly$α$ forest AP measurements can help improve constraints on the dark energy equation of state, and are expected to play an important role in upcoming DESI analyses.

Cuceu, Andrei [LBL, Berkeley; Chicago U., KICP] (O

The QDP/PLT user's guide

PLT is a high level plotting package. A Programmer can create a default plot suited for the data being displayed. At run times, users can then interact with the plot overriding any or all of these defaults. The user is also provided the capability to fit functions to the displayed data. This ability to display, interact with, and to fit the data make PLT a useful tool in the analysis of data. The Quick and Dandy Plotter (QDP) program will read ASCII text files that contain PLT commands and data. Thus, QDP provides and easy way to use the PLT software QPD files provide a convenient way to exchange data. The QPD/PLT software is written in standard FORTRAN 77 and has been ported to VAX VMS, SUN UNIX, IBM AIX, NeXT NextStep, and MS-DOS systems.

Tennant, Allyn F.

NTRE extended life feasibility assessment

Results of a feasibility analysis of a long life, reusable nuclear thermal rocket engine are presented in text and graph form. Two engine/reactor concepts are addressed: the Particle Bed Reactor (PBR) design and the Commonwealth of Independent States (CIS) concept. Engine design, integration, reliability, and safety are addressed by various members of the NTRE team from Aerojet Propulsion Division, Energopool (Russia), and Babcock & Wilcox.

Source record

Topic Modeling Tool for PeTaL (Periodic Table of Life)

A topic modeling tool is constructed for the purpose of providing insights from biology to the engineer within the framework of PeTaL (Periodic Table of Life). The machine learning text mining tools–latent Dirichlet allocation (LDA) and nonnegative matrix factorization (NMF) with Kullback-Leibler (KL) divergence—are used to provide topic clusters to the user. Topic clusters are the underlying themes of a paper. For the text modeling problem, NMF-KL is the equivalent of probabilistic latent semantic analysis. Both LDA and NMF-KL are top-performing modeling tools. These tools are used to identify biological specimens relevant to the user. Various organisms solve a particular survival problem in nature differently. The topic clusters allow people without domain expertise to find these cross-topic themes in the body of documents and then branch out and examine papers whose target organisms solve the engineer’s problem. Abstracts from the Journal of Experimental Biology were used as input for the clustering tool in addition to a curated set of articles for validation. The tool is able to accept alternate input sources.

Machine learning

Applications of Automation Methods for Nonlinear Fracture Test Analysis

As fracture mechanics material testing evolves, the governing test standards continue to be refined to better reflect the latest understanding of the physics of the fracture processes involved. The traditional format of ASTM fracture testing standards, utilizing equations expressed directly in the text of the standard to assess the experimental result, is self-limiting in the complexity that can be reasonably captured. The use of automated analysis techniques to draw upon a rich, detailed solution database for assessing fracture mechanics tests provides a foundation for a new approach to testing standards that enables routine users to obtain highly reliable assessments of tests involving complex, non-linear fracture behavior. Herein, the case for automating the analysis of tests of surface cracks in tension in the elastic-plastic regime is utilized as an example of how such a database can be generated and implemented for use in the ASTM standards framework. The presented approach forms a bridge between the equation-based fracture testing standards of today and the next generation of standards solving complex problems through analysis automation.

Allen, Phillip A.

Jackson, L., Johnson, M.B., Latrach, A., Grimes, D., Martinez, C., and Mclaughlin, J.F., 2024, Multidisciplinary geotechnical data collection, curation, and analysis for conformity with the regulatory framework for geologic carbon storage in Wyoming, USA: Geological Society of America Abstracts with Programs. Vol. 56, No. 5, 2024, doi: 10.1130/abs/2024AM-405024

Title: Multidisciplinary Geotechnical Data Collection, Curation, and Analysis for Conformity with the Regulatory Framework for Geologic Carbon Storage in Wyoming, USA. Text: Construction and operation of wells for geologic sequestration of carbon dioxide necessitate that they are permitted under the Environmental Protection Agency’s Underground Injection Control Class VI requirements. Class VI wells conform to stringent requirements to ensure long-term safety and integrity of the storage site and the protection of Underground Sources of Drinking Water. Entities pursuing Class VI permitting must provide comprehensive geologic site characterization, including regional geologic structure and stratigraphy, aquifer information, reservoir and confining unit geomechanical properties, geochemical analyses, assessment of trapping capacity and mechanisms, and a variety of other of multidisciplinary geotechnical data. The Wyoming Class VI Site Characterization Database Project is focused on developing a geologic site characterization database of geotechnical information, which has been compiled and verified from established, public databases/entities and scientific literature to expedite Class VI permitting in Sweetwater County within the Greater Green River Basin of southern Wyoming. The preliminary suite of compiled data from 14,000 wells includes 8,000 wells with logs and 7,250 wells with formation tops, ~70 wells with core data (e.g., X-Ray diffraction, petrographic, and petrophysical data), ~2,500 water analyses, ~740 seismic events data, and ~520 bottom-hole temperature measurements. Future work on—and stemming from—this project will include new core analyses, calculation and interpolation of subsurface temperature gradients, mechanical earth models, geochemical simulations, storage capacity estimation, stratigraphic column generation and correlation, and construction of subsurface maps. Finally, this work will help to inspire and facilitate subsurface data compilation and curation beyond Sweetwater County, Wyoming.

42 ENGINEERING

Constraints on anomalous Higgs boson couplings from its production and decay using the WW channel in proton–proton collisions at $\sqrt{s} = 13~\text {TeV}$

A study of the anomalous couplings of the Higgs boson to vector bosons, including ${\textit{CP}}$-violation effects, has been conducted using its production and decay in the WW channel. This analysis is performed on proton–proton collision data collected with the CMS detector at the CERN LHC during 2016–2018 at a center-of-mass energy of 13 TeV, and corresponds to an integrated luminosity of 138$\,\text {fb}^{-1}$. The different-flavor dilepton $({\textrm{e}} {{\upmu }})$ final state is analyzed, with dedicated categories targeting gluon fusion, electroweak vector boson fusion, and associated production with a W or Z boson. Kinematic information from associated jets is combined using matrix element techniques to increase the sensitivity to anomalous effects at the production vertex. A simultaneous measurement of four Higgs boson couplings to electroweak vector bosons is performed in the framework of a standard model effective field theory. All measurements are consistent with the expectations for the standard model Higgs boson and constraints are set on the fractional contribution of the anomalous couplings to the Higgs boson production cross section.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Hayabusa Recovery, Curation and Preliminary Sample Analysis: Lessons Learned from Recent Sample Return Mission

I describe lessons learned from my participation on the Hayabusa Mission, which returned regolith grains from asteroid Itokawa in 2010 [1], comparing this with the recently returned Stardust Spacecraft, which sampled the Jupiter Family comet Wild 2. Spacecraft Recovery Operations: The mission Science and Curation teams must actively participate in planning, testing and implementing spacecraft recovery operations. The crash of the Genesis spacecraft underscored the importance of thinking through multiple contingency scenarios and practicing field recovery for these potential circumstances. Having the contingency supplies on-hand was critical, and at least one full year of planning for Stardust and Hayabusa recovery operations was necessary. Care must be taken to coordinate recovery operations with local organizations and inform relevant government bodies well in advance. Recovery plans for both Stardust and Hayabusa had to be adjusted for unexpectedly wet landing site conditions. Documentation of every step of spacecraft recovery and deintegration was necessary, and collection and analysis of launch and landing site soils was critical. We found the operation of the Woomera Text Range (South Australia) to be excellent in the case of Hayabusa, and in many respects this site is superior to the Utah Test and Training Range (used for Stardust) in the USA. Recovery operations for all recovered spacecraft suffered from the lack of a hermetic seal for the samples. Mission engineers should be pushed to provide hermetic seals for returned samples. Sample Curation Issues: More than two full years were required to prepare curation facilities for Stardust and Hayabusa. Despite this seemingly adequate lead time, major changes to curation procedures were required once the actual state of the returned samples became apparent. Sample databases must be fully implemented before sample return for Stardust we did not adequately think through all of the possible sub sampling and analytical activities before settling on a database design - Hayabusa has done a better job of this. Also, analysis teams must not be permitted to devise their own sample naming schemes. The sample handling and storage facilities for Hayabusa are the finest that exist, and we are now modifying Stardust curation to take advantage of the Hayabusa facilities. Remote storage of a sample subset is desirable. Preliminary Examination (PE) of Samples: There must be some determination of the state and quantity of the returned samples, to provide a necessary guide to persons requesting samples and oversight committees tasked with sample curation oversight. Hayabusa s sample PE, which is called HASPET, was designed so that late additions to the analysis protocols were possible, as new analytical techniques became available. A small but representative number of recovered grains are being subjected to in-depth characterization. The bulk of the recovered samples are being left untouched, to limit contamination. The HASPET plan takes maximum advantage of the unique strengths of sample return missions

Zolensky, Michael E.

Experiences with Text Mining Large Collections of Unstructured Systems Development Artifacts at JPL

Often repositories of systems engineering artifacts at NASA's Jet Propulsion Laboratory (JPL) are so large and poorly structured that they have outgrown our capability to effectively manually process their contents to extract useful information. Sophisticated text mining methods and tools seem a quick, low-effort approach to automating our limited manual efforts. Our experiences of exploring such methods mainly in three areas including historical risk analysis, defect identification based on requirements analysis, and over-time analysis of system anomalies at JPL, have shown that obtaining useful results requires substantial unanticipated efforts - from preprocessing the data to transforming the output for practical applications. We have not observed any quick 'wins' or realized benefit from short-term effort avoidance through automation in this area. Surprisingly we have realized a number of unexpected long-term benefits from the process of applying text mining to our repositories. This paper elaborates some of these benefits and our important lessons learned from the process of preparing and applying text mining to large unstructured system artifacts at JPL aiming to benefit future TM applications in similar problem domains and also in hope for being extended to broader areas of applications.

text mining

Multicriteria Measures to Assess the Sustainability of Diets: A Systematic Review

Abstract Context Assessing the overall sustainability of a diet is a challenging undertaking requiring a holistic approach capable of addressing the multicriteria nature of this concept. Objective The aim was to identify and summarize the multicriteria measures used to assess the sustainability characteristics of diets reported at the individual level by healthy adults. Data Sources Articles were identified via PubMed, Scopus, and Web of Science. The search strategy consisted of key words and MeSH terms, and was concluded in September 2022, covering references in English, Spanish, and Portuguese. Data Extraction This systematic review followed the PRISMA guidelines. The search identified 5663 references, from which 1794 were duplicates. Two reviewers independently screened the titles and abstracts of each of the 3869 records and the full-text of the 144 references selected. Of these, 7 studies met the inclusion criteria. Data Analysis A total of 6 multicriteria measures were identified: 3 different Sustainable Diet Indices, the Quality Environmental Costs of Diet, the Quality Financial Costs of Diet, and the Environmental Impact of Diet. All of these incorporated a health/nutrition dimension, while the environmental and economic dimensions were the second and the third most integrated, respectively. A sociocultural sustainability dimension was included in only 1 of the measures. Conclusion Despite some methodological concerns in the development and validation process of the identified measures, their inclusion is considered indispensable in assessing the transition towards sustainable diets in future studies. Systematic Review Registration PROSPERO registration no. CRD42022358824.

Rei, Mariana (ORCID:0000000189453708)

A Solid State Zwitterionic Plastic Crystal with High Static Dielectric Constant

The dielectric data in Figure 3, Figure 4, Figure S6 of the published paper was extracted from 2EOIMTSA-BDS-DATA .txt file. This file can be directly opened using a text file editor. It can also be imported to Excel/ Origin for further plotting and analysis. The G' and G'' in Figure 3 of the publihsed paper was plotted from data in file 2EOImTSA-temperature-sweep.xlsx. This file can be directly opend using Excel. The details of DFT simulations mentioned in Figure 2, Figure 7, and Figure S9 of the published paper are included in the DFT.zip file.

Huang, Zitan [Pennsylvania State University]