Search NASA⌕ Search

SEARCH · Search NASA

Results for “knowledge extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

New developments of a knowledge based system (VEG) for inferring vegetation characteristics

An extraction technique for inferring physical and biological surface properties of vegetation using nadir and/or directional reflectance data as input has been developed. A knowledge-based system (VEG) accepts spectral data of an unknown target as input, determines the best strategy for inferring the desired vegetation characteristic, applies the strategy to the target data, and provides a rigorous estimate of the accuracy of the inference. Progress in developing the system is presented. VEG combines methods from remote sensing and artificial intelligence, and integrates input spectral measurements with diverse knowledge bases. VEG has been developed to (1) infer spectral hemispherical reflectance from any combination of nadir and/or off-nadir view angles; (2) test and develop new extraction techniques on an internal spectral database; (3) browse, plot, or analyze directional reflectance data in the system's spectral database; (4) discriminate between user-defined vegetation classes using spectral and directional reflectance relationships; and (5) infer unknown view angles from known view angles (known as view angle extension).

Kimes, D. S.↗

Improvement and generalization of ABCD method with Bayesian inference

To find New Physics or to refine our knowledge of the Standard Model at the LHC is an enterprise that involves many factors, such as the capabilities and the performance of the accelerator and detectors, the use and exploitation of the available information, the design of search strategies and observables, as well as the proposal of new models. We focus on the use of the information and pour our effort in re-thinking the usual data-driven ABCD method to improve it and to generalize it using Bayesian Machine Learning techniques and tools. We propose that a dataset consisting of a signal and many backgrounds is well described through a mixture model. Signal, backgrounds and their relative fractions in the sample can be well extracted by exploiting the prior knowledge and the dependence between the different observables at the event-by-event level with Bayesian tools. We show how, in contrast to the ABCD method, one can take advantage of understanding some properties of the different backgrounds and of having more than two independent observables to measure in each event. In addition, instead of regions defined through hard cuts, the Bayesian framework uses the information of continuous distribution to obtain soft-assignments of the events which are statistically more robust. To compare both methods we use a toy problem inspired by pp\to hh\to b\bar b b \bar b p p → h h → b b ‾ b b ‾ , selecting a reduced and simplified number of processes and analysing the flavor of the four jets and the invariant mass of the jet-pairs, modeled with simplified distributions. Taking advantage of all this information, and starting from a combination of biased and agnostic priors, leads us to a very good posterior once we use the Bayesian framework to exploit the data and the mutual information of the observables at the event-by-event level. We show how, in this simplified model, the Bayesian framework outperforms the ABCD method sensitivity in obtaining the signal fraction in scenarios with 1% and 0.5% true signal fractions in the dataset. We also show that the method is robust against the absence of signal. We discuss potential prospects for taking this Bayesian data-driven paradigm into more realistic scenarios.

Alvarez, Ezequiel↗

Gaseous By-Product of Thermal Vacuum Processing of Lunar Highland Simulant

Scientists at Kennedy Space Center are advancing technologies to achieve oxygen extraction on the Lunar surface. One such technology is Molten Regolith Electrolysis (MRE) in which regolith simulant is melted under high temperatures to perform oxygen extraction through electrolysis of the melt pool. A study was conducted to understand the molten formation and proper-ties of Lunar Highlands Simulants (LHS-1) under high vacuum environments ranging between 10-6 Torr. These studies include investigating the regolith melt behavior and by-product gas analysis systems to prepare for pilot plant development and operations. The off-gassing rates and pressure build-up, gaseous compounds, and material compatibility may be a risk to MRE systems. The KSC findings are reported here for inclusion into future mis-sion architecture and operational planning. The four stages of the experiment were vacuum, resistive heating, regolith melt and gas detection. The regolith mass of 70 g was melted in a 40 cm tall x 50 cm diameter vacuum chamber at temperature increase rate 47 ℃ /min and ramp rate of 1 ampere/3 min up to an 18-amp maximum. The test is conducted for about 60 minutes and the constituent gases produced during heating of regolith is monitored with a re-sidual gas analyzer. During the molten formation, primary off-gassing volatiles included water vapor (18 amu), a peak at 44 amu (FeO2 – expected), and atomic oxygen (16 amu), among a series of other compounds that could create compatibility issues such as magnesium, chlorine, and silicon oxide. These results add to the knowledge required for successful oxygen extraction on the moon.

Molten Regolith Electrolysis↗

Agentic traffic intelligence: Augmented human-in-the-loop scenario generation for microscopic traffic simulation

Traditional microscopic traffic simulation generation often relies on static datasets and manual design, limiting its ability to simulate complex conditions easily. This paper presents a novel framework, Agentic Traffic Intelligence, which combines human approval large language models (LLMs), the Real-Twin tool, and multi-agent systems to perform realistic microscopic traffic simulation scenario generation. The proposed framework incorporates human-in-the-loop (HIL) control, retrieval-augmented generation (RAG), and multi-agent control mechanisms. HIL mechanisms are used to guide multiple LLMs focused on attributes for microscopic simulation generation and to improve the interpretability and transparency of LLM execution for users. RAG enhances context extraction by dynamically integrating external knowledge sources for traffic scenario generation foundations. A multi-agent architecture with supervisory control coordinates the interaction of simulation components, including traffic simulators, control logic, and calibration tools. This enables the synthesis of simulation-ready scenarios that reflect dynamic demand profiles and behavior controls. Furthermore, the framework fuses multisource traffic data with unstructured context and supports iterative refinement through interactive user feedback. Validated through microscopic simulation using Simulation of Urban Mobility, the generated scenarios demonstrate high-fidelity network generation with inflow and turn movement and behavioral calibration, offering a robust and efficient tool for stress-testing and optimizing urban mobility systems.

Hierarchical multi-agent control↗

Shuttle entry trajectory reconstruction using inflight accelerometer and gyro measurements

An error analysis has been made of a Shuttle postflight entry trajectory reconstruction process to obtain trajectory state estimation errors and to assess the impact of these errors on Shuttle aerodynamic force coefficient extraction. In this analysis, the entry trajectory is assumed to be reconstructed via numerical integration of onboard accelerometer and gyro measurements and constrained to satisfy ground-based radio tracking. The trajectory state estimation errors are calculated using a Kalman-Schmidt sequential filter assuming various measurement error models and combinations of ground-based tracking. The resultant trajectory estimation errors are analyzed in a simplified perturbation process to establish the accuracy to which postflight aerodynamic force coefficients can be determined. Results are presented which show that the principal error sources affecting the trajectory reconstruction and thus the force coefficient extraction, assuming perfect atmospheric density knowledge, are the accelerometer and gyro resolution, acceleration-sensitive gyro drifts, and the alignment uncertainties associated with integration on the Shuttle.

Compton, H. R.↗

An industrial perspective of the LANDSAT opportunity

The feasibility of enhancing LANDSAT products to provide the greatest usability low cost data possible can be determined through government sponsorship and finance of one or more task forces composed of a critical number of experts in multiple disciplines from many industries and academia. The synergism of multiple minds addressing singular problems without the creation of permanent or perpetual structures must yield output in the form of implementable specifications, even if presented as alternatives. Changes are needed within the spacecraft in order to account for Sun angle changes. The use of pointing accuracy to make geometric corrections (and possible radiometric corrections, is needed more than onboard data reduction and information extraction, which assume a proper knowledge of application and reduce potential utilization. Multilinear arrays need to be investigated and methods for sensor calibration and for determining the effects of atmospheric inversion, as well as the best way to back out the modulation transfer function must be determined.

Williams, B. F.↗

Transferring Knowledge from Observations and Models to Decision Makers: An Overview and Challenges

Over the last 25 years, a tremendous progress has been made in the Earth science space-based remote sensing observations, technologies and algorithms. Such advancements have improved the predictability by providing lead-time and accuracy of forecast in weather, climate, natural hazards, and natural resources. It has further reduced or bounded the overall uncertainties by partially improving our understanding of planet Earth as an integrated system that is governed by non-linear and chaotic behavior. Many countries such US, European Community, Japan, China and others have invested billions of dollars in developing and launching space-based assets in the low earth (LEO) and geostationary (GEO) orbits. However, the wealth of this scientific knowledge that has potential of extracting monumental socio-economic benefits from such large investments have been slow in reaching to public and decision makers. For instance, there are a number of areas such as energy forecasting, aviation safety, agricultural competitiveness, disaster management, security, air quality and public health can directly take advantage. Nevertheless, we all live in a global economy that depends on access to the best available Earth Science information for all inhabitants of this planet. This paper surveys and examines a number such applications in terms of their architecture, maturity and economic applicability as they apply to the societal needs. A detailed analysis is also presented of various challenges and issues that pertain to a number of areas such as: (1) difficulties in making a speedy transition of data and information from observations and models to relevant Decision Support Systems (DSS) or tools, (2) data and models inter-operability issues, (3) limitations of spatial, spectral and temporal resolution, (4) communication limitations as dictated by the availability of image processing and data compression techniques. Additionally, the most critical element amongst all is the organizational and management boundaries that must be resolved at local, state, national and international levels to implement and realize free flow of such vital information. This paper also makes attempts to address this topic and discuss possible approaches to deal with this quandary.

Habib, Shahid↗

Utilizing Earth Observations for Societal Issues

Over the last four decades a tremendous progress has been made in the Earth science space-based remote sensing observations, technologies and algorithms. Such advancements have improved the predictability by providing lead-time and accuracy of forecast in weather, climate, natural hazards, and natural resources. It has further reduced or bounded the overall uncertainties by partially improving our understanding of planet Earth as an integrated system that is governed by non-linear and chaotic behavior. Many countries such as the US, European Community, Japan, China, Russia, India has and others have invested billions of dollars in developing and launching space-based assets in the low earth (LEO) and geostationary (GEO) orbits. However, the wealth of this scientific knowledge that has potential of extracting monumental socio-economic benefits from such large investments have been slow in reaching the public and decision makers. For instance, there are a number of areas such as water resources and availability, energy forecasting, aviation safety, agricultural competitiveness, disaster management, air quality and public health, which can directly take advantage. Nevertheless, we all live in a global economy that depends on access to the best available Earth Science information for all inhabitants of this planet. This presentation discusses a process to transition Earth science data and products for societal needs including NASA's experience in achieving such objectives. It is important to mention that there are many challenges and issues that pertain to a number of areas such as: (1) difficulties in making a speedy transition of data and information from observations and models to relevant Decision Support Systems (DSS) or tools, (2) data and models inter-operability issues, (3) limitations of spatial, spectral and temporal resolution, (4) communication limitations as dictated by the availability of image processing and data compression techniques. Additionally, the most critical element amongst all is the organizational and management boundaries that must be resolved at local, state, national and international levels to implement and realize free flow of such vital information.

Habib, Shahid↗

Elucidating the Interfacial Barriers in Lanthanide Back-Extraction: From Water to Oil and Back Again

Recovery of critical rare earth elements from complex mixtures has long been realized via solvent extraction, where ions in an aqueous phase are separated into an organic phase using amphiphilic ligands. While a great deal of effort has been placed on understanding this forward reaction, substantial knowledge gaps in the back-extraction process remain. This includes the mechanism of interfacial dissociation and transport back into a highly acidic aqueous phase for further processing. In this work, we connect back-extraction kinetics made in realistic solvent extraction systems to salient interfacial chemistry and structure that represent bottlenecks in the back-extraction of lanthanide ions. We show that the interface between the two liquid phases varies dramatically based on the composition of both phases. Water stretching signals are shown to report on the population of lingering interfacial complexes and are thus used as a reporter of competitive adsorption from excess free ligands in solution for limited interfacial vacancies. We show that excess free ligands, often used to improve forward extractions, set up interfacial blockades inhibiting back-extraction both kinetically and thermodynamically. In conclusion, this insight opens up avenues to tune interfacial properties to facilitate a more dynamic, exchangeable interface to speed up back-extractions while using less energy intensive chemical swings.

Interfaces↗

Semantic Analysis of Email Using Domain Ontologies and WordNet

The problem of capturing and accessing knowledge in paper form has been supplanted by a problem of providing structure to vast amounts of electronic information. Systems that can construct semantic links for natural language documents like email messages automatically will be a crucial element of semantic email tools. We have designed an information extraction process that can leverage the knowledge already contained in an existing semantic web, recognizing references in email to existing nodes in a network of ontology instances by using linguistic knowledge and knowledge of the structure of the semantic web. We developed a heuristic score that uses several forms of evidence to detect references in email to existing nodes in the Semanticorganizer repository's network. While these scores cannot directly support automated probabilistic inference, they can be used to rank nodes by relevance and link those deemed most relevant to email messages.

Berrios, Daniel C.↗

Expert Seeker: A People-Finder Knowledge Management System

The first objective for this report was to perform a comprehensive research of industry models currently being used for similar purposes, in order to provide the Center with ideas of what is being done in area by private companies and government agencies. The second objective was to evaluate the use of taxonomies or ontologies to describe and catalog the areas of expertise at GSFC. The creation of a knowledge taxonomy is necessary for information extraction in order for The Expert Seeker to adequately search and find experts in a particular area of expertise. The requirements to develop a taxonomy are: provide minimal descriptive text; have the appropriate level of abstration; facilitate browsing; ease of use and speed of data entry are critical for success; customized to the organization and its culture; extent of knowledge areas; expandable, so new skills could be develop; could be complemented with free text fields to allow users the option to describe their knowledge in detail.

Becerra-Fernandez, Irma↗

Calibration of Viking imaging system pointing, image extraction, and optical navigation measure

Pointing control and knowledge accuracy of Viking Orbiter science instruments is controlled by the scan platform. Calibration of the scan platform and the imaging system was accomplished through mathematical models. The calibration procedure and results obtained for the two Viking spacecraft are described. Included are both ground and in-flight scan platform calibrations, and the additional calibrations unique to optical navigation.

Breckenridge, W. G.↗

Robustness of topological persistence in knowledge distillation for wearable sensor data

Topological data analysis (TDA) has shown great success in various applications involving wearable sensor data. However, there are difficulties in leveraging topological features in machine learning and wearable sensors because of the large time consumption and computational resources required to extract the features. To address this problem, knowledge distillation (KD) is utilized to generate a small model and accommodate topological features with persistence image (PI) representations from the raw time series data. Deploying topological knowledge in KD enables the student to achieve better performance compared to the one trained solely on raw time series data. However, it is not yet known if there are coherent characteristics for topological features in PI, which can aid in improving the performance during KD. In this paper, we investigate the suitability and challenges of utilizing topological features in KD for wearable sensor data, thereby contributing to the advancement of the field. Our study explores the impact of transferred topological features by comparing the Teacher-to-Student framework with Multiple Teachers-to-Student where teachers utilize both time series data and persistence images obtained by TDA as inputs. Additionally, we conduct a rigorous examination of topological knowledge effects by testing under various corruptions, knowledge types, and learning strategies in the context of human activity recognition tasks. Our analysis of topological features in KD presents the optimal strategy for incorporating these features. This study includes datasets of varying scales, window lengths, and activity classes, providing a comprehensive evaluation. Our results demonstrate that leveraging topological features in KD to enhance performance across databases.

97 MATHEMATICS AND COMPUTING↗

Research relative to a model of the Orion nebula

Research basically has been directed along two avenues. First of all, there is a long-standing interest in modeling H II regions in order to understand the physical processes involved and to extract important astrophysical information. This includes knowledge of chemical elemental abundances and properties of the exciting stars, such as their far ultraviolet spectrum. The Orion Nebula is a prime candidate to study because it is nearby, bright, the extinction is not large, and its appearance is reasonably circular. Such a shape is consistent with a geometrical structure that is spherically symmetric (1-D) or one that is axisymmetric (2-D) seen nearly face-on. Previously all detailed modeling of the ionization and thermal structure of H II regions has been confined to 1-D because of computational complexity. There is now the capability to treat the 2-D case with much of the level of physical sophistication as the 1-D case. The interpretation of the spectra of gaseous nebulae in terms of underlying physical processes requires measurements of line intensities at different points in the object and over as wide a set of excitation and ionization conditions as practical.Observations of nebulae have been made for many years in the optical and radio but only recently in the infrared and ultraviolet. These relatively new windows allow observations of lines of ionization states not available in the radio or optical--providing a much more complete set of known quantities to undertake a meaningful modeling effort.

Zuckerman, B. M.↗

A petabyte size electronic library using the N-Gram memory engine

A model library containing petabytes of data is proposed by Triada, Ltd., Ann Arbor, Michigan. The library uses the newly patented N-Gram Memory Engine (Neurex), for storage, compression, and retrieval. Neurex splits data into two parts: a hierarchical network of associative memories that store 'information' from data and a permutation operator that preserves sequence. Neurex is expected to offer four advantages in mass storage systems. Neurex representations are dense, fully reversible, hence less expensive to store. Neurex becomes exponentially more stable with increasing data flow; thus its contents and the inverting algorithm may be mass produced for low cost distribution. Only a small permutation operator would be recalled from the library to recover data. Neurex may be enhanced to recall patterns using a partial pattern. Neurex nodes are measures of their pattern. Researchers might use nodes in statistical models to avoid costly sorting and counting procedures. Neurex subsumes a theory of learning and memory that the author believes extends information theory. Its first axiom is a symmetry principle: learning creates memory and memory evidences learning. The theory treats an information store that evolves from a null state to stationarity. A Neurex extracts information data without a priori knowledge; i.e., unlike neural networks, neither feedback nor training is required. The model consists of an energetically conservative field of uniformly distributed events with variable spatial and temporal scale, and an observer walking randomly through this field. A bank of band limited transducers (an 'eye'), each transducer in a bank being tuned to a sub-band, outputs signals upon registering events. Output signals are 'observed' by another transducer bank (a mid-brain), except the band limit of the second bank is narrower than the band limit of the first bank. The banks are arrayed as n 'levels' or 'time domains, td.' The banks are the hierarchical network (a cortex) and transducers are (associative) memories. A model Neurex was built and studied. Data were 50 MB to 10 GB samples of text, data base, and images: black/white, grey scale, and high resolution in several spectral bands. Memories at td, S(m(sub td)), were plotted against outputs of memories at td-1. S(m(sub td)) was Boltzman distributed, and memory frequencies exhibited self-organized criticality (SOC); i.e., 'l/f(sup beta)' after long exposures to data. Whereas output signals from level n may be encoded with B(sub output) = O(-log(2)f(sup beta)) bits, and input data encoded with B(sub input) = O((S(td)/S(td-1))(sup n)), B(sup output)/B(sub input) is much less than 1 always, the Neurex determines a canonical code for data and it is a lossless data compressor. Further tests are underway to confirm these results with more data types and larger samples.

Bugajski, Joseph M.↗

NASA's online machine aided indexing system

This report describes the NASA Lexical Dictionary, a machine aided indexing system used online at the National Aeronautics and Space Administration's Center for Aerospace Information (CASI). This system is comprised of a text processor that is based on the computational, non-syntactic analysis of input text, and an extensive 'knowledge base' that serves to recognize and translate text-extracted concepts. The structure and function of the various NLD system components are described in detail. Methods used for the development of the knowledge base are discussed. Particular attention is given to a statistically-based text analysis program that provides the knowledge base developer with a list of concept-specific phrases extracted from large textual corpora. Production and quality benefits resulting from the integration of machine aided indexing at CASI are discussed along with a number of secondary applications of NLD-derived systems including on-line spell checking and machine aided lexicography.

Silvester, June P.↗

Knowledge Discovery and Data Mining: An Overview

The process of knowledge discovery and data mining is the process of information extraction from very large databases. Its importance is described along with several techniques and considerations for selecting the most appropriate technique for extracting information from a particular data set.

data mining knowledge discovery data search↗

Linking Threat Agents to Targeted Organizations: A Pipeline for Enhanced Cybersecurity Risk Metrics

In this study, we present a methodology leveraging Large Language Models (LLMs) to transform Cybersecurity Threat Intelligence (CTI) narratives into actionable insights for individual organizations. Our approach automates the extraction of machine-readable adversary SKRAM (Skills, Knowledge, Resources, Authorities, and Motivation) attributes from open-source reports, extending LLM utility beyond typical interactions. This innovation enables precise, automated assessments of cybersecurity risks posed by various adversaries. Using a chain-of-thought and multi-shot prompting strategy, our methodology advances the automation of cybersecurity feature extraction for new machine-learning models that predict the risk of adversary targeting. This approach is refined using a substantial dataset of over 150 analyst-validated threat reports and synthetic organizational data from 900 companies. Here, by bootstrapping the training data with a rule-based heuristic over synthetic data, we have developed a high-accuracy machine-learning model that allows entities to dynamically prioritize threats and defensive actions.

Cyber Threat Intelligence↗