Search NASA⌕ Search

SEARCH · Search NASA

Results for “Information systems →Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

SPICE: A Geometry Information System Supporting Planetary Mapping, Remote Sensing and Data Mining

SPICE is an information system providing space scientists ready access to a wide assortment of space geometry useful in planning science observations and analyzing the instrument data returned therefrom. The system includes software used to compute many derived parameters such as altitude, LAT/LON and lighting angles, and software able to find when user-specified geometric conditions are obtained. While not a formal standard, it has achieved widespread use in the worldwide planetary science community

planetary science investigations↗

Scalable Pattern Matching in Metadata Graphs via Constraint Checking

Pattern matching is a fundamental tool for answering complex graph queries. Unfortunately, existing solutions have limited capabilities: They do not scale to process large graphs and/or support only a restricted set of search templates or usage scenarios. Moreover, the algorithms at the core of the existing techniques are not suitable for today’s graph processing infrastructures relying on horizontal scalability and shared-nothing clusters, as most of these algorithms are inherently sequential and difficult to parallelize. In this article we present an algorithmic pipeline that bases pattern matching on constraint checking. The key intuition is that each vertex and edge participating in a match has to meet a set of constraints implicitly specified by the search template. These constraints can be verified independently and typically are less expensive to compute than searching the full template. The pipeline we propose generates these constraints and iterates over them to eliminate all the vertices and edges that do not participate in any match, thus reducing the background graph to a subgraph that is the union of all template matches—the complete set of all vertices and edges that participate in at least one match. Additional analysis can be performed on this annotated, reduced graph, such as full match enumeration, match counting, or computing vertex/edge centrality. Furthermore, a vertex-centric formulation for constraint checking algorithms exists, and this makes it possible to harness existing high-performance, vertex-centric graph processing frameworks. This technique (i) enables highly scalable pattern matching in metadata (labeled) graphs; (ii) supports arbitrary patterns with 100% precision; (iii) enables tradeoffs between precision and time-to-solution, while always selects all vertices and edges that participate in matches, thus offering 100% recall; and (iv) supports a set of popular data analytics scenarios. We implement our approach on top of HavoqGT, an open-source asynchronous graph processing framework, and demonstrate its advantages through strong and weak scaling experiments on massive scale real-world (up to 257 billion edges) and synthetic (up to 4.4 trillion edges) labeled graphs, respectively, and at scales (1,024 nodes / 36,864 cores), orders of magnitude larger than used in the past for similar problems. This article serves two purposes: First, it synthesises the knowledge accumulated during a long-term project. Second, it presents new system features, usage scenarios, optimizations, and comparisons with related work that strengthen the confidence that pattern matching based on iterative pruning via constraint checking is an effective and scalable approach in practice. The new contributions include the following: (i) We demonstrate the ability of the constraint checking approach to efficiently support two additional search scenarios that often emerge in practice, interactive incremental search and exploratory search. (ii) We empirically compare our solution with two additional state-of-the-art systems, Arabsque and TriAD. (iii) We show the ability of our solution to accommodate a more diverse range of datasets with varying properties, e.g., scale, skewness, label distribution, and match frequency. (iv) We introduce or extend a number of system features (e.g., work aggregation, load balancing, and the ability to cap the generated traffic) and design optimizations and demonstrate their advantages with respect to improving performance and scalability. (v) We present bottleneck analysis and insights into artifacts that influence performance. (vi) We present a theoretical complexity argument that motivates the performance gains we observe.

97 MATHEMATICS AND COMPUTING↗

Reusing Information Management Services for Recommended Decadal Study Missions to Facilitate Aerosol and Cloud Studies

NASA Earth Sciences Division (ESD) has made great investments in the development and maintenance of data management systems and information technologies, to maximize the use of NASA generated Earth science data. With information management system infrastructure in place, mature and operational, very small delta costs are required to fully support data archival, processing, and data support services required by the recommended Decadal Study missions. This presentation describes the services and capabilities of the Goddard Space Flight Center (GSFC) Earth Sciences Data and Information Services Center (GES DISC) and the reusability for these future missions. The GES DISC has developed a series of modular, reusable data management components currently in use. They include data archive and distribution (Simple, Scalable, Script-based, Science [S4] Product Archive aka S4PA), data processing (S4 Processor for Measurements aka S4PM), data search (Mirador), data browse, visualization, and analysis (Giovanni), and data mining services. Information management system components are based on atmospheric scientist inputs. Large development and maintenance cost savings can be realized through their reuse in future missions.

Kempler, Steve↗

Comparative Assessment of Data-driven Process Models in Health Information Technology

Process mining for conformance analysis consists of comparing a reference process model against a data-driven process model generated via log files from information technology systems. However, in the absence of a complete reference process model, we found no suggested approaches in the literature to address the need for evaluating process conformance among different healthcare facilities to assess standardization of care. Our goal is to find similarities and dissimilarities in data-driven process models among US Veterans Health Administration (VHA) facilities that can be indicative of patient safety issues. Our hypothesis was that the analysis would not produce statistically significant differences in outcome. We present a unique implementation of conformance analysis in process mining that consists of combining process mining, process mapping and statistical metrics. We illustrate our approach by applying it to the analysis of two clinical radiology order process models generated from healthcare data provided by two similar facilities in the VHA. The comparative assessment showed that about 70% of the orders completed successfully and 30% were not completed due to policy and duplications. Our analysis found a good statistical correlation between both facilities, as the Spearman’s correlation coefficient between facilities for the frequency of cases per total hours was 0.87879, for the frequency of cases by state transition was 0.79702 and for the throughput time per state transition was 0.63582. Additional statistical analyses using the Mann-Whitney U test and the root mean square error both produced values that were not significant. The foregoing approach validated our hypothesis by demonstrating a good statistical correlation of data describing the flow of clinical radiology orders absent a credible reference model. Finding good agreement between both facilities was important in confirming that the clinical orders flow in a similar manner, suggesting standardization of care.

97 MATHEMATICS AND COMPUTING↗

Advances in scientific literature mining for interpreting materials characterization

Abstract Using synchrotron light sources, such as the National Synchrotron Light Source II at Brookhaven National Laboratory, scientists in fields as diverse as physics, biology, and materials science, identify the atomic structure, chemical composition, or other important properties of varied specimens. x-ray spectroscopy from light sources is particularly valuable for materials research with vast information available about reference spectra in the scientific literature. However, as the technique is applicable to many science domains, searching for information about select x-ray spectroscopy spectra is impeded by the sheer number of publications. Moreover, useful information about the context of an experiment or figures presented in papers can be buried among the details, which takes time to assess. This work presents a scientific literature mining system that supports data acquisition, information extraction, and user interaction for referencing x-ray spectra identification and spectral interpretation. The goal is to provide efficient access to useful spectral data to researchers who may spend only a few days at a synchrotron light source. With this system, users browse a classification tree for papers arranged according to x-ray spectroscopic methods, chemical elements, and x-ray absorption spectroscopy edges. Relevant figures are extracted with sentences from the paper that explain them, known as ‘figure explanatory text.’ Notably, this system focuses on semantic aspects (logical analysis) to find figure explanatory text using deep contextualized word embeddings techniques and contains an interface to obtain labeled data from domain experts that is used to evaluate and improve the model.

Park, Gilchan (ORCID:0000000201536646)↗

Robust Informatics Infrastructure Required For ICME: Combining Virtual and Experimental Data

With the increased emphasis on reducing the cost and time to market of new materials, the need for robust automated materials information management system(s) enabling sophisticated data mining tools is increasing, as evidenced by the emphasis on Integrated Computational Materials Engineering (ICME) and the recent establishment of the Materials Genome Initiative (MGI). This need is also fueled by the demands for higher efficiency in material testing; consistency, quality and traceability of data; product design; engineering analysis; as well as control of access to proprietary or sensitive information. Further, the use of increasingly sophisticated nonlinear, anisotropic and or multi-scale models requires both the processing of large volumes of test data and complex materials data necessary to establish processing-microstructure-property-performance relationships. Fortunately, material information management systems have kept pace with the growing user demands and evolved to enable: (i) the capture of both point wise data and full spectra of raw data curves, (ii) data management functions such as access, version, and quality controls;(iii) a wide range of data import, export and analysis capabilities; (iv) data pedigree traceability mechanisms; (v) data searching, reporting and viewing tools; and (vi) access to the information via a wide range of interfaces. This paper discusses key principles for the development of a robust materials information management system to enable the connections at various length scales to be made between experimental data and corresponding multiscale modeling toolsets to enable ICME. In particular, NASA Glenn's efforts towards establishing such a database for capturing constitutive modeling behavior for both monolithic and composites materials

Mutli-scale models↗

Science Breakthroughs 2030. Final report

Agriculture is a fundamental societal activity, characterized by many different landscapes, crops, markets, and participants. Food, agricultural, and biofuels products are central to the daily life of all citizens, though most do not recognize the fragility of the environment that brings forth this abundance. As is noted in the 2012 report from the President's Council of Advisors on Science and Technology, Agricultural Preparedness and the United States Agricultural Research Enterprise (PCAST, 2012) the food and agricultural system faces constant challenges in: Managing new pests, pathogens, and invasive plants. Increasing the efficiency of water use. Growing food in a changing climate. Reducing the environmental footprint of agriculture. Managing the production of bioenergy. Producing safe and nutritious food. Assisting with global food security and maintaining abundant yields. Science Breakthroughs 2030 was organized to identify the most compelling research directions in food and agriculture, in particular those empowered by the application of insights and tools from disciplines of science and engineering not typically associated with food and agricultural research. A committee appointed by the Chairman of the National Research Council explored ideas for research directions with input from the scientific community, with the objective of producing a report describing ambitious and achievable scientific pathways to address major problems and create new opportunities in food and agriculture. Following numerous meetings, a jamboree, and town hall, the appointed committee prepared a report that has subsequently become a reference for federal agencies supporting research in the food and agricultural space. It highlights five key areas for research investment with broad application across food and agriculture: integrated systems research; sensor development; data mining and information sciences, genomics; and the microbiome.

09 BIOMASS FUELS↗

MCS+: An Efficient Algorithm for Crawling the Community Structure in Multiplex Networks

In this article, we consider the problem of crawling a multiplex network to identify the community structure of a layer-of-interest. A multiplex network is one where there are multiple types of relationships between the nodes. In many multiplex networks, some layers might be easier to explore (in terms of time, money etc.). We propose MCS+, an algorithm that can use the information from the easier to explore layers to help in the exploration of a layer-of-interest that is expensive to explore. We consider the goal of exploration to be generating a sample that is representative of the communities in the complete layer-of-interest. This work has practical applications in areas such as exploration of dark (e.g., criminal) networks, online social networks, biological networks, and so on. For example, in a terrorist network, relationships such as phone records, e-mail records, and so on are easier to collect; in contrast, data on the face-to-face communications are much harder to collect, but also potentially more valuable. We perform extensive experimental evaluations on real-world networks, and we observe that MCS+ consistently outperforms the best baseline—the similarity of the sample that MCS+ generates to the real network is up to three times that of the best baseline in some networks. We also perform theoretical and experimental evaluations on the scalability of MCS+ to network properties, and find that it scales well with the budget, number of layers in the multiplex network, and the average degree in the original network.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Development of sensitized pick coal interface detector system

One approach for detection of the coal interface is measurement of pick cutting loads and shock through the use of pick strain gage load cells and accelerometers. The cutting drum of a long wall mining machine contains a number of cutting picks. In order to measure pick loads and shocks, one pick was instrumented and telemetry used to transmit the signals from the drum to an instrument-type tape recorder. A data system using FM telemetry was designed to transfer cutting bit load and shock information from the drum of a longwall shearer coal mining machine to a chassis mounted data recorder. The design of components in the test data system were finalized, the required instruments were assembled, the instrument system was evaluated in an above-ground simulation test, and an underground test series to obtain tape recorded sensor data was conducted.

Burchill, R. F.↗

Earth Science Mining Web Services

To allow scientists further capabilities in the area of data mining and web services, the Goddard Earth Sciences Data and Information Services Center (GES DISC) and researchers at the University of Alabama in Huntsville (UAH) have developed a system to mine data at the source without the need of network transfers. The system has been constructed by linking together several pre-existing technologies: the Simple Scalable Script-based Science Processor for Measurements (S4PM), a processing engine at he GES DISC; the Algorithm Development and Mining (ADaM) system, a data mining toolkit from UAH that can be configured in a variety of ways to create customized mining processes; ActiveBPEL, a workflow execution engine based on BPEL (Business Process Execution Language); XBaya, a graphical workflow composer; and the EOS Clearinghouse (ECHO). XBaya is used to construct an analysis workflow at UAH using ADam components, which are also installed remotely at the GES DISC, wrapped as Web Services. The S4PM processing engine searches ECHO for data using space-time criteria, staging them to cache, allowing the ActiveBPEL engine to remotely orchestras the processing workflow within S4PM. As mining is completed, the output is placed in an FTP holding area for the end user. The goals are to give users control over the data they want to process, while mining data at the data source using the server's resources rather than transferring the full volume over the internet. These diverse technologies have been infused into a functioning, distributed system with only minor changes to the underlying technologies. The key to the infusion is the loosely coupled, Web-Services based architecture: All of the participating components are accessible (one way or another) through (Simple Object Access Protocol) SOAP-based Web Services.

Pham, Long↗

Large Scale Data Mining to Improve Usability of Data: An Intelligent Archive Testbed

Research in certain scientific disciplines - including Earth science, particle physics, and astrophysics - continually faces the challenge that the volume of data needed to perform valid scientific research can at times overwhelm even a sizable research community. The desire to improve utilization of this data gave rise to the Intelligent Archives project, which seeks to make data archives active participants in a knowledge building system capable of discovering events or patterns that represent new information or knowledge. Data mining can automatically discover patterns and events, but it is generally viewed as unsuited for large-scale use in disciplines like Earth science that routinely involve very high data volumes. Dozens of research projects have shown promising uses of data mining in Earth science, but all of these are based on experiments with data subsets of a few gigabytes or less, rather than the terabytes or petabytes typically encountered in operational systems. To bridge this gap, the Intelligent Archives project is establishing a testbed with the goal of demonstrating the use of data mining techniques in an operationally-relevant environment. This paper discusses the goals of the testbed and the design choices surrounding critical issues that arose during testbed implementation.

Ramapriyan, Hampapuram↗

Development of sensitized pick coal interface detector system

One approach for detection of the coal interface is measurement of the pick cutting hoads and shock through the use of pick strain gage load cells and accelerometers. The cutting drum of a long wall mining machine contains a number of cutting picks. In order to measure pick loads and shocks, one pick was instrumented and telementry used to transmit the signals from the drum to an instrument-type tape recorder. A data system using FM telemetry was designed to transfer cutting bit load and shock information from the drum of a longwall shearer coal mining machine to a chassis mounted data recorder.

Burchill, R. F.↗

An Approach for Defining IASMS Services, Functions, and Capabilities

Assuring safety in the NAS with the inclusion of new entrants, such as Advanced Air Mobility (AAM), will require overcoming unique safety challenges that result from combining innovative technologies with novel airspace concepts for moving people and cargo using autonomous vehicles. The focus of the In-time Aviation Safety Management System (IASMS) is to overcome AAM’s safety assurance challenges. The IASMS Concept of Operations (ConOps) describes an interconnected set of services, functions, and capabilities (SFCs) designed to manage operational risks, identify unknown risks, and inform system designs. This paper describes an approach for defining SFCs based on technology trends in research, assessment of known and unknown risks in voluntary safety reports, and causal and contributing factors in aviation accidents and incidents. This approach would identify potential SFCs that further expand the Monitor, Assess, and Mitigate (M-A-M) functionality that represents the enabling framework of the IASMS. Safety implications that will result from integration of AAM in the transformation of the National Airspace System (NAS) were addressed in National Academies committees reports on AAM and IASMS. Development of a ConOps for IASMS was a top recommendation and can be represented as a reframing of safety assurance that builds on real-time alerting such as the Traffic Alert and Collision Avoidance System, and adds the more encompassing in-time temporal parameter in recognition of the different timelines for collecting and assessing safety data for risk mitigations. For example, mining for safety trends from data sources such as the Aviation Safety Information Analysis and Sharing system occurs over a longer time period. Research on AAM operations poses that SFCs can be designed to monitor the safety margin appropriate for AAM including with regards to the distance between current flight parameters and nominal ideal conditions. These in-time comparisons will become more complex as the density of operations increases at least in certain areas and can include planned and actual 4D trajectory, and in-time comparisons having implications on conflict modeling and prediction including expected and actual departure time, fix/waypoint crossing times, and arrival time. These comparisons would be integrated as part of SFCs that redefine and inform new safety margin. An increased safety margin improves management of operational risks while reducing the potential for anomalies. An increased safety margin also has implications for operator confidence in the certainty of its operations and trust in automation. Technology trends in research could be used to refine existing SFCs and define needs for additional SFCs that provide safety improvements to the design and operation of vehicles, airspace design, and operator performance requirements. NASA is developing innovative approaches to safeguard against major accidents and incidents that have occurred in the NAS and those anticipated with the inclusion of envisioned AAM operations. The innovations use operational performance data to monitor, detect, and predict flight variations exceeding safe nominal patterns, such as would be caused by navigational error, severe weather complications, or hijacking of UAS controls. These innovative approaches have high potential to prevent accidents and incidents in the new AAM era. It is anticipated that elements of the innovations will evolve into SFCs for the IASMS. Voluntary safety reports can be monitored to identify anomalies related to design or operational performance risks. Reports could be periodically monitored and assessed for specific topics. Reports might serve as weak signals or precursors indicative of emergent risk such as when combined with other safety information. The architecture could include SFCs that are based on voluntary safety reports recognizing the periodic temporal nature of data analysis. As previously mentioned, aviation accidents with their causal and contributing precursors can inform the need for SFCs in the IASMS. Accidents and incidents at San Francisco International Airport such as Asiana 214 and Air Canada 759 illustrate how combinations of different factors lead to increased risk. These types of precursors and different factors have implications on the types of SFCs that could be needed to monitor and manage different sources and types of design and operational risk. Continuing to assure the safety of AAM as designs and operations gain in complexity can be accompanied by defining SFCs that also increase in complexity. These SFCs can leverage information from findings and recommendations synthesized across on-going research, voluntary safety reports, and accident and incident reports. These SFCs can serve to refine accuracy of algorithms and resolve limitations with current practices. The IASMS architecture represents the framework for the SFCs and their critical role in safety assurance.

In-Time Aviation Safety Management System↗

RHSEG and Subdue: Background and Preliminary Approach for Combining these Technologies for Enhanced Image Data Analysis, Mining and Knowledge Discovery

Under a project recently selected for funding by NASA's Science Mission Directorate under the Applied Information Systems Research (AISR) program, Tilton and Cook will design and implement the integration of the Subdue graph based knowledge discovery system, developed at the University of Texas Arlington and Washington State University, with image segmentation hierarchies produced by the RHSEG software, developed at NASA GSFC, and perform pilot demonstration studies of data analysis, mining and knowledge discovery on NASA data. Subdue represents a method for discovering substructures in structural databases. Subdue is devised for general-purpose automated discovery, concept learning, and hierarchical clustering, with or without domain knowledge. Subdue was developed by Cook and her colleague, Lawrence B. Holder. For Subdue to be effective in finding patterns in imagery data, the data must be abstracted up from the pixel domain. An appropriate abstraction of imagery data is a segmentation hierarchy: a set of several segmentations of the same image at different levels of detail in which the segmentations at coarser levels of detail can be produced from simple merges of regions at finer levels of detail. The RHSEG program, a recursive approximation to a Hierarchical Segmentation approach (HSEG), can produce segmentation hierarchies quickly and effectively for a wide variety of images. RHSEG and HSEG were developed at NASA GSFC by Tilton. In this presentation we provide background on the RHSEG and Subdue technologies and present a preliminary analysis on how RHSEG and Subdue may be combined to enhance image data analysis, mining and knowledge discovery.

Tilton, James C.↗

Autonomous, Context-Sensitive, Task Management Systems and Decision Support Tools II: Contextual Constraints and Information Sources

Recent advances in artificial intelligence, machine learning, data mining and sensor technology have resulted in the availability of a vast amount of digital data and information and the development of advanced automated reasoners. This creates the opportunity for the development of a robust dynamic task manager and decision support tool that is context sensitive and integrates information from a wide array of on-board and off aircraft sourcesa tool that monitors systems and the overall flight situation, anticipates information needs, prioritizes tasks appropriately, keeps pilots well informed, and is nimble and able to adapt to changing circumstances. This is the second of two companion reports exploring issues associated with autonomous, context-sensitive, task management and decision support tools. In the first report, we explored fundamental issues associated with the development of such a system. In this report, we extend this work to focus on two critical aspects of these systems: 1) the constraints and conditions that drive the dynamic prioritization and presentation of data and information to the pilots, and 2) specific data and information to be accessed, monitored, integrated, and displayed in such a system.

context-sensitive↗

The Oklahoma Geographic Information Retrieval System

The Oklahoma Geographic Information Retrieval System (OGIRS) is a highly interactive data entry, storage, manipulation, and display software system for use with geographically referenced data. Although originally developed for a project concerned with coal strip mine reclamation, OGIRS is capable of handling any geographically referenced data for a variety of natural resource management applications. A special effort has been made to integrate remotely sensed data into the information system. The timeliness and synoptic coverage of satellite data are particularly useful attributes for inclusion into the geographic information system.

Blanchard, W. A.↗

Second Eastern Regional Remote Sensing Applications Conference

Participants from state and local governments share experiences in remote sensing applications with one another and with users in the Federal government, universities, and the private sector during technical sessions and forums covering agriculture and forestry; land cover analysis and planning; surface mining and energy; data processing; water quality and the coastal zone; geographic information systems; and user development programs.

Imhoff, M. L.↗

Improve Data Mining and Knowledge Discovery Through the Use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(R) (MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykhian, Gholam Ali↗