Search NASASearch

SEARCH · Search NASA

Results for “Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Enriching the Twitter Stream Increasing Data Mining Yield and Quality Using Machine Learning

Social media data streams are important sources of real-time and historical global information for science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we are exploring the Twitter data stream for its potential in augmenting the validation program of NASA Earth science missions, specifically the Global Precipitation Measurement (GPM) mission. We have implemented a tweet processing infrastructure that outputs classified precipitation tweets. Inputs are "passive" tweets, along with a smaller number of tweets from "active" participants, i.e., those knowingly contributing to our effort. The "active" tweets, presumably of higher quality, enrich the Twitter stream. "Active" sources include data scraped from other social media (e.g., public Facebook posts) and data from existing crowdsourcing programs (e.g., mPING reports). In addition, there is likely relevant precipitation information in images and documents that are the end points of links often included in tweets. Information derived from these "active" sources could then be tweeted into the Twitter stream, thus enriching its quality. The objective of our current work is to mine these tweet­ linked images and documents, using neural networks, to increase the information content and quality related to precipitation. For images, we classified them as either precipitation-related or not. For training and validation, we used images obtained via the Google custom search API. We created two models: (1) by training a simple Convolutional Neural Network and (2) by using transfer learning principles to adapt a pre-trained object recognition model. For documents, both those linked to tweets and the tweet contents, we trained Hierarchical Attention Networks to determine precipitation occurrence, type, and intensity. For training and validation, we used a keyword-filtered tweet data set labelled with ground truth data from Dark Sky (an API to retrieve weather-related labels) and the National Severe Storms Laboratory's Multi­ Radar/Multi-Sensor (MRMS) system. Our results demonstrated the efficacy of our machine learning approaches for enriching the Twitter stream, to derive information potentially useful for validation of earth science satellite data.

Albayrak, Arif

Data Mining of Groundwater to Identify MAGs with Methane, Propane and Toluene Monooxygenases

Whole genome sequencing datasets, involving more than 600 groundwater samples, from nine countries, were analyzed to identify metagenome assembled genomes (MAGs) containing full operons for propane monooxygenase, soluble methane monooxygease, toluene monooxygenase and particulate ammonia/methane monooxygenase. The enzymes encoded by these genes are a focus of interest because of their ability to degrade common groundwater contaminants. Due to the large amount of data, sequence analyses involved more than 80 individual KBase narratives. The approach followed the KBase tutorial called "Metagenome-Assembled Genome Extraction from a Compost Microbiome Enrichment" The generated MAGs were exported from each individual narrative into separate summary KBase narratives for each monooxygenase. Three KBase narratives were generated for particulate ammonia/methane monooxygenase, due to the large number of MAGs identified.

59 BASIC BIOLOGICAL SCIENCES

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science

Grids for Dummies: Featuring Earth Science Data Mining Application

This viewgraph presentation discusses the concept and advantages of linking computers together into data grids, an emerging technology for managing information across institutions, and potential users of data grids. The logistics of access to a grid, including the use of the World Wide Web to access grids, and security concerns are also discussed. The potential usefulness of data grids to the earth science community is also discussed, as well as the Global Grid Forum, and other efforts to establish standards for data grids.

Hinke, Thomas H.

Grist : grid-based data mining for astronomy

The Grist project is developing a grid-technology based system as a research environment for astronomy with massive and complex datasets. This knowledge extraction system will consist of a library of distributed grid services controlled by a workflow system, compliant with standards emerging from the grid computing, web services, and virtual observatory communities. This new technology is being used to find high redshift quasars, study peculiar variable objects, search for transients in real time, and fit SDSS QSO spectra to measure black hole masses. Grist services are also a component of the 'hyperatlas' project to serve high-resolution multi-wavelength imagery over the Internet. In support of these science and outreach objectives, the Grist framework will provide the enabling fabric to tie together distributed grid services in the areas of data access, federation, mining, subsetting, source extraction, image mosaicking, statistics, and visualization.

grid computing

Monitoring Bone Health after Spaceflight: Data Mining to Support an Epidemiological Analysis of Age-related Bone Loss in Astronauts

Through the epidemiological analysis of bone data, HRP is seeking evidence as to whether the prolonged exposure to microgravity of low earth orbit predisposes crewmembers to an earlier onset of osteoporosis. While this collaborative Epidemiological Project may be currently limited by the number of ISS persons providing relevant spaceflight medical data, a positive note is that it compares medical data of astronauts to data of an age-matched (not elderly) population that is followed longitudinally with similar technologies. The inclusion of data from non-ISS and non-NASA crewmembers is also being pursued. The ultimate goal of this study is to provide critical information for NASA to understand the impact of low physical or minimal weight-bearing activity on the aging process as well as to direct its development of countermeasures and rehabilitation programs to influence skeletal recovery. However, in order to optimize these results NASA needs to better define the requirements for long term monitoring and encourage both active and retired astronauts to contribute to a legacy of data that will define human health risks in space.

Baker, K. S,

TRMM Data Mining Service at the Goddard Earth Sciences (GES) DISC DAAC Tropical Rainfall Measuring Mission (TRMM)

TRMM has acquired more than four years of data since its launch in November 1997. All TRMM standard products are processed by the TRMM Science Data and Information System (TSDIS) and archived and distributed to general users by the GES DAAC. Table 1 shows the total archive and distribution as of February 28, 2002. The Utilization Ratio (UR), defined as the ratio of the number of distributed files to the number of archived files, of the TRMM standard products has been steadily increasing since 1998 and is currently at 6.98.

Source record

SPICE: A Geometry Information System Supporting Planetary Mapping, Remote Sensing and Data Mining

SPICE is an information system providing space scientists ready access to a wide assortment of space geometry useful in planning science observations and analyzing the instrument data returned therefrom. The system includes software used to compute many derived parameters such as altitude, LAT/LON and lighting angles, and software able to find when user-specified geometric conditions are obtained. While not a formal standard, it has achieved widespread use in the worldwide planetary science community

planetary science investigations

Systems Development, Data Mining, and Knowledge Discovery

The primary role of the Technical Integration Office is to provide technical solutions and services to different branches at KSC (Kennedy Space Center) and NASA program customers. The Technical Integration Office helps support KSC's operational needs by providing services such as digital connectivity, data center services, modelling and simulation tools, and communication video services. To learn the necessary technology and processes for my internship, I am working on two projects: learning C# (C Sharp programming language) with SQL and developing requirements for a PX (Communication and Public Engagement) inventory management system. To learn how to efficiently program with C#, my mentor assigned me to complete a sports informatics application that would let users discover facts and rules about various sports. The sports informatics application comes with search capabilities, report generating features, rule lists that users can modify, and diagrams for various sport strategies. To further build upon this project, I also developed a sport simulation game with the application. Once I begin more SQL-based projects, I will have the opportunity to learn how to manage databases and link SQL servers with C# programs. To develop requirements for the inventory management system, I have met with PX representatives and toured their storage facilities to see how they organize and store their items and equipment. I will also be meeting with representatives from the budget office to find out what information must be in a system budget report. The main components the system must have are customer request management, a search feature for items and equipment, report generation capabilities, and automated system warnings when item quantities reach or go below administrator-specified threshold levels. I have drafted questions and shall statements that will ultimately become part of the inventory management system requirements document.

Espinosa, Gabriel

Data Mining of Network Logs

The statement of purpose is to analyze network monitoring logs to support the computer incident response team. Specifically, gain a clear understanding of the Uniform Resource Locator (URL) and its structure, and provide a way to breakdown a URL based on protocol, host name domain name, path, and other attributes. Finally, provide a method to perform data reduction by identifying the different types of advertisements shown on a webpage for incident data analysis. The procedures used for analysis and data reduction will be a computer program which would analyze the URL and identify and advertisement links from the actual content links.

Collazo, Carlimar

Data Mining for Vortices on the Earth's Magnetosphere - Algorithm Application for Detection and Analysis

Unsteady processes in the solar wind– magnetosphere interaction, such as vortices developed at the magnetopause boundary by the Kelvin–Helmholtz instability, may contribute to the process of mass, momentum and energy transfer into the Earth’s magnetosphere. The research described in this paper validates an algorithm to automatically detect and characterize vortices based on velocity data from simulations. The vortex identification algorithm (VIA) systematically searches the 3-D velocity fields to identify critical points where the magnitude of the velocity vector vanishes. The velocity gradient tensor is computed and its invariants are used to assess vortex structure in the flow field. We use the Community Coordinated Modeling Center (CCMC) Runs on Request capability to create a series of model runs initialized from the conditions observed by the Cluster mission in the Hwang et al. (2011) analysis of Kelvin–Helmholtz vortices observed during southward interplanetary magnetic field (IMF) conditions. We analyze further the properties of the vortices found in the runs, including the velocity changes within their motion across the magnetosheath. We also demonstrate the potential of our tool to identify and characterize other transient features (e.g., flux transfer events, FTEs) with vortical internal structures. We find that the vortices are associated with flows on the magnetosheath side of the magnetopause that reach speeds greater than the solar wind speed at the bow shock.

Collado-Vega, Yaireska M.

A Data Mining Project to Identify Cardiovascular Related Factors That May Contribute to Changes in Visual Acuity Within the US Astronaut Corps

Many of the cardiovascular-related adaptations that occur in the microgravity environment are due, in part, to a well-characterized cephalad-fluid shift that is evidenced by facial edema and decreased lower limb circumference. It is believed that most of these alterations occur as a compensatory response necessary to maintain a "normal" blood pressure and cardiac output while in space. However, data from both flight and analog research suggest that in some instances these microgravity-induced alterations may contribute to cardiovascular-related pathologies. Most concerning is the potential relation between the vision disturbances experienced by some long duration crewmembers and changes in cerebral blood flow and intra-ocular pressure. The purpose of this project was to identify cardiovascular measures that may potentially distinguish individuals at risk for visual disturbances after long duration space flight. Toward this goal, we constructed a dataset from Medical Operation tilt/stand test evaluations pre- (days L-15-L-5) and immediate post-flight (day R+0) on 20 (3 females, 17 males). We restricted our evaluation to only crewmembers who participated in both shuttle and space station missions. Data analysis was performed using both descriptive and analytical methods (Stata 11.2, College Station, TX) and are presented as means +/- 95% CI. Crewmembers averaged 5207 (3447 - 8934) flight hours across both long (MIR-23 through Expedition16) and short (STS-27 through STS-101) duration missions between 1988 and 2008. The mean age of the crew at the time of their most recent shuttle flight was 41 (34-44) compared to 47 (40-54) years during their time on station. In order to focus our analysis (we did not have codes to separate out subjects by symptomotology) , we performed a visual inspection of each cardiovascular measures captured during testing and plotted them against stand time, pre- to post-flight, and between mission duration. It was found that pulse pressure most clearly differentiated the two mission types. Statistical analysis confirmed that pulse pressure was significantly higher before [45.6; (42.1 to 49.1)] and after [50.7; (46.9 to 54.6)] time on station compared with their most recent shuttle flight [31.6 (27.8 to 35.4), and 32.2 (28.3 to 36.0) respectively] even after correcting differences in age and cumulative number of mission hours. Without knowing the identity of which long duration crewmembers demonstrated visual changes, we were limited to examining whether certain crew regulate components of pulse pressure, systolic and diastolic blood pressure, differently due to microgravity exposure. To that end, we stratified crew into tertiles based on either their pre-flight measure of systolic or diastolic blood pressure. Those crew in the highest tertile for both systolic (lower tertile (n=8; 103-111), middle tertile (n=7; 113-121), and upper tertile (n=5; 125-136) and diastolic blood pressure (lower tertile (n=8; 58-64), middle tertile (n=7; 67-73), and upper tertile (n=5; 75-81) demonstrated less variability in pulse pressure between R+0 and L-10 (Figure 2). Interestingly, those crewmembers with the highest resting systolic blood pressure demonstrated either no change or in some instances an increase in total peripheral resistance, where those in the lower tertiles had lower values of total peripheral resistance compared to pre-flight levels. In this study, it was found that crewmembers in the highest tertile for both systolic and diastolic blood pressure demonstrated less variability in pulse pressure and that the decrease in variability was due in part to lower levels of compliance as indicated by similar or higher levels of total peripheral resistance after compared with before flight levels. Whether there is a relation between blood pressure regulation and total peripheral resistance in crew presenting with negative changes in visual acuity remains unknown.

Westby, Christian M.

GeneLab for High Schools: Data Mining for the Next Generation

Modern biological sciences have become increasingly based on molecular biology and high-throughput molecular techniques, such as genomics, transcriptomics, and proteomics. NASA Scientists and the NASA Space Biology Program have aimed to examine the fundamental building blocks of life (RNA, DNA and protein) in order to understand the response of living organisms to space and aid in fundamental research discoveries on Earth. In an effort to enable NASA funded science to be available to everyone, NASA has collected the data from omics studies and curated them in a data system called GeneLab. Whilst most college-level interns, academics and other scientists have had some interaction with omics data sets and analysis tools, high school students often have not. Therefore, the Space Biology Program is implementing a new Summer Program for high-school students that aims to inspire the next generation of scientists to learn about and get involved in space research using GeneLabs Data System. The program consists of three main components core learning modules, focused on developing students knowledge on the Space Biology Program and Space Biology research, Genelab and the data system, and previous research conducted on model organisms in space; networking and team work, enabling students to interact with guest lecturers from local universities and their fellow peers, and also enabling them to visit local universities and genomics centers around the Bay area; and finally an independent learning project, whereby students will be required to form small groups, analyze a dataset on the Genelab platform, generate a hypothesis and develop a research plan to test their hypothesis. This program will not only help inspire high-school students to become involved in space-based research but will also help them develop key critical thinking and bioinformatics skills required for most college degrees and furthermore, will enable them to establish networks with their peers and connections with university Professors that may help them achieve their educational goals.

genelab

Development and Testing of Data Mining Algorithms for Earth Observation

The new algorithms developed under this project included a principled procedure for classification of objects, events or circumstances according to a target variable when a very large number of potential predictor variables is available but the number of cases that can be used for training a classifier is relatively small. These "high dimensional" problems require finding a minimal set of variables -called the Markov Blanket-- sufficient for predicting the value of the target variable. An algorithm, the Markov Blanket Fan Search, was developed, implemented and tested on both simulated and real data in conjunction with a graphical model classifier, which was also implemented. Another algorithm developed and implemented in TETRAD IV for time series elaborated on work by C. Granger and N. Swanson, which in turn exploited some of our earlier work. The algorithms in question learn a linear time series model from data. Given such a time series, the simultaneous residual covariances, after factoring out time dependencies, may provide information about causal processes that occur more rapidly than the time series representation allow, so called simultaneous or contemporaneous causal processes. Working with A. Monetta, a graduate student from Italy, we produced the correct statistics for estimating the contemporaneous causal structure from time series data using the TETRAD IV suite of algorithms. Two economists, David Bessler and Kevin Hoover, have independently published applications using TETRAD style algorithms to the same purpose. These implementations and algorithmic developments were separately used in two kinds of studies of climate data: Short time series of geographically proximate climate variables predicting agricultural effects in California, and longer duration climate measurements of temperature teleconnections.

Glymour, Clark

Virtual Sensors: Using Data Mining Techniques to Efficiently Estimate Remote Sensing Spectra

Various instruments are used to create images of the Earth and other objects in the universe in a diverse set of wavelength bands with the aim of understanding natural phenomena. These instruments are sometimes built in a phased approach, with some measurement capabilities being added in later phases. In other cases, there may not be a planned increase in measurement capability, but technology may mature to the point that it offers new measurement capabilities that were not available before. In still other cases, detailed spectral measurements may be too costly to perform on a large sample. Thus, lower resolution instruments with lower associated cost may be used to take the majority of measurements. Higher resolution instruments, with a higher associated cost may be used to take only a small fraction of the measurements in a given area. Many applied science questions that are relevant to the remote sensing community need to be addressed by analyzing enormous amounts of data that were generated from instruments with disparate measurement capability. This paper addresses this problem by demonstrating methods to produce high accuracy estimates of spectra with an associated measure of uncertainty from data that is perhaps nonlinearly correlated with the spectra. In particular, we demonstrate multi-layer perceptrons (MLPs), Support Vector Machines (SVMs) with Radial Basis Function (RBF) kernels, and SVMs with Mixture Density Mercer Kernels (MDMK). We call this type of an estimator a Virtual Sensor because it predicts, with a measure of uncertainty, unmeasured spectral phenomena.

Srivastava, Ashok N.

Data Mining Activity for Bone Discipline: Calculating a Factor of Risk for Hip Fracture in Long-Duration Astronauts

The factor-of-risk (Phi), defined as the ratio of applied load to bone strength, is a biomechanical approach to hip fracture risk assessment that may be used to identify subjects who are at increased risk for fracture. The purpose of this project was to calculate the factor of risk in long duration astronauts after return from a mission on the International Space Station (ISS), which is typically 6 months in duration. The load applied to the hip was calculated for a sideways fall from standing height based on the individual height and weight of the astronauts. The soft tissue thickness overlying the greater trochanter was measured from the DXA whole body scans and used to estimate attenuation of the impact force provided by soft tissues overlying the hip. Femoral strength was estimated from femoral areal bone mineral density (aBMD) measurements by dual-energy x-ray absorptiometry (DXA), which were performed between 5-32 days of landing. All long-duration NASA astronauts from Expedition 1 to 18 were included in this study, where repeat flyers were treated as separate subjects. Male astronauts (n=20) had a significantly higher factor of risk for hip fracture Phi than females (n=5), with preflight values of 0.83+/-0.11 and 0.36+/-0.07, respectively, but there was no significant difference between preflight and postflight Phi (Figure 1). Femoral aBMD measurements were not found to be significantly different between men and women. Three men and no women exceeded the theoretical fracture threshold of Phi=1 immediately postflight, indicating that they would likely suffer a hip fracture if they were to experience a sideways fall with impact to the greater trochanter. These data suggest that male astronauts may be at greater risk for hip fracture than women following spaceflight, primarily due to relatively less soft tissue thickness and subsequently greater impact force.

Ellman, R.