Search NASA⌕ Search

SEARCH · Search NASA

Results for “data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Knowledge Driven Image Mining with Mixture Density Mercer Kernels

This paper presents a new methodology for automatic knowledge driven image mining based on the theory of Mercer Kernels; which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly infinite dimensional feature space. In that high dimensional feature space, linear clustering, prediction, and classification algorithms can be applied and the results can be mapped back down to the original image space. Thus, highly nonlinear structure in the image can be recovered through the use of well-known linear mathematics in the feature space. This process has a number of advantages over traditional methods in that it allows for nonlinear interactions to be modelled with only a marginal increase in computational costs. In this paper, we present the theory of Mercer Kernels, describe its use in image mining, discuss a new method to generate Mercer Kernels directly from data, and compare the results with existing algorithms on data from the MODIS (Moderate Resolution Spectral Radiometer) instrument taken over the Arctic region. We also discuss the potential application of these methods on the Intelligent Archive, a NASA initiative for developing a tagged image data warehouse for the Earth Sciences.

Srivastava, Ashok N.↗

Knowledge Driven Image Mining with Mixture Density Mercer Kernals

This paper presents a new methodology for automatic knowledge driven image mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly infinite dimensional feature space. In that high dimensional feature space, linear clustering, prediction, and classification algorithms can be applied and the results can be mapped back down to the original image space. Thus, highly nonlinear structure in the image can be recovered through the use of well-known linear mathematics in the feature space. This process has a number of advantages over traditional methods in that it allows for nonlinear interactions to be modelled with only a marginal increase in computational costs. In this paper we present the theory of Mercer Kernels; describe its use in image mining, discuss a new method to generate Mercer Kernels directly from data, and compare the results with existing algorithms on data from the MODIS (Moderate Resolution Spectral Radiometer) instrument taken over the Arctic region. We also discuss the potential application of these methods on the Intelligent Archive, a NASA initiative for developing a tagged image data warehouse for the Earth Sciences.

Srivastava, Ashok N.↗

Development of remote sensing techniques for assessment of hydrologic conditions in coal mining regions of Appalachia

In December of 1974 the John F. Kennedy Space Center, NASA, and the Water Resources Division, United States Geological Survey (USGS), acquired photographic, thermal, and multispectral data over the Cumberland region of eastern Tennessee. This data was effectively used to delineate ground water sources, and surface water runoff into river systems in the Cumberlands. The data, coupled with an overview of the area from the Earth Resources Technology Satellite (ERTS), could be useful in determining hydrologic conditions in coal mining regions of the Appalachians.

Pope, C. D.↗

Machine processing of remotely sensed data; Proceedings of the Fifth Annual Symposium, Purdue University, West Lafayette, Ind., June 27-29, 1979

Papers are presented on techniques and applications for the machine processing of remotely sensed data. Specific topics include the Landsat-D mission and thematic mapper, data preprocessing to account for atmospheric and solar illumination effects, sampling in crop area estimation, the LACIE program, the assessment of revegetation on surface mine land using color infrared aerial photography, the identification of surface-disturbed features through a nonparametric analysis of Landsat MSS data, the extraction of soil data in vegetated areas, and the transfer of remote sensing computer technology to developing nations. Attention is also given to the classification of multispectral remote sensing data using context, the use of guided clustering techniques for Landsat data analysis in forest land cover mapping, crop classification using an interactive color display, and future trends in image processing software and hardware.

Tendam, I. M.↗

LANDSAT inventory of surface-mined areas using extendible digital techniques

Multispectral LANDSAT imagery was analyzed to provide a rapid and accurate means of identification, classification, and measurement of strip-mined surfaces in Western Maryland. Four band analysis allows distinction of a variety of strip-mine associated classes, but has limited extendibility. A method for surface area measurements of strip mines, which is both geographically and temporally extendible, has been developed using band-ratioed LANDSAT reflectance data. The accuracy of area measurement by this method, averaged over three LANDSAT scenes taken between September 1972 and July 1974, is greater than 93%. Total affected acreage of large (50 hectare/124 acre) mines can be measured to within 1.0%.

Anderson, A. T.↗

Monitoring of environmental effects of coal strip mining from satellite imagery

This paper evaluates satellite imagery as a means of monitoring coal strip mines and their environmental effects. The satellite imagery employed is Skylab EREP S-190A and S-190B from SL-2, SL-3 and SL-4 missions; a large variety of camera/film/filter combinations has been reviewed. The investigation includes determining the applicability of satellite imagery for detection of disturbed acreage in areas of coal surface mining as well as the much more detailed monitoring of specific surface-mining operations, including: active mines, inactive mines, highwalls, ramp roads, pits, water impoundments and their associated acidity, graded areas and types of grading, and reclamed areas. Techniques have been developed to enable mining personnel to utilize this imagery in a practical and economic manner, requiring no previous photo-interpretation background and no purchases of expensive viewing or data-analysis equipment. To corroborate the photo-interpretation results, on-site observations were made in the very active mining area near Madisonville, Kentucky.

Brooks, R. L.↗

Provenance Representation in the Global Change Information System (GCIS)

Global climate change is a topic that has become very controversial despite strong support within the scientific community. It is common for agencies releasing information about climate change to be served with Freedom of Information Act (FOIA) requests for everything that led to that conclusion. Capturing and presenting the provenance, linking to the research papers, data sets, models, analyses, observation instruments and satellites, etc. supporting key findings has the potential to mitigate skepticism in this domain. The U.S. Global Change Research Program (USGCRP) is now coordinating the production of a National Climate Assessment (NCA) that presents our best understanding of global change. We are now developing a Global Change Information System (GCIS) that will present the content of that report and its provenance, including the scientific support for the findings of the assessment. We are using an approach that will present this information both through a human accessible web site as well as a machine readable interface for automated mining of the provenance graph. We plan to use the developing W3C PROV Data Model and Ontology for this system.

Tilmes, Curt↗

NEWTS Integrated Dataset (version 1.0)

The National Energy Water Treatment and Speciation (NEWTS) Integrated Dataset v1.0 provides water researchers, community leaders, and regulators with a unified and standardized energy-related wastewater stream database. This resource is derived from 27 state and federal entities, and scientific publications, and contains more than 400,000 sample records, many of which also provide geospatial information. The dataset includes data for several different energy-related wastewater types including produced water, other oil and gas wastewaters, mine drainage, coal ash leachate, and power plant wastewater. The NEWTS Integrated Dataset was built to support environmentally prudent decision-making, explore treatment opportunities, and identify potential critical mineral sources. A subset of this novel resource is also featured on NETL NEWTS State-Level Database Dashboard. Additional data can be found in the NEWTS EDX Group and the NEWTS Federal Database Dashboard.

abandoned mine drainage↗

CO2 Storage Economic Analysis: CarbonSAFE Use Case

Poster on “CO2 Storage Economic Analysis: CarbonSAFE Use Case” for the CCUS 2025 conference held in Houston, Texas March 3-5, 2025. The cost of designing, permitting, constructing, operating, and closing a CO2 storage project is of vital importance to project developers. The National Energy Technology Laboratory has developed the NRAP/SMART Technoeconomic and Liability Evaluation for Storage (TALES) Model to provide quantitative cost-based insights to support developers planning CO2 injection and storage projects. This study presents a collaborative economic analysis applying TALES with data from the San Juan Basin CarbonSAFE Phase III project led by the New Mexico Institute of Mining and Technology to estimate potential costs incurred during the implementation of a real-world commercial-scale carbon storage project. Scenario analysis was implemented in which different operational and cost attributes were varied and the associated cost implications observed. Key results data and project cost summary metrics, first-year breakeven price of CO2 ($/tonne) and net present value (NPV), are presented for base and alternative cases. Output provides a unique perspective for project stakeholders towards evaluating the influence of different operational strategies and financing approaches on overall project cost and financial viability.

carbon storage↗

BiG-SCAPE 2.0 and BiG-SLiCE 2.0: scalable, accurate and interactive sequence clustering of metabolic gene clusters

Microbial metabolic gene clusters encode the biosynthesis or catabolism of metabolites that facilitate ecological specialization, mediate microbiome interactions and constitute a major source of medicines and crop protection agents. Here, we present BiG-SCAPE and BiG-SLiCE 2.0, next-generation methods that facilitate scalable, accurate and interactive gene cluster analyses. BiG-SCAPE 2.0 updates its classification, alignment methods, and visualizations, enabling more accurate analysis, up to 8x faster runtimes and halved memory requirements. BiG-SLiCE 2.0 updates its distance metric, pHMM database, and classification logic, resulting in increased sensitivity nearing that of BiG-SCAPE. Analysis of 260,630 biosynthetic gene clusters from publicly available genomes reveals that both tools generate concurring estimates of gene cluster diversity, thus providing significantly extended methodological support for recent evidence indicating that the vast majority of natural product diversity remains unexplored. Together, these updates will facilitate global genome mining efforts for natural product discovery and microbiome analyses scalable with current data sizes.

Draisma, Arjan [Wageningen University & Research (↗

MINE: a new way to design genetics experiments for discovery

Abstract The Maximally Informative Next Experiment or MINE is a new experimental design approach for experiments, such as those in omics, in which the number of effects or parameters p greatly exceeds the number of samples n (p > n). Classical experimental design presumes n > p for inference about parameters and its application to p > n can lead to over-fitting. To overcome p > n, MINE is an ensemble method, which makes predictions about future experiments from an existing ensemble of models consistent with available data in order to select the most informative next experiment. Its advantages are in exploration of the data for new relationships with n < p and being able to integrate smaller and more tractable experiments to replace adaptively one large classic experiment as discoveries are made. Thus, using MINE is model-guided and adaptive over time in a large omics study. Here, MINE is illustrated in two distinct multiyear experiments, one involving genetic networks in Neurospora crassa and a second one involving a genome-wide association study in Sorghum bicolor as a comparison to classic experimental design in an agricultural setting.

Biochemistry & Molecular Biology↗

The application of remote sensing technology to the solution of problems in the management of resources in Indiana

The use of satellite remote sensing for resources management was investigated in Indiana. The technique was applied to strip mining and reclamation, highway planning, and the detection of dolomite reefs. A data base was created and used to produce land characteristics and suitability maps for land use planning. In addition, a three dimensional model was developed which provides a cross-sectional profile of the thermal plumes emitted by point sources of thermal pollution into rivers and lakes; this model may be used for the design and site selection of electric power plants.

Landgrebe, D. A.↗

Multidisciplinary applications of ERTS and Skylab data in Ohio

Experimental studies of ERTS-1 and Skylab earth resources data, in combination with correlative aircraft and on-site data, for environmental quality, land use, and resource management applications in Ohio show several areas of operational promise. Prime data use candidates demonstrated to date include definition and enforcement of surface mining (all minerals) legislation; Lake Erie modeling/management; land use classification and mapping studies at state, regional, and localized levels; and resources' inventories particularly of forested areas on both regional (multicounty) and localized scales.

Sweet, D. C.↗

Assessment of a 40-kilowatt stirling engine for underground mining applications

An assessment of alternative power souces for underground mining applications was performed. A 40-kW Stirling research engine was tested to evaluate its performance and emission characteristics when operated with helium working gas and diesel fuel. The engine, the test facility, and the test procedures are described. Performance and emission data for the engine operating with helium working gas and diesel fuel are reported and compared with data obtained with hydrogen working gas and unleaded gasoline fuel. Helium diesel test results are compared with the characteristics of current diesel engines and other Stirling engines. External surface temperature data are also presented. Emission and temperature results are compared with the Federal requirements for diesel underground mine engines. The durability potential of Stirling engines is discussed on the basis of the experience gaind during the engine tests.

Cairelli, J. E.↗

MERRA-2 Data and Analytic Services at NASA GES DISC for Climate Extremes Study

NASA's climate reanalysis datasets from the Modern Era Retrospective-analysis for Research and Applications, Version 2 (MERRA-2) contains numerous long-term atmosphere, land, and ocean data products from 1980-present. MERRA-2 datasets, such as precipitation, soil moisture, and temperature, have been used widely to study extreme events. The native archived MERRA-2 data files are day-file (hourly time interval) and month-file, containing up to 125 parameters in one file. Due to the large number of data files and volumes, it is challenging for users, especially the applications research community, to handle the original hourly data files for long time periods to analyze extreme events. In this presentation, we review MERRA-2 data for studies of extreme conditions, and demonstrate analytic services at the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). One of the current operational services, 'subsetter', allows users to download only specific data of interest, i.e. data selected by parameter, region, and time period. New services are under development that will provide more 'on-the-fly' statistical calculations when downloading data; improve efficiency when accessing long time-series data. We will provide additional "How-to" resources that include step-by-step instructions on data access and usage. We have tested restructuring of day-files in an optimized data cube, which has significantly improved system performance for accessing long time-series. Overall performance is associated with cube size and structure, data compression method, and how the data are accessed. The optimized data cube structure will enable better online analytic services for statistical analysis and extreme events mining. To demonstrate the service, we use an extreme drought associated with the anomalous 2016 monsoon over southern Asia. This prototype time-series service may be augmented in the cloud infrastructure in the future.

data access↗

Enhanced Resistance Pines for Improved Renewable Biofuel and Chemical Production (Technical Report)

We completed phenotyping constitutive and inducible oleoresin flow across two seasons, constitutive resin canal number and density and wood terpene content in our ADEPT2 and CCLONES populations. We completed genetic association between 19 oleoresin phenotypes and a total of 523,192 SNP markers from ADEPT2 and 13,883 SNP markers in CCLONES using four mixed linear models. A total of 293 significant SNPs (FDR = 0.20) were identified. We used the MENTOR tool to mine mechanistic connections from a multiplex network constructed from poplar multi-omic data to construct a conceptual model for a subset of these significant SNPs. Our model contains 6 transcriptional regulators in addition to 3 monoterpene synthases. To generate more lines of evidence for these significant SNPs, we completed a time course RNAseq experiment after inducing vascular zone cells to differentiate into new resin canals with a methyl jasmonate treatment, a single nuclei RNAseq that identified differentiating resin canal epithelial cells and are completing analysis for a QTL study in a hybrid pine population. The time course identified 4634 significantly down and 1890 significantly up regulated transcripts after treatment with methyl jasmonate, an inducer of new resin canal formation in the vascular cambial meristem. To analyze this large set of differentially regulated genes, we created a predictive expression network and analyzed it with random walk restart using 6 seed genes coding for transcription factors regulating xylem differentiation in poplar. Of the top ranked 200 transcripts, 119 transcripts were significant differentially expressed supporting these transcripts as potential candidates regulating resin canal formation. Analysis of single nuclei sequencing of shoot tips that contain differentiating resin canals, identified 10 clusters. One cluster was highly enriched in transcripts coding for 9 of the enzymes in the MEP pathway 3 prenyl synthetases, and 3 monoterpene synthases strongly suggesting that this cluster represents resin canal epithelial cells. We are mining the additional transcripts to create a trajectory analysis. In summary, we have identified > 10 novel genes that are strongly supported candidates for further analysis in breeding lines and for genetic engineering over- and under- expressing lines to increase wood terpene content to improve resistance to insect and fungal pathogens while simultaneously increasing terpene supplies for renewable chemicals and biofuels.

59 BASIC BIOLOGICAL SCIENCES↗

NEWTS Integrated Dataset (version 2.0)

The National Energy Water Treatment and Speciation (NEWTS) Integrated Dataset v2.0 provides water researchers, community leaders, regulators, and industry stakeholders with a unified and standardized energy-process wastewater chemistry database. This resource is derived from 39 state and federal entities, and scientific publications, and contains more than 700,000 sample records, many of which also provide geospatial information. The dataset includes chemistry data for several different energy-process wastewater types including produced water, other oil and gas wastewaters, mine drainage, coal ash leachate, power plant wastewater, and geothermal fluids. The NEWTS Integrated Dataset was built to support prudent decision-making, characterization of potential critical mineral sources, and modeling of treatment and valorization options. A subset of this novel resource is also featured on the NEWTS State-Level Database Dashboard. Additional data can be found in the NEWTS EDX Group and the NEWTS Federal Database Dashboard.

AMD↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗