Search NASA⌕ Search

SEARCH · Search NASA

Results for “data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Voltage Mining for (De)lithiation-Stabilized Cathodes and a Machine Learning Model for Li-Ion Cathode Voltage

Advances in lithium-metal anodes have inspired interest in discovery of Li-free cathodes, most of which are natively found in their charged state. This is in contrast to today's commercial lithium-ion battery cathodes, which are more stable in their discharged state. In this study, we combine calculated cathode voltage information from both categories of cathode materials, covering 5577 and 2423 total unique structure pairs, respectively. The resulting voltage distributions with respect to the redox pairs and anion types for both classes of compounds emphasize design principles for high-voltage cathodes, which favor later Period 4 transition metals in their higher oxidation states and more electronegative anions like fluorine or polyanion groups. Generally, cathodes that are found in their charged, delithiated state are shown to exhibit voltages lower than those that are most stable in their lithiated state, in agreement with thermodynamic expectations. Deviations from this trend are found to originate from different anion distributions between redox pairs. In addition, a machine learning model for voltage prediction based on chemical formulas is trained and shows state-of-the-art performance when compared to two established composition-based ML models for material properties predictions, Roost and CrabNet.

25 ENERGY STORAGE↗

An assessment of AVIRIS data for hydrothermal alteration mapping in the Goldfield Mining District, Nevada

Airborne Visible and Infrared Imaging Spectrometer (AVIRIS) data were acquired over the Goldfield Mining District, Nevada, in September 1987. Goldfield is one of the group of large epithermal precious metal deposits in Tertiary volcanic rocks, associated with silicic volcanism and caldera formation. Hydrothermal alteration consists of silicification along fractures, advanced agrillic and argillic zones further away from veins and more widespread propylitic zones. An evaluation of AVIRIS data quality was performed. Faults in the data, related to engineering problems and a different behavior of the instrument while on-board the U2, were encountered. Consequently, a decision was made to use raw data and correct them only for dark current variations and detector read-out-delays. New software was written to that effect. Atmospheric correction was performed using the flat field correction technique. Analysis of the data was then performed to extract spectral information, mainly concentrating on the 2 to 2.45 micron window, as the alteration minerals of interest have their distinctive spectral reflectance features in this region. Principally kaolinite and alunite spectra were clearly obtained. Mapping of the different minerals and alteration zones was attempted using ratios and clustering techniques. Poor signal-to-noise performance of the instrument and the lack of appropriate software prevented the production of an alteration map of the area. Spectra extracted locally from the AVIRIS data were checked in the field by collecting representative samples of the outcrops.

Carrere, Veronique↗

Eddy covariance towers as sentinels of abnormal radioactive material releases

Ensuring accurate detection and attribution of abnormal releases of radioactive material is critical for protecting human health and safety. Most commonly, such detection is accomplished via active monitoring approaches involving the collection of physical samples. Further, this is labor intensive and limits the temporal and spatial resolution of any detected events to a relatively coarse level. As an alternative first step towards passive monitoring, we developed an approach using eddy flux tower data records to identify signals from a known abnormal release and quantify the extent to which that signal also occurs at other times in the data record. Through two case studies, one of which targeted the Fukushima nuclear disaster and the other targeting an abnormal release event at a radioisotope production facility in Fleurus, Belgium, we tested our approach and identified several potential heretofore unidentified abnormal events that were consistent with atmospheric circulation patterns and/or wind direction from known release sites. Because our approach is relatively simple and is resistant to systematic errors in the observational record, it has broad applicability beyond specific constituents and ecosystem types to identify a wide variety of limited-duration anomalies in flux tower data to ensure human health and industrial safety.

54 ENVIRONMENTAL SCIENCES↗

Microbial spies and bloggers: programming cells to convert environmental information into discernible signals

Microbes regulate their dynamic behaviors using the chemical and physical characteristics of their environment. The ability of microbes to continuously convert this physicochemical information into biochemical information and to use organic matter in the environment as a power source makes these organisms attractive as chassis for building sensors. However, most biosensors have severe limitations when considering applications in hard-to-image settings like soils, sediments, and wastewater. Emerging technologies at the interface of biomolecular design, microbiome engineering, and synthetic biology offer new tools to program cells and communities as biosensors for these settings. Here, in this review, we describe innovations in biosensor outputs that are enabling new applications in complex environments, including reporters that are read out using electrochemical, gas chromatography, hyperspectral imaging, and next-generation sequencing methods. We also discuss computational advances that are accelerating the diversification of sensing components by mining metagenomics data for new transcriptional regulators and by designing allosteric protein switches that directly regulate reporter outputs using analytes. We highlight emerging opportunities for programming undomesticated microbes in communities to function as distributed sensors in the environment. Finally, we discuss the need for responsible biosensor development and to modernize regulatory frameworks to support evidence-based assessment of environmental biosensors.

analyte↗

Unlocking Solutions: Innovative Approaches to Identifying and Mitigating the Environmental Impacts of Undocumented Orphan Wells in the United States

In the United States, hundreds of thousands of undocumented orphan wells have been abandoned, leaving the burden of managing environmental hazards to governmental agencies or the public. These wells, a result of over a century of fossil fuel extraction without adequate regulation, lack basic information like location and depth, emit greenhouse gases, and leak toxic substances into groundwater. For most of these wells, basic information such as well location and depth is unknown or unverified. Addressing this issue necessitates innovative and interdisciplinary approaches for locating, characterizing, and mitigating their environmental impacts. Our survey of the United States revealed the need for tools to identify well locations and assess conditions, prompting the development of technologies including machine learning to automatically extract information from old records (95%+ accuracy), remote sensing technologies like aero-magnetometers to find buried wells, and cost-effective methods for estimating methane emissions. Notably, fixed-wing drones equipped with magnetometers have emerged as cost-effective and efficient for discovering unknown wells, offering advantages over helicopters and quadcopters. Efforts also involved leveraging local knowledge through outreach to state and tribal governments as well as citizen science initiatives. These initiatives aim to significantly contribute to environmental sustainability by reducing greenhouse gases and improving air and water quality.

54 ENVIRONMENTAL SCIENCES↗

A Survey of Open Source Software Repositories in the U.S. Department of Energy’s National Laboratories

There are 17 national laboratory systems in the United States operating under the auspices of the U.S. Department of Energy (DOE). These government labs employ tens of thousands of people engaging in research software engineering activities across a variety of missions. To support this work, many open source projects are maintained. Further, many of these projects have broad utility to the computing community at large and domain scientists in a variety of fields. However, the complexity and decentralized nature of the laboratory system has resulted in a situation where no one entity even knows about all the open source software projects in this ecosystem, let alone crude metrics of their health. In this article, we do the first external inventory of open source software repositories with a nexus to DOE labs. We posit that a project’s need for sustainability support can be determined by comparing measures of active use to measures of active maintenance.

97 MATHEMATICS AND COMPUTING↗

Visual Analytics of Multivariate Networks With Representation Learning and Composite Variable Construction

Multivariate networks are commonly found in real-world data-driven applications. Uncovering and understanding the relations of interest in multivariate networks is not a trivial task. This article presents a visual analytics workflow for studying multivariate networks to extract associations between different structural and semantic characteristics of the networks (e.g., what are the combinations of attributes largely relating to the density of a social network?). The workflow consists of a neural-network-based learning phase to classify the data based on the chosen input and output attributes, a dimensionality reduction and optimization phase to produce a simplified set of results for examination, and finally an interpreting phase conducted by the user through an interactive visualization interface. A key part of our design is a composite variable construction step that remodels nonlinear features obtained by neural networks into linear features that are intuitive to interpret. We demonstrate the capabilities of this workflow with multiple case studies on networks derived from social media usage and also evaluate the workflow with qualitative feedback from experts.

97 MATHEMATICS AND COMPUTING↗

Software Development Cost Estimation Executive Summary

Identify simple fully validated cost models that provide estimation uncertainty with cost estimate. Based on COCOMO variable set. Use machine learning techniques to determine: a) Minimum number of cost drivers required for NASA domain based cost models; b) Minimum number of data records required and c) Estimation Uncertainty. Build a repository of software cost estimation information. Coordinating tool development and data collection with: a) Tasks funded by PA&E Cost Analysis; b) IV&V Effort Estimation Task and c) NASA SEPG activities.

data mining↗

Software Development Cost: How Much? You Sure? Technical Summary

This slide presentation reports on findings from analyzing a NASA COCOMO 81 dataset with 93 records. The current tool is called COSEEKMO, using a methodology can be applied to any set of cost models and data. The effort was part of a research initiative funded by the NASA Office of Safety and Mission and Assurance (OSMA) aimed at improving software reliability.

software effort estimation↗

Studies in Software Cost Model Behavior: Do We Really Understand Cost Model Performance?

While there exists extensive literature on software cost estimation techniques, industry practice continues to rely upon standard regression-based algorithms. These software effort models are typically calibrated or tuned to local conditions using local data. This paper cautions that current approaches to model calibration often produce sub-optimal models because of the large variance problem inherent in cost data and by including far more effort multipliers than the data supports. Building optimal models requires that a wider range of models be considered while correctly calibrating these models requires rejection rules that prune variables and records and use multiple criteria for evaluating model performance. The main contribution of this paper is to document a standard method that integrates formal model identification, estimation, and validation. It also documents what we call the large variance problem that is a leading cause of cost model brittleness or instability.

data mining↗