Search NASA⌕ Search

DOE OSTI · 1649534

Advanced data science toolkit for non-data scientists – A user guide

Abstract

Emerging modern data analytics attracts much attention in materials research and shows great potential for enabling data-driven design. Data populated from the high-throughput CALPHAD approach enables researchers to better understand underlying mechanisms and to facilitate novel hypotheses generation, but the increasing volume of data makes the analysis extremely challenging. Here in this paper, we introduce an easy-to-use, versatile, and open-source data analytics frontend, ASCENDS (Advanced data SCiENce toolkit for Non-Data Scientists), designed with the intent of accelerating data-driven materials research and development. The toolkit is also of value beyond materials science as it can analyze the correlation between input features and target values, train machine learning models, and make predictions from the trained surrogate models of any scientific dataset. Various algorithms implemented in ASCENDS allow users performing quantified correlation analyses and supervised machine learning to explore any datasets of interest without extensive computing and data science background. The detailed usage of ASCENDS is introduced with an example of experimental high-temperature alloy data.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Peng, Jian, Lee, Sangkeun (Matt), Williams, Andrew, Haynes, James A., Shin, Dongwon. 2020-02-01. Advanced data science toolkit for non-data scientists – A user guide. https://doi.org/10.1016/j.calphad.2019.101733

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Ontologies for Intelligent Data Science

As anyone even vaguely aware of current technology can tell you, machine learning (ML) and artificial intelligence (AI) have made exceptional breakthroughs in recent years. Generative artificial intelligence (GAI) emerged circa 2022 dominated by Large Language Models (LLMs) and generative tools for images emerged at about the same time.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Verification and Validation of Performance with Dissemination of Best Practices in District Energy and CHP for Enhanced Resiliency, Energy Efficiency, and Cybersecurity

This report contains the results of the International District Energy Association’s work to analyze, validate, and verify performance data of existing district energy systems and identify industy best practices for the purpose of improving system reliability, resiliency, and efficiency, and to accellerate decarbonization. In addition to a technical evaluation of the surveyed systems and identification of a series of technical performance metrics, the report illustrates the accompanying operations and financial best practices employed by surveyed systems to fully serve their customer base. Additionally, the third chapter of the report describes the current landscape of cybersecurity threats and counteracting measures, and recommends a series of steps for effectively guarding highly networked district energy systems against cybersecurity attacks.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗