Search NASASearch

SEARCH · Search NASA

Results for “data mining system”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Large language models for batteries

Large Language Models (LLMs) are advanced artificial intelligence systems capable of solving diverse tasks using language, reasoning, and external tools. Despite their growing deployment in academia and industry, their potential remains underexplored in battery research. This review presents a comprehensive overview of existing and emerging applications of LLMs in batterie field, addressing two critical questions: What can LLMs offer to support battery-related tasks, and how to develop more effective models for this purpose. We begin by outlining the principles of LLMs and criteria for selecting appropriate models and tools for battery research and development. We then explore their roles in text-mining, data interpretation, and the development of intelligent battery systems. In parallel, we discuss technical challenges, such as data standardizing and sharing, model evaluation, and tool integration. Lastly, we propose future research directions with short-, medium-, and long-term goals and highlight more broad perspectives for connecting experts and cross-disciplinary collaborations.

SoC

Foil Bearing Supported Compressor-Expander

R&D Dynamics Corporation (RDD) is developing a fuel cell system compressor-expander designed for heavy-duty vehicle applications. The system emphasizes high reliability, versatility, long service life, and cost-effective mass production. Its design integrates a high-speed centrifugal compressor/expander, oil-free foil air/gas bearings, a direct current permanent magnet (DCPM) motor, and an inverter utilizing silicon carbide (SiC) switches. Key advantages include high efficiency, compact size and reduced weight, enhanced reliability, and elimination of oil contamination risks. The technology targets multiple sectors such as long-haul semi-trucks, mining and construction equipment, power generation, data centers, marine systems, and other heavy machinery. Evaluation units are ready for shipment, and manufacturing capabilities are currently being developed. Contract Number: DE-EE0009617 (2021).

30 DIRECT ENERGY CONVERSION

LIB Design Module for Grid Energy System Application

We will employ a machine learning approach with intelligent data mining and database construction to analyze enormous data repositories for identifying and extracting geographic-dependent cell design specifications from publicly accessible grid-scale energy storage usage databases in an automated way at scale.

Liu, Dianying [Pacific Northwest National Laborato

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING

Application of PRIM for understanding patterns in carbon dioxide model-observation differences

Reducing uncertainties in regional carbon balances requires a better understanding of CO 2 transport in synoptic weather systems. Here, we apply the Patient Rule Induction Method (PRIM), a data-mining method to identify high-density regions for a target-class within an input parameter space, to airborne observations of potential temperature, wind speed, water vapor mixing ratio, and CO 2 dry mol fraction gathered during the Atmospheric Carbon and Transport (ACT)-America Summer 2016 and Winter 2017 campaigns. ACT observations were targeted at expert-designated cases of fair weather and near-frontal warm and cold sector air at atmospheric boundary-layer, lower-, and higher free tropospheric levels (ABL, LFT, and HFT, respectively). We investigate atmospheric characteristics of these pre-defined cases and associated CO 2 model-observation-differences in the mesoscale WRF-Chem model. PRIM results separate winter- and summertime observations as well as observations from ABL, LFT, and HFT with enrichment factors of 4.0–20.5 inside the PRIM box compared to the entire dataset but cannot distinguish between near-frontal warm and cold sector observations in the higher free troposphere. Analyzing of the parameter space constrained by PRIM, we find that large magnitude model observation differences preferentially associated with times when atmospheric conditions are less typical. This association suggests that PRIM could provide a useful tool for isolating atmospheric conditions with large-magnitude and non-Gaussian CO 2 -residuals for targeted transport model evaluation and to potentially improve inversion results during synoptically active periods.

Gerken, Tobias [James Madison Univ., Harrisonburg,

Autonomous Synthesis and Inverse Design of Electrochromic Polymers with High Efficiency and Accuracy

Here, the design and synthesis of functional polymers, aimed at targeted properties through specific structures, have long been challenged by their complex and often nonlinear structure–property relationships. Key processes, including knowledge accumulation for predictive design and experimental refinement and validation, are traditionally labor-insensitive and time-consuming, making it difficult to balance accuracy and efficiency. Here, we introduce an accelerated, autonomous system for the on-demand synthesis of electronic polymers that achieves the desired electrochromic functionality with high accuracy and efficiency. Our approach leverages large language model-assisted data mining, a physics-informed copolymer machine learning model, and an AI-driven autonomous robotic workflow in the Polybot lab. Within 72 h, Polybot autonomously synthesized electrochromic polymers (ECPs) with targeted, previously-unreported color values, including green polymers with specific absorption profiles, precisely fine-tuning copolymer structures with a 5% step size in comonomer composition within a three-monomer system. A publicly accessible ECP informatics database has also been created to foster knowledge exchange.

AI-driven Robotic Lab

Rapid Adaptation of Chemical Named Entity Recognition Using Few-Shot Learning and LLM Distillation

Named entity recognition (NER) has been widely used in chemical text mining for the automatic identification and extraction of chemical entities. However, existing chemical NER systems primarily focus on scenarios with abundant training data, requiring significant human effort on annotations. This poses challenges for applications in the chemical field, such as catalysis, where many advancements have traditionally relied on trial-and-error investigations and incremental adjustment of variables. This hinders catalysis science and technology progress in addressing emerging energy and environmental crises. In this work, we propose a few-shot NER model that can quickly adapt to extract new types of chemical entities by using only a limited number of annotated examples. Our model employs a metric-learning approach to transfer entity similarity knowledge from high-resource chemical domains (with abundant annotations) to enable effective entity recognition in low-resource specialized domains (limited annotation). We validate the effectiveness of our model on a few-shot chemical NER benchmark built based on six existing chemical NER data sets. Experiments show that the proposed few-shot NER model can achieve reasonable performance with only 5 examples per entity type and shows consistent improvement as the number of examples increases. Furthermore, we demonstrate how the proposed model can be trained with large language model (LLM) annotated data, opening a new pathway for rapid adaptation of NER systems. Furthermore, our approach leverages the knowledge broadness of large language models for chemistry while distilling this knowledge into a lightweight model suitable for efficient and in-house use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

NETL Coal Energy Atlas: A Collection of Coal/Energy Related Maps

The NETL Coal Energy Atlas contains a comprehensive collection of coal and energy-related maps and graphics curated by the National Energy Technology Laboratory (NETL) Systems Analysis group. It serves as a living document providing an overview of the U.S. coal and energy sectors. The volume is structurally organized into six key thematic areas. Ultimately, the atlas functions as a modular baseline for data integration, allowing researchers to drill down into specific regional locations or customize geographic base layers for advanced systems analysis.

bituminous coal

Auger@TA: In-situ Cross-Calibration of the World's Largest Cosmic Ray Observatories

The Pierre Auger Observatory (Auger) and the Telescope Array (TA) are the world's two largest ultra-high-energy cosmic ray (UHECR) observatories. They operate in the Southern and Northern hemispheres, respectively, at similar latitudes but with distinct surface detector (SD) designs. A significant challenge in studying UHECR physics across the full sky is the apparent discrepancy in flux measurements between the two experiments. This discrepancy could arise from astrophysical differences and/or systematic effects related to their detector designs and sensitivities to extensive air shower components. To address this, the Auger@TA working group aims to cross-calibrate the two observatories with a self-triggering micro-Auger array within the TA array. This micro-array consists of eight Auger Surface Detector (SD) stations equipped with Water Cherenkov Detectors (WCDs) and AugerPrime Surface Scintillator Detectors. Seven SD stations, configured with a centered-1-PMT design, are arranged in a hexagonal pattern with one station in the center, with 1.5 km spacing, mirroring the Auger layout. The eighth station, which features a standard 3-PMT Auger station, is located in conjunction with a TA detector at the center of the hexagon, forming a triplet for high-statistics and low-uncertainty cross-calibration. A custom communication system that uses readily available components enables seamless communication between stations and remote access to each station through a central computer. The micro-array is now fully deployed, and initial data-taking is about to start. This presentation will detail the instrumentation, communication systems, central data acquisition system, expected performance of the micro-array, and preliminary results as appropriate.

Mocellin, Adriel G.B. [Colorado School of Mines]

Tethys Water Demand Data

U.S. water demand varies sharply by sector and region as land use, population, weather patterns, and economic activity co-evolve. High-resolution water demand data is required to capture these dynamics, support integrated energy-water-land modeling, and local-to-regional water scarcity assessments. This dataset contains gridded (1/8 degree), monthly, multi-sector water demand dataset for the contiguous United States (CONUS) covering 1980-2100 across eight future scenarios of human-Earth system change. The dataset covers irrigation, thermoelectric, municipal (public-supply and domestic), livestock, manufacturing, and mining demands, separately for withdrawals and consumption, and includes per-cell renewable vs. non-renewable water source attributions. The dataset is validated against the latest USGS 2010-2020 water-use data for the three largest water demand sectors (Domestic, Electricity, and Irrigation), with correlations ranging from 0.73-0.95 at the HUC6 scale. The two datasets largely agree on an aggregate basis with per-sector bias falling within +/-7%, but they disagree on the spatial allocation of water with individual HUC6 basins having normalized RMSE from 68-171% and median absolute percent difference from 37-86%. This dataset advances prior global products by combining state-resolved sectoral demands from GCAM-USA, future power-plant siting from the CERF model, and scenario-consistent high-resolution climate and population forcing data across the eight scenarios.

GCAM-USA

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY

North America’s Potential for an Environmentally Sustainable Nickel, Manganese, and Cobalt Battery Value Chain

The Detroit Big Three General Motors (GMs), Ford, and Stellantis predict that electric vehicle (EV) sales will comprise 40–50% of the annual vehicle sales by 2030. Among the key components of LIBs, the LiNixMnyCo1−x−yO2 cathode, which comprises nickel, manganese, and cobalt (NMC) in various stoichiometric ratios, is widely used in EV batteries. This review reveals NMC cathodes from laboratory research. Furthermore, this study examines the environmental effect of NMC cathode production for EV batteries (including coating technologies), encompassing aspects such as energy consumption, water usage, and air emissions. Although gaps persist in NMC cathode environmental assessments (NMC111, NMC532, NMC622, and NMC811), limited life cycle assessments “(LCA)” have been conducted. Most available data originate from Asia (primarily China), accounting for 85% of the production of EV LIB cathode materials. The concept of battery passports for data collection on LIB components has been proposed to facilitate material traceability as a system for ensuring a sustainable supply chain for critical minerals. The automotive industry’s shift to electrification necessitates a sustainable supply chain from mine to vehicle end-of-life. As the critical mineral supply moves from Asia to North America, environmentally friendly industrial methods must be studied to provide this supply chain direction.

25 ENERGY STORAGE

Verifiable Fuel Cycle Simulation

The Verifiable Fuel Cycle Simulation (VISION) is a system-dynamics model of the entire nuclear fuel cycle, from mining to disposal, using user-defined deployment scenarios, fuel recipes, and recycling parameters. It is written in the PowerSim simulation environment and utilized MS Excel to manage data inputs and outputs.

Hays, RossD (0000000184587539)

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING

Anion Data for the East River Watershed, Colorado (2014-2025)

The anion data for the East River Watershed, Colorado, consist of fluoride, chloride, sulfate, nitrate, and phosphate concentrations collected at multiple, long-term monitoring sites that include stream, groundwater, and spring sampling locations. These locations represent important and/or unique end-member locations for which solute concentrations can be diagnostic of the connection between terrestrial and aquatic systems. Such locations include drainages underlined entirely or largely by shale bedrock, land covered dominated by conifers, aspens, or meadows, and drainages impacted by historic mining activity and the presence of naturally mineralized rock. Developing a long-term record of solute concentrations from a diversity of environments is a critical component of quantifying the impacts of both climate change and discrete climate perturbations, such as drought, forest mortality, and wildfire, on the riverine export of multiple anionic species. Such data may be combined with stream gauging stations co-located at each monitoring site to directly quantify the seasonal and annual mass flux of these anionic species out of the watershed. This data package contains (1) a zip file (anion_data_2014_2025.zip) containing a total of 386 files: 387 data files of anion data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v7_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v7_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; and (4) a anion MDL fact sheet (anion_MDLs_202608 in PDF and docx formats). Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 47 locations containing anion data. Update on 2022-06-10: versioned updates to this dataset was made along with these changes: (1) updated anion data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, and (5) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2022-12-20: Updates were made to both the data files and reporting format specific files. Conversion issues affecting ER-PLM locations for anion data was resolved for the data files. Additionally, the flmd and dd files were updated to reflect the updated versions of these files. Available data was added up until 2022-03-14. Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-05-19. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-09-11. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until the end of WY2025 (September 30, 2025). An anion MDL document was included in this update.

54 ENVIRONMENTAL SCIENCES

Benchtop Autonomous Electrochemical Characterization System for Combinatorial Thin-Film Solid Oxide Electrodes

The design of materials for electrochemical energy conversion is complicated by a vast search space of candidate materials and multifaceted property requirements: multicarrier conductivity, stability, and catalytic activity are all necessary but rarely intersect. Although self-driving laboratories are rapidly rising to address such material optimization problems, the required infrastructure for integrated, large-scale robotic facilities can be cost-prohibitive. Here we develop and evaluate a closed-loop measurement system for efficient screening of proton-conducting oxide electrodes for ceramic fuel cells and electrolyzers, building on top of an existing benchtop instrument and integrating techniques for rapid impedance measurement and automated analysis. This system exemplifies a “minimum viable” self-driving implementation that can deliver substantial benefits with relatively simple infrastructure. Combinatorial thin-film microelectrode libraries are characterized with a recently developed joint time-domain and frequency-domain impedance measurement technique, which provides an order-of-magnitude acceleration relative to conventional impedance spectroscopy. The distribution of relaxation times is extracted from impedance data and analyzed without human intervention. These results feed an active learning and Bayesian optimization process that learns to predict electrochemical impedance as a function of material composition, measurement temperature, oxygen partial pressure, and electrical bias, which further reduces the screening time by tenfold with optimized experimental sequences. We apply this system to Ba⁡(Co,Fe,Zr,Y)⁢O 3−𝛿 combinatorial libraries and evaluate its effectiveness for learning material property trends and optimizing expensive-to-evaluate properties such as activation energy. This offers insights into key methodological aspects of practical autonomous experimentation, including surrogate model validation, cost-aware acquisition functions, and high-throughput data interpretation. Our results demonstrate the efficacy of the system for rapidly gathering information, but also highlight real-world experimental challenges of thin-film degradation and numerical instability in surrogate models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Techno Economic Analysis for Production of Renewable Natural Gas and Value-Added Chemicals from Forest Biomass Residues (CRADA Final Report)

West Biofuels - in collaboration with the University of California San Diego, the University of California Davis, the National Laboratory of the Rockies (NLR), the Colorado School of Mines’ Center for Hydrate Research, Placer County Air Pollution Control District, the Sierra Business Council, and the Southern California Gas Company (SoCalGas)-will demonstrate an innovative pathway to convert forest biomass to renewable gas (RG). The project, Production of Pipeline Grade Renewable Natural Gas and Value-Added Chemicals from Forest Biomass Residues, will utilize an existing pilot-scale gasification system and data and lessons learned from a lab-scale catalytic reactor design, develop, and demonstrate an integrated pilot-scale RG process with a scaled-up catalytic reactor and hydrate-based gas separation process. This robust and simple process uses a commercial proven gasifier, a commercially available catalyst that can tolerate gas contaminants, and a gas separation process that simplifies the downstream processing. Current testing has shown that the process can yield a large amount of methane rich renewable gas in addition to a mixture of propanol, ethanol and other higher alcohols and the relative amounts can be controlled by shifting the process conditions. This process is novel and groundbreaking because the product mixture of RG and valuable alcohols byproducts makes RG production economically feasible at $\$$12 per million British thermal units (MMBtu) or less using forest biomass from high hazard zones.

09 BIOMASS FUELS