Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning for engineering applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Deep Generative Models in Energy System Applications: Review, Challenges, and Future Directions

In recent years, with the advent of mature machine learning products like ChatGPT, Stable Diffusion, and Sora, the world has witnessed tremendous changes driven by the rapid development of generative artificial intelligence (GAI). Beyond applications in text, speech, image, and video creation, deep generative models (DGMs) underpinning these cutting-edge technologies have also been employed by domain researchers to address scientific and engineering challenges. This paper aims to fill a gap in the research community by providing a systematic review of how DGMs have been utilized in energy system applications. After introducing four most popular DGMs, we review and categorize 196 research articles into five focus areas: data generation, forecasting, situational awareness, modeling, and optimal decision-making. Through this classification, we uncover trends in how DGMs are employed for each type of problem, highlighting GAI techniques that contribute to breakthroughs over traditional methods. We discuss limitations in existing literature, engineering challenges, and propose future directions, all tailored to the unique nature of problems in energy system engineering. Our goal is to offer insights for energy system domain researchers, providing a comprehensive view of existing studies and potential future opportunities.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Deep learning based x-ray spectrometer for high repetition rate characterization of betatron radiation

Betatron radiation produced from a laser-wakefield accelerator is a broadband, hard x-ray (>1 keV) source that has been used in a variety of applications in medicine, engineering, and fundamental science. Further development and optimization of stable, high repetition rate (HRR) (>1 Hz) betatron sources will provide a means to extend their application base to include single-shot dynamical measurements of ultrafast processes or dense materials. Recent advances in laser technology used in such experiments have enabled increases in shot-rate and system stability, providing improved statistical analysis and detailed parameter scans. However, unique challenges exist at high repetition rate, where data throughput and source optimization are now limited by diagnostic acquisition rates and analysis. Here, we present the development of a machine-learning algorithm for the real-time analysis of betatron radiation. We report on the fielding of this deep learning algorithm for online source characterization at the Institut National de la Recherche Scientifique's Advanced Laser Light Source. By fine-tuning an algorithm originally trained on a fully synthetic dataset using a subset of experimental data, the algorithm can predict the betatron critical energy with a percent error of 7.2 % with a reconstruction time of 1.5 ms, providing a valuable tool for real-time, multi-objective optimization at HRR.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Overview of RFID Applications Utilizing Neural Networks

As Radio Frequency Identification (RFID) methods continue to evolve to higher levels of complexity, one form of machine learning is making its appearance. The use of Neural Networks (NN) in the RFID field is steadily increasing, and in the fields of localization and activity recognition, promising results are being shown from a variety of research. RFID applications fall primarily under two types of problems including regression and classification. We analyze RIFD localization techniques which fall under regression, and activity recognition which falls under classification. Many works don’t classify themselves as activity recognition methods, but because they fall under the classification category, we still consider them as activity recognition techniques. This research overviews the Neural Network models in the localization field based on whether they can perform independently of the environment in which they were tested. For activity recognition and accessory fields, the major methods involve tag-based and tag-free approaches. In conclusion, after the models are surveyed, a comparison study is given to examine what may be the cause for increased accuracy between different Neural Network models.

42 ENGINEERING↗

Recent Progress on Surface Water Quality Models Utilizing Machine Learning Techniques

Surface waterbodies are heavily exposed to pollutants caused by natural disasters and human activities. Empowering sensor technologies in water quality monitoring, sufficient measurements have become available to develop machine learning (ML) models. Numerous ML models have quickly been adopted to predict water quality indicators in various surface waterbodies. This paper reviews 78 recent articles from 2022 to October 2024, categorizing water quality models utilizing ML into three groups: Point-to-Point (P2P), which estimates the current target value based on other measurements at the same time point; Sequence-to-Point (S2P), which utilizes previous time series data to predict the target value at one time point ahead; and Sequence-to-Sequence (S2S), which uses previous time series data to forecast sequential target values in the future. The ML models used in each group are classified and compared according to water quality indicators, data availability, and model performance. Widely used strategies for improving performance, including feature engineering, hyperparameter tuning, and transfer learning, are recognized and described to enhance model effectiveness. The interpretability limitations of ML applications are discussed. This review provides a perspective on emerging ML for surface water quality models.

machine learning (ML)↗

A Workflow to Optimize Fast Neutron Irradiation in A Thermal Neutron Spectrum Test Reactor Leveraging Open-Source Tools

The Advanced Test Reactor (ATR) located at Idaho National Laboratory (INL) is one of the key nuclear engineering research and testing facilities within the US Department of Energy (DOE). The ATR is one of few high-power research reactors in the world with different application including accelerated testing of nuclear fuel, materials irradiation in a very high neutron flux environment, and medical radioisotope production [1]. Also, the ATR offers opportunities for testing fast spectrum fission and fusion reactor materials. The key challenges in this area are in further detailing and optimizing a fast spectrum environment within a thermal test reactor. This challenge involves researching, developing, and testing novel concepts for the multiplying of neutron populations into ever higher energy spectra in high flux test reactors like ATR. The main objective of this work is to investigate candidate materials for establishing a fast neutron experiment irradiation in thermal neutron spectrum test reactors which can be accomplished by filtering thermal and epithermal neutrons and boosting fast neutrons at designated irradiation positions. However, adding these filters will render the neutron spectrum and the criticality of the system. The selection of the thickness and material layers should be accomplished by developing an optimization design algorithm that is applicable for ATR to enhance the fast neutron spectrum irradiation utilizing high-fidelity Monte Carlo methods along with advanced machine learning capabilities. This paper presents workflow for design optimization to enhance fast neutron irradiation in the ATR. The workflow leverages open-source tools to develop an algorithm that is viable to ATR and can be leveraged in other reactors. The following sections discuss the development of the experiment design optimization workflow and its application to ATR irradiation positions.

42 - ENGINEERING↗

Convergent Manufacturing of Large-Scale Components for Nuclear Applications, via Additive Manufacturing and Powder Metallurgy Hot Isostatic Pressing

Powder metallurgy (PM)–hot isostatic pressing (PM-HIP) has long been recognized as a powerful route for producing fully dense, near net shape metallic components. By consolidating powders under high temperature and pressure, HIP provides isotropic properties, uniform microstructures, and scalability to complex geometries that are vital for sectors such as aerospace, energy, and nuclear power. Yet despite these advantages, the technology has remained constrained by costly trial and error canister fabrication, limitations of conventional forging, and incomplete knowledge about how the canister design influences final part properties. Additive manufacturing (AM), by contrast, thrives on design freedom and geometric flexibility but struggles with speed, scalability, and cost when applied to very large structures. The research presented in this report investigated how a convergent manufacturing approach, combining AM with PM-HIP, can merge the strengths of both technologies, leveraging AM’s flexibility for canister design and HIP’s consolidation capability to deliver reliable, large, and complex parts. The work progressed through three case studies that built on one another in scale and complexity. Small cylindrical canisters fabricated by conventional methods, laser powder bed fusion, and directed energy deposition were filled with stainless steel powders and subjected to HIP. The resulting parts demonstrated near-full density and mechanical properties on par with wrought stainless steel, showing for the first time that AM canisters can be a direct substitute for conventional ones without sacrificing quality. The next step involved a medium-scale, noncentrosymmetric T-valve, which is an enclosed, multibranch geometry that tested the limits of AM + PM-HIP integration. The T-valve achieved predictable shrinkage and uniform densification, confirming feasibility for enclosed designs. However, this study also revealed oxide inclusions and interfacial challenges at the AM + HIP boundary, underscoring the critical importance of controlling interface chemistry and employing robust, in situ strategies, such as melt pool monitoring and thermal monitoring, coupled with nondestructive evaluation techniques such as x-ray computed tomography. Finally, the effort culminated in fabricating a large-scale impeller weighing nearly 2000 lb and spanning 5 ft in diameter. Produced via multirobot wire arc AM and hot isostatic pressed to near-full density, the impeller validated industrial-scale feasibility. Predictive models closely matched experimental shrinkage, tensile properties were spatially uniform across the component, and the AM + PM-HIP interface proved mechanically sound despite the presence of oxide-decorated prior particle boundaries. This large-scale demonstration is a major milestone, showing that hybrid AM + PM‑HIP can reliably deliver components at reactor-relevant scales. Collectively, these studies charted a logical pathway: small-scale work built scientific confidence, medium-scale work highlighted opportunities and challenges, and large-scale work proved industrial impact. The overarching conclusion of this report is that AM + PM-HIP should not be seen as a replacement for forging but as a complementary pathway that provides the US with flexibility, resilience, and new options for manufacturing nuclear-grade components. Looking ahead, several directions emerge as critical to sustaining progress. Predictive modeling must become faster, more accessible, and more accurate, with digital twins and machine learning reducing reliance on trial and error. Powders and alloys must be optimized for HIP, with improved cleanliness, reduced oxides, and tailored chemistries that enhance creep, fatigue, and irradiation resistance. Interfaces between AM and HIP regions must be better engineered through coatings, machining strategies, and surface treatments to mitigate oxide formation and ensure reliable bonding to explore opportunities for HIP of targeted compositional parts, as well as multimaterial HIP cladding applications. Monitoring and nondestructive evaluation need to expand, incorporating multimodal sensors, x-ray computed tomography, and real-time data integration through platforms such as Pelican. At the same time, the pathway to industrial adoption requires techno-economic analysis, machinability studies, and qualification frameworks aligned with industry and regulatory standards. Finally, workforce and academic engagement must be strengthened. Programs that train technicians and engineers for US Navy and US Department of Energy manufacturing challenges should be paired with academic partnerships to support fundamental research, with open sharing of non-export-controlled data to accelerate innovation and build the next generation of experts. In conclusion, this report demonstrates that hybrid AM + PM-HIP is scientifically viable and strategically important. By combining the design agility of AM with the consolidation strength of HIP and embedding modeling, monitoring, and workforce development, this approach provided a transformative new capability for US manufacturing. The path forward is clear: hybrid AM + PM-HIP is not just a promising research direction but is also potentially an industrially relevant pathway that can reshape how nuclear-grade components are designed, qualified, and deployed.

36 MATERIALS SCIENCE↗

Benchmarking machine learning strategies for phase-field problems

Abstract We present a comprehensive benchmarking framework for evaluating machine-learning approaches applied to phase-field problems. This framework focuses on four key analysis areas crucial for assessing the performance of such approaches in a systematic and structured way. Firstly, interpolation tasks are examined to identify trends in prediction accuracy and accumulation of error over simulation time. Secondly, extrapolation tasks are also evaluated according to the same metrics. Thirdly, the relationship between model performance and data requirements is investigated to understand the impact on predictions and robustness of these approaches. Finally, systematic errors are analyzed to identify specific events or inadvertent rare events triggering high errors. Quantitative metrics evaluating the local and global description of the microstructure evolution, along with other scalar metrics representative of phase-field problems, are used across these four analysis areas. This benchmarking framework provides a path to evaluate the effectiveness and limitations of machine-learning strategies applied to phase-field problems, ultimately facilitating their practical application.

36 MATERIALS SCIENCE↗

VoroClust

SAND2025-11465O VoroClust, also known as Voronoi Clustering, is a fast, density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. It operates as quickly as distance-based clustering methods while effectively capturing complex regional geometries, matching the performance of current density-based methods. VoroClust employs a data-centered sphere cover to reduce computational demands while preserving data topology. It propagates clusters outward from local density peaks. Although supervised machine learning is powerful for applications like image classification and segmentation, it requires comprehensive, consistent datasets, which many applications lack. Unsupervised clustering algorithms analyze the structure of each dataset rather than relying on similarities with other examples, making them well-suited for practical applications with insufficient or inappropriate data for supervised learning. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ebeida, Mohamed [Sandia National Lab. (SNL-CA), Li↗

Charting the state of GEMs in microalgae: progress, challenges, and innovations

Genome-scale metabolic models (GEMs) provide a systems-level framework for understanding and engineering microalgal metabolism. This review explores the evolution of GEMs in microalgae, highlighting advances in light modeling, automation, and multi-omics integration. Special emphasis is placed on Chlamydomonas reinhardtii as a model species. Limitations of current models, particularly for microalgae, are discussed, alongside promising developments in dynamic modeling and machine learning. Together, these innovations chart a path toward more predictive, adaptable GEMs that can accelerate biotechnological applications of microalgae in sustainable production systems.

Plant Sciences↗

Application of physics-informed neural networks (PINNs) solution to coupled thermal and hydraulic processes in silty sands

Abstract The accurate modeling of water and heat transport in soils is crucial for both geo-environmental and geothermal engineering. Traditional modeling methods are problematic because they require well-defined boundaries and initial conditions. Recently, physics-informed neural networks (PINNs), which incorporate partial differential equations (PDEs) to solve forward and inverse problems, have attracted increasing attention in machine learning research. In this study, we applied PINNs to tackle hydraulic and thermal transport coupling forward problems in silty sands. A fully connected deep neural network was utilized for training. This neural network model leverages automatic differentiation to apply the governing equations as constraints, based on the mathematical approximations established by the neural network itself. We conducted forward problems and compared the solutions derived from PINNs with those from Finite Element Method (FEM) simulations. The forward problem results demonstrate the PINNs model’s capability in predicting hydraulic transport, heat transport, and thermal–hydraulic coupling in silty sands under various boundary conditions. The PINNs exhibited great performance in simulating the thermal–hydraulic coupling problem. The accuracy of the PINNs solutions shows its potential for simulation in geotechnical engineering.

Feng, Yuan↗

Multiparameter optical fiber sensing for energy infrastructure through nanoscale light–matter interactions: From hardware to software, science to commercial opportunities

Monitoring of energy infrastructure through robust yet economical sensing platforms is becoming an area of increased importance, with ubiquitous applications including the electrical grid, natural gas and oil transportation pipelines, H2 infrastructure (storage and transportation), carbon storage, power generation, and subsurface environments. Plasmonic and functional nanomaterial enabled fiber optic sensors show excellent promise for a wide range of sensing applications due to their versatility to be engineered for specific analytes of interest while retaining inherent advantages of the optical fiber sensor platform. Through the design of novel sensing layers, the optical transduction mechanism and wavelength dependence can also be tailored for ease of integration with low-cost interrogation systems enabling an inexpensive yet highly functional optical fiber sensing platform. In addition, recent advances in artificial intelligence and machine learning theoretical methods have been leveraged to simultaneously extract multiple parameters through multi-wavelength interrogation such that unique wavelengths can also serve as unique sensing elements, analogous to electronic nose sensor technologies. The concept of an optical fiber based “photonic nose” via multiple interrogation wavelengths and/or sensor nodes offers a compelling platform technology to realize multiparameter speciation of chemical analytes within complex gas mixtures. In this Perspective, we further generalize the notion of multiparameter sensing through the novel “photonic nervous system” concept based upon low-cost, functionalized optical fiber sensor probes monitoring a variety of distinct analyte classes (physical, chemical, electromagnetic, etc.) simultaneously to provide broad situational awareness via integrated sensors.

Su, Yang-Duan (ORCID:0000000214820902)↗

Gradient flow based phase-field modeling using separable neural networks

Allen–Cahn equation is a reaction–diffusion equation and is widely used for modeling phase separation. Machine learning methods for solving the Allen–Cahn equation in its strong form suffer from inaccuracies in collocation techniques, errors in computing higher-order spatial derivatives, and the large system size required by the space–time approach. To overcome these challenges, we propose solving the gradient flow of the Ginzburg–Landau free energy functional, which is equivalent to the Allen–Cahn equation, thereby avoiding the second-order spatial derivatives associated with the Allen–Cahn equation. A minimizing movement scheme is employed to solve the gradient flow problem, eliminating the complexities of a space–time approach. We utilize a separable neural network that efficiently represents the phase field through low-rank tensor decomposition. As we use the minimizing movement scheme to numerically solve the gradient flow problem, we thus, refer to the proposed method as the Separable Deep Minimizing Movement (SDMM) method. The evaluation of the functional in the minimizing movement scheme using the Gauss quadrature technique bypasses the inaccuracies associated with collocation techniques traditionally used to solve partial differential equations. A hyperbolic tangent transformation is introduced on the phase field prior to the evaluation of the functional to ensure that it remains strictly bounded within the values of the two phases. For this transformation, theoretical guarantee for energy stability of the minimizing movement scheme is established. Our results suggest that this transformation helps to improve the accuracy and efficiency significantly. The proposed method resolves the challenges faced by state-of-the-art machine learning techniques, outperforming them in both accuracy and efficiency. It is also the first machine learning method to achieve an order of magnitude speed improvement over the finite element method. In addition to its formulation and computational implementation, several case studies illustrate the applicability of the proposed method.

42 ENGINEERING↗

Micropolar deep material network

This study extends the Deep Material Network (DMN), a physics-informed machine learning framework, to predict the homogenized mechanical response of composite materials with micropolar (Cosserat-type) constitutive behavior. This extension incorporates microstructure-dependent size effects, enabling accurate, efficient, and size-aware predictions for composites with complex internal architectures. While traditional, direct numerical simulation micropolar models effectively capture size effects by introducing extra local degrees of freedom, they bring significant computational challenges, particularly for multiscale analyses relevant to engineering applications. The micropolar DMN developed in this paper achieves high accuracy while significantly reducing computation time compared to micropolar direct numerical simulations. This advancement enables multiscale analyses and parameter studies that were previously impractical, such as high-cycle fatigue simulations and comprehensive investigations of internal length scale effects notably in size-dependent plastic response and the optimization of lattice structures. By uniting microstructure-sensitive modeling, physics-driven learning, and scalable surrogate modeling, the micropolar DMN paves the way for accelerated material design, large-scale parametric studies, and the reliable incorporation of size-dependent effects across a wide range of engineering applications, including optimization and next-generation composite design.

36 MATERIALS SCIENCE↗

Sandia Image Labeling Tool (SILT)

SAND2025-01840O Sandia Image Labeling Tool (SILT) is a Python tool that labels images for machine learning and other applications. SILT uploads JSON files to create a template for labeling an image. SILT can then upload images, including images that are tens of GB large and dynamically loads them in a manner that a user, with a very modest spec laptop, can handle. It then saves the label as a JSON using the template. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Pitts, Todd↗

AnisONet: A deep neural operator-based anisotropic permeability upscaler from pore to Darcy scale

Directional permeability variations, which govern directional fluid flow in porous media with anisotropy, are important to accurately predict flow behavior, reactive transport, and fluid–solid interactions for various processes such as enhanced geothermal systems, energy storage devices, and biological systems. However, the intricate architecture of porous media makes it difficult to predict directional permeabilities. In this work, we present a novel machine learning (ML) framework, AnisONet, built upon an integration of a convolutional neural network, Swin transformer, and the deep operator network architecture, designed to predict anisotropic permeability and upscale predictions to larger spatial domains. First, AnisONet was evaluated with three classes of two-dimensional (2D) porous media, including synthetic circular and elliptical grains and natural sandstone grains from micro-computed tomography images. A lattice Boltzmann model (LBM) was used to calculate directional permeabilities at every 10° angle, producing 19 data points per image of porous media. AnisONet is then trained to predict permeability as a function of rotation angle. AnisONet showed strong predictive capability of directional permeability. Second, we tested our model for five upscaling cases with a large image size in the finite-element method (FEM) for 2D Darcy flow with various permeability tensor construction methods. Overall, upscaled permeability tensors in FEM simulations produce a reasonably good match with LBM results, highlighting the importance of selecting appropriate tensor formation strategies for accurate permeability upscaling. AnisONet, as a directional permeability estimator, could be further developed for more complex geometries, with the potential to develop a foundational ML model for various applications in porous media.

42 ENGINEERING↗

Python Library for Monte Carlo Simulations with Ab Initio and Machine-Learned Interatomic Potentials

There is a growing need in the simulation community for software that provides a transparent, reproducible, usable, and extensible (TRUE) Monte Carlo (MC) simulation framework employing energies from ab initio methods and machine-learning interatomic potentials (MLIPs). We introduce a Python library (ASE-MC) that adds Monte Carlo functionality to the Atomic Simulation Environment (ASE) package. Now, we can combine the powerful tools used to build systems and perform ab initio and MLIP in ASE with MC simulation algorithms to sample the configurational space with a concise Python script. After presenting the design philosophy, we demonstrate the flexibility of our approach using selected examples. These example simulations include liquid water described with a message-passing MLIP in the canonical and isothermal–isobaric ensembles, sampling the characteristic dihedral angle of biphenyl and comparing an MLIP to first-principles calculations, and a grand canonical Monte Carlo simulation of ammonia adsorption on Pt(111). These examples showcase the main features of the software, which include flexibility in the choice of ab initio or MLIP engine, ab initio or MLIP grand canonical MC with cavity bias insertions and deletions, the ability to add custom MC moves to the move set, and how users can condense complex MC workflows into a single Python script. Finally, this library serves as a framework for reproducible Monte Carlo simulations, facilitating easy reproduction of the work and application to new systems.

97 MATHEMATICS AND COMPUTING↗

Data-driven Community-centered Resilient Assessment and Planning Toolkit for Nexus of Energy and Water (DCRAPT-NEW)

Urban areas, including Detroit and Pittsburgh, have suffered significant dual outages of the electrical and water infrastructure in the past decade due, in part, to the increasing number of extreme weather events. With increasing temperatures and rainfall intensity, these regions need to prepare for increasing extreme events through community-based energy and water resilience analysis, planning, and enhancement. This project developed a suite of open-source, open-access, community-centered, data-driven assessment and distributed energy resource (DER) and planning tools for energy and water resilience enhancement in urban areas. Through establishing a multi-level community awareness and engagement mechanism and a comprehensive collection of power outage and flooding data, an innovative group of community energy and water resilience assessment and planning tools have been developed for a wide range of users with differing and variable sets of data available to them. The developed tools include (1) DOE EAGLE-I data-driven, deep-learning assisted resilience assessment and DER planning tools at the county level with socioeconomic factors incorporated; (2) Utility annual power outage data-driven tools for long term resilience assessment and DER planning and 15-min power outage data-driven tools for short term resilience assessment and planning; (3) Detailed engineering tools for energy and water systems resilience assessment and planning when the system topology and component fragility curves are available; (4) Alternative Resiliency Metric Calculation that extracts and separates outage and restoration processes; and (5) Co-optimization tools that evaluate the resilience of the power and sewage system and allow users to conduct joint planning with energy and wastewater systems. The developed tools provide planners, decision-makers, and stakeholders with powerful capabilities to systematically evaluate system/community resilience and optimal and actionable guidance for enhancing resilience while prioritizing DER investments. The tools have been used and validated in Detroit and Pittsburgh and can be used in other areas of the nation. In addition, this project will (1) advance the knowledge and applications of machine-learning methods in analyzing and fusing different layers of information and generating meaningful data points such as generating rare weather events; (2) significantly improve the energy and water resilience of the identified communities in Detroit and Pittsburgh and prepare for more frequent and severe weather conditions; (3) help communities assess extreme weather event impacts and address short-term and long-term resilience-related issues The developed tools have been made public via GitHub and demonstrated to community stakeholders and utility companies via the two annual workshops and numerous community engagement meetings. The project outcomes are also disseminated through publications in various journals and conference proceedings, and presentations at top conferences.

13 HYDRO ENERGY↗

Molecular property prediction for very large databases with natural language processing: a case study in ionic liquid design

The prospect of using artificial intelligence (AI) to accurately screen very large databases of compounds for multiple properties has yet to be realized. Here, we explore this possibility using ionic liquids (ILs) which offer unique physicochemical properties and excellent tunability, making them highly versatile solvents for various research applications. Screening millions of potential ILs for the best perfomance for use in specific tasks with experimental methods alone however, is impractical. Further, traditional’ physics-based computational chemistry is hindered by high computational cost. To address this challenge, we leverage a natural language processing (NLP)-based molecular embedding technique with advanced machine learning (ML) models to predict seven key IL properties: viscosity, density, ionic conductivity, surface tension, melting temperature, toxicity, and water solubility. Comprehensive datasets for these properties are obtained, then NLP featurization with Mol2vec is compared with other featurization techniques such as 2D Morgan fingerprints, and 3D quantum chemistry-derived sigma profiles. NLP-based featurization exhibited the best predictive performance, achieving the highest R 2 and lowest RMSE values for all the studied IL properties. Further, we present case studies of how ILs might be screened using combined property criteria for practical cases – lignocellulosic biomass processing, CO 2 capture, and optimal electrolytes for batteries – screening a novel database of ∼10.6 million generated feasible ILs. The results introduce NLP as a powerful tool for engineering many designer solvents with desirable properties for task specific applications.

Mohan, Mood [Oak Ridge National Laboratory (ORNL),↗