Search NASA⌕ Search

SEARCH · Search NASA

Results for “SMILES”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Preface to the Special Issue on Modeling and Data Analysis Methods for the SMILE mission

The SMILE (Solar wind Magnetosphere Ionosphere Link Explorer) project (http://www.nssc.cas.cn/smile/, https://www.cosmos.esa.int/web/smile/mission) is a joint spacecraft mission of the European Space Agency (ESA) and the Chinese Academy of Sciences (CAS) with an expected launch in 2025. SMILE aims to study the global interactions of solar wind–magnetosphere–ionosphere innovatively by imaging the Earth’s magnetosheath and cusps in soft X-rays and the northern auroral region in ultraviolet (UV) while simultaneously measuring plasma and magnetic field parameters in the solar wind and magnetosheath along a highly-elliptical and highly-inclined orbit. This special issue is composed of 22 articles, presenting recent progress in modeling and data analysis techniques developed for the SMILE mission. In this preface, we categorize the articles into the following seven topics and provide brief summaries: (1) instrument descriptions of the Soft X-ray Imager (SXI), (2) numerical modeling of the X-ray signals, (3) data processing of the X-ray images, (4) boundary tracing methods from the simulated images, (5) physical phenomena and a mission concept related to the scientific goals of SMILE-SXI, (6) studies of the aurora, and (7) ground-based support for SMILE.

SMILE↗

Astro Smile

This is a humorous look at life aboard the Space Shuttle.

Source record↗

Hybrid-Vlasov Simulation of Soft X-Ray Emissions at the Earth’S Dayside Magnetospheric Boundaries

Solar wind charge exchange produces emissions in the soft X-ray energy range which can enable the study of near-Earth space regions such as the magnetopause, the magnetosheath and the polar cusps by remote sensing techniques. The Solar wind–Magnetosphere–Ionosphere Link Explorer (SMILE) and Lunar Environment heliospheric X-ray Imager (LEXI) missions aim to obtain soft X-ray images of near-Earth space thanks to their Soft X-ray Imager (SXI) instruments. While earlier modeling works have already simulated soft X-ray images as might be obtained by SMILE SXI during its mission, the numerical models used so far are all based on the magnetohydrodynamics description of the space plasma. To investigate the possible signatures of ion-kinetic-scale processes in soft X-ray images, we use for the first time a global hybrid-Vlasov simulation of the geospace from the Vlasiator model. The simulation is driven by fast and tenuous solar wind conditions and purely southward interplanetary magnetic field. We first produce global X-ray images of the dayside near-Earth space by placing a virtual imaging satellite at two different locations, providing meridional and equatorial views. We then analyze regional features present in the images and show that they correspond to signatures in soft X-ray emissions of mirror mode wave structures in the magnetosheath and flux transfer events (FTEs) at the magnetopause. Our results suggest that, although the time scales associated with the motion of those transient phenomena will likely be significantly smaller than the integration time of the SMILE and LEXI imagers, mirror-mode structures and FTEs can cumulatively produce detectable signatures in the soft X-ray images. For instance, a local increase by 30% in the proton density at the dayside magnetopause resulting from the transit of multiple FTEs leads to a 12% enhancement in the line-of-sight- and time integrated soft X-ray emissivity originating from this region. Likewise, a proton density increase by 14% in the magnetosheath associated with mirror-mode structures can result in an enhancement in the soft X-ray signal by 4%. These are likely conservative estimates, given that the solar wind conditions used in the Vlasiator run can be expected to generate weaker soft X-ray emissions than the more common denser solar wind. These results will contribute to the preparatory work for the SMILE and LEXI missions by providing the community with quantitative estimates of the effects of small-scale, transient phenomena occurring on the dayside.

magnetosphere↗

Ensemble Spread Behavior in Coupled Climate Models: Insights From the Energy Exascale Earth System Model Version 1 Large Ensemble

AbstractAssessing uncertainty in future climate projections requires understanding both internal climate variability and external forcing. For this reason, single‐model initial condition large ensembles (SMILEs) run with Earth System Models (ESMs) have recently become popular. Here we present a new 20‐member SMILE with the Energy Exascale Earth System Model version 1 (E3SMv1‐LE), which uses a “macro” initialization strategy choosing coupled atmosphere/ocean states based on inter‐basin contrasts in ocean heat content (OHC). The E3SMv1‐LE simulates tropical climate variability well, albeit with a muted warming trend over the twentieth century due to overly strong aerosol forcing. The E3SMv1‐LE's initial climate spread is comparable to other (larger) SMILEs, suggesting that maximizing inter‐basin ocean heat contrasts may be an efficient method of generating ensemble spread. We also compare different ensemble spread across multiple SMILEs, using surface air temperature and OHC. The Community Earth system Model version 1, the only ensemble which utilizes a “micro” initialization approach perturbing only atmospheric initial conditions, yields lower spread in the first ∼30 years. The E3SMv1‐LE exhibits a relatively large spread, with some evidence for anthropogenic forcing influencing spread in the late twentieth century. However, systematic effects of differing “macro” initialization strategies are difficult to detect, possibly resulting from differing model physics or responses to external forcing. Notably, the method of standardizing results affects ensemble spread: control simulations for most models have either large background trends or multi‐centennial variability in OHC. This spurious disequlibrium behavior is a substantial roadblock to understanding both internal climate variability and its response to forcing.

Stevenson, Samantha↗

Thermochemical Data for Furan-based Monomer Candidates for Frontal Ring-Opening Metathesis Polymerization (FROMP)

This dataset includes 471 furan-based monomer candidates for frontal ring-opening metathesis polymerization (FROMP) and relevant thermochemistry as calculated with density functional theory (DFT). The monomer candidates were combinatorically enumerated using Diels-Alder reactions of furan derivatives as dienes and four types of dienophiles (alkenes, alkynes, allenes, and benzynes). Common substituents were enumerated for the dienophile classes, and methyl substitution on the diene was explored. We used the SMILES arbitrary target specification (SMARTS) language to produce monomers and ring-opened structures from diene and dienophile precursor SMILES, and we studied the ring-opening reaction using a homodesmotic equation with ethene. RDKit conformers were initially generated from SMILES, then optimized with GFN2-xTB. The two conformers lowest in energy were then optimized with DFT using the wb97x-D3 functional, def2-TZVP basis set, and def2/J auxiliary basis set. Gibbs free energy corrections were obtained through frequency calculations. Structures with imaginary frequencies below -50 cm^{-1} were excluded from this work, and smaller imaginary modes were flipped to be positive for free energy calculations. Modes below 50 cm^{-1} were treated with the modified rigid rotor approximation, and all thermochemical values were calculated at T=200C. The CSV file contains the monomer SMILES, the free energy of reaction for Diels-Alder addition (G_DA_200), and the enthalpy of the ring-opening reaction (H_RO_200). All energies are given in kcal/mol. An interactive HTML is also included to visualize the monomers in this dataset.

Chua, Lauren↗

Mshpy23: A User-Friendly, Parameterized Model of Magnetosheath Conditions

Lunar Environment heliospheric X-ray Imager (LEXI) and Solar wind – Magnetosphere - Ionosphere Link Explorer (SMILE) will observe magnetosheath and its boundary motion in soft X-rays for understanding magnetopause reconnection modes under varioussolar wind conditions after their respective launches in 2024 and 2025. Magnetosheath conditions, namely, plasma density, velocity, and temperature, are key parameters for predicting and analyzing soft X-ray images from the LEXI and SMILE missions. We developed a user-friendly model of magnetosheath that parameterizes number density, velocity, temperature, and magnetic field by utilizing the global Magnetohydrodynamics (MHD) model as well as the pre-existing gas-dynamic and analytic models. Using this parameterized magnetosheath model, scientists can easily reconstruct expected soft X-ray images and utilize them for analysis of observed images of LEXI and SMILE without simulating the complicated global magnetosphere models. First, we created an MHD based magnetosheath model by running a total of 14 OpenGGCM global MHD simulations under 7 solar wind densities (1, 5, 10, 15, 20, 25, and 30cm−3) and 2 interplanetary magnetic field BZ components (± 4nT), and then parameterizing the results in new magnetosheath conditions. We compared the magnetosheath model result with THEMIS statistical data and it showed good agreement with a weighted Pearson correlation coefficient greater than 0.77, especially for plasma density and plasma velocity. Second, we compiled a suite of magnetosheath models incorporating previous magnetosheath models (gas-dynamic, analytic), and did two case studies to test the performance. The MHD based model was comparable to or better than the previous models while providing self consistency among the magnetosheath parameters. Third, we constructed a tool to calculate a soft X-ray image from any given vantage point, which can support the planning and data analysis of the aforementioned LEXI and SMILE missions. A release of the code has been uploaded to a Github repository.

Magnetosheath↗

A Comprehensive Machine Learning Model for Metal–Ligand Binding Prediction: Applications in Chemistry and Biology

A machine-learning (ML) model that predicts metal–ligand binding constants was developed using the open-source Chemprop software. The model was trained on over 30,000 experimental log K 1 values, which include both protonation and metal–ligand stability constants, comprising over 3500 ligands and 10 2 metal ions from 73 total elements, thus generalizing beyond existing limited approaches, which focus only on specific metals or ligand families. The best-performing model included a combination of SMILES-based molecular representations along with descriptors for the metal ion and experimental conditions. It had an external test R 2 value of 0.942, and MAE value of 0.834. A “SMILES-only” simpler version also produced accurate predictions and preserved the binding trends, serving as a quick and easily accessible alternative for users without computational expertise. The SMILES-only model performed comparably to density functional theory (DFT) calculations but utilized a fraction of the computational resources. The model was successfully applied across diverse domains, including bioinorganic chemistry, heavy metal remediation, and sensor development and demonstrated its effectiveness as a rapid and reliable screening tool for both academic and industrial uses.

Ligands↗

Effects of Orientation on Recognition of Facial Affect

The ability to discriminate facial features is often degraded when the orientation of the face and/or the observer is altered. Previous studies have shown that gross distortions of facial features can go unrecognized when the image of the face is inverted, as exemplified by the 'Margaret Thatcher' effect. This study examines how quickly erect and supine observers can distinguish between smiling and frowning faces that are presented at various orientations. The effects of orientation are of particular interest in space, where astronauts frequently view one another in orientations other than the upright. Sixteen observers viewed individual facial images of six people on a computer screen; on a given trial, the image was either smiling or frowning. Each image was viewed when it was erect and when it was rotated (rolled) by 45 degrees, 90 degrees, 135 degrees, 180 degrees, 225 degrees and 270 degrees about the line of sight. The observers were required to respond as rapidly and accurately as possible to identify if the face presented was smiling or frowning. Measures of reaction time were obtained when the observers were both upright and supine. Analyses of variance revealed that mean reaction time, which increased with stimulus rotation (F=18.54, df 7/15, p (is less than) 0.001), was 22% longer when the faces were inverted than when they were erect, but that the orientation of the observer had no significant effect on reaction time (F=1.07, df 1/15, p (is greater than) .30). These data strongly suggest that the orientation of the image of a face on the observer's retina, but not its orientation with respect to gravity, is important in identifying the expression on the face.

Cohen, M. M.↗

Evaluating the Use of Foundational Chemical Language Models in Multimodal Graph Fusion

Rapid and accurate prediction of the physicochemical properties of molecules given their structures remains a key challenge in cheminformatics. Machine learning approaches offer high-throughput options, but the optimality of inductive biases and data representations are up for debate. For example, BERT-based masked language models (MLMs) can be trained in a self-supervised way on hundreds of millions to billions of readily available SMILES strings. Another option is graph neural networks (GNNs), which can operate directly on molecular structures. Yet, generating accurate molecular geometry is computationally expensive, leading to a relative scarcity in data compared to SMILES strings. It is attractive to combine these two paradigms by pre-training an LM on a large corpus of SMILES strings and embedding these representation into a geometric graph neural network. Despite the promise of such an approach, and contrary to previous studies, we find mixed results with the combination of the LMs and GNNs on several molecule datasets. In particular, we found evidence for improvement on the FreeSolv and QM7 benchmarks, but degraded performance on the ESOL, LIPO and QM9 datasets compared to a GNN baseline.

Francel, Collin [University of Alabama]↗

Comparison of Machine Learning Approaches for Prediction of the Equivalent Alkane Carbon Number for Microemulsions Based on Molecular Properties

The chemical properties of oils are vital in the design of microemulsion systems. The hydrophilic–lipophilic difference equation used to predict microemulsions’ phase behavior expresses the oils’ physiochemical properties as the equivalent alkane carbon number (EACN). The experimental determination of EACN requires knowledge of the temperature dependence of the microemulsion system and the effects of different surfactant concentrations. Thus, the experimental determination is time-intensive and tedious, requiring days to months for proper separations. Furthermore, the experiments require high purity of chemicals because microemulsions are sensitive to impurities. Our work focuses on the quick and reliable predictions of the EACN with machine learning (ML) models. Due to the immaturity of ML chemical predictions, we compare three graph neural networks (GNNs) and a gradient-boosted tree algorithm, known as XGBoost. The GNNs use the molecular structures represented as simplified molecular-input line-entry system (SMILES) codes for the initial input, which allows us to assess whether geometry optimization is necessary for reliable results. The XGBoost model also begins with the SMILES representations of the molecules but uses molecular descriptors instead of geometry optimizations. As a result, the best model tested (crystal graph convolutional neural network with Merck molecular force field-94) has an error of 1.15 EACN units of the true EACN for unknown data with the errors skewed toward zero and an R² score of 0.9

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Changing windstorm characteristics over the US Northeast in a single model large ensemble

Abstract Extreme windstorms pose a significant hazard to infrastructure and public safety, particularly in the highly populated US Northeast (NE). However, the influence climate change and changing land use will have on these events remains unclear. A large ensemble generated using the Max-Planck Institute (MPI) Earth system model is used to generate projections of NE windstorms under different shared socioeconomic pathways (SSPs) and to attribute changes to projected land use land cover (LULC) change, externally forced changes and internal climate variability. To reduce the influence of coarse grid cell resolution and uncertainties in surface roughness lengths, windstorms are identified using simultaneous widespread exceedance of local 99th percentile 10 m wind speeds (U 99 ). Projected declines in forest cover in the NE and the resulting reductions in surface roughness length under SSP3-7.0 lead to projections of large increases in U 99 and derived windstorm intensity and scale. However, these projected changes in regional LULC under SSP3-7.0 are unprecedented in a historical context and may not be realistic. After corrections are applied to remove the influence of LULC on wind speeds, regionally averaged U 99 exhibit declines for most of the single model initial-condition large ensemble (SMILE) members which are broadly proportional to the radiative forcing and global air temperature increase in the SSPs, with a median value of −0.15 ms −1 °C −1 . While weak cyclones are projected to decline in frequency in the NE, intense cyclones and the resulting windstorms and indices of socioeconomic loss do not. Where present, significant trends in these loss indices are positive, and some MPI SMILE members generate future windstorms that are unprecedented in the historical period.

Coburn, Jacob (ORCID:0000000309538117)↗

Evaluation of historical precipitation interannual variability in CMIP6 over the United States

Interannual precipitation variability profoundly influences society via its effects on agriculture, water resources, infrastructure, and disaster risks. In this study, we use daily in situ precipitation observations from the global historical climatology network-daily (GHCN-D) to assess the ability of 21 Coupled Model Intercomparison Project Phase 6 (CMIP6) models, including the 50-member fifth-generation Canadian Earth System Model single model initial-condition large ensemble (CanESM5_SMILE), to realistically simulate historical interannual precipitation variability trends within 17 regions of the contiguous United States (CONUS). We assess how accurately the CMIP6 simulations align with observational data across annual, summer, and winter periods, focusing on four key hydrometeorological metrics, including interannual precipitation variability, relative interannual precipitation variability (coefficient of variation), annual mean precipitation, and annual wet day frequency. Our findings reveal that CMIP6 ensemble members generally reproduce the spatial patterns of observed trends in annual mean precipitation. In most regions, models agree well with the signs of observed changes in annual mean precipitation, though discrepancies in trend magnitude are evident. Further, observed trends in winter mean precipitation broadly exhibit a spatial pattern similar to that of the observed annual mean. However, analysis of the CanESM5_SMILE shows that trends in precipitation variability may primarily be the result of model-simulated internal variability, suggesting caution in interpreting multi-model single-realization ensemble results. Challenges in accurately simulating interannual precipitation variability underscore the need for ongoing model refinement and validation to enhance climate projections, especially in regions vulnerable to extreme precipitation events.

54 ENVIRONMENTAL SCIENCES↗

A Variational Autoencoder Model Toward Molecular Structure Representation Learning of Fuels

Here, in this work, a Variational Autoencoder (VAE)-based data-driven modeling framework is developed with the overarching goal of enabling fuel design. The VAE model is trained on a large dataset with several chemical species to learn a compressed latent space molecular representation. Chemical structure in the form of Simplified Molecular Input Line Entry System (SMILES) string is fed as input, encoded into the VAE latent space, and decoded back to the SMILES string using Long Short-Term Memory (LSTM) networks. Complexities of the VAE training loss function are thoroughly examined by varying the weightage (beta (𝜷) parameter) of the latent space regularization term, thereby assessing the balance between reconstruction accuracy and validity, and focusing on both accurate molecular structure reconstruction and latent space consistency. Two different strategies for 𝜷 variation are evaluated: linear annealing and cyclic annealing. In addition, the impact of total correlation adjustment and hierarchical priors is also studied with regard to the balance between reconstruction fidelity and latent space regularization, and potential issues such as posterior collapse, over-regularization, and poor disentanglement of latent variables. Overall, the best performance of the model is achieved with hierarchical priors and incrementally increasing 𝜷 from 0 to a threshold value of 0.25 over 75 epochs. The generative VAE model can be readily coupled with Quantitative Structure–Property Relationship (QSPR) analysis to develop an integrated end-to-end framework for fuel-property prediction and molecular design of novel promising fuels.

fuel design↗

Equivariant Graph Attention Network - 3D Conformers & Feature Fusion

EGAN-3F (Equivariant Graph Attention Network - 3D Conformers & Feature Fusion) presents an innovative approach for predicting binding affinity between small molecules and protein targets, a fundamental task in drug discovery. Traditional structure-based methods often depend on protein-ligand complex structures obtained from crystallography or molecular docking. In contrast, ligand-only machine learning models using 1D or 2D representations such as SMILES have been developed to predict binding affinity without structural information about the target; however, their accuracy is often limited due to the lack of 3D ligand information. EGAN-3F addresses this limitation by integrating spatially aware graph learning with traditional descriptor-based features. We systematically investigate how combining 2D and 3D molecular representations enhances binding affinity prediction from SMILES strings. This approach underscores the importance of modeling conformational diversity and incorporating chemically meaningful descriptors to improve predictive accuracy. The key innovation of EGAN-3F lies in its ability to achieve robust ligand-based binding affinity predictions without requiring protein-ligand complex structures, effectively bridging the gap between purely structural and ligand-only modeling paradigms.

Shim, Heesung [Lawrence Livermore National Laborat↗

HT Model Dataset

This website contains the dataset that was used for writing the manuscript "HT Model: Using the Molecular Transformer for predicting hydrotreating reactions" (PNNL-SA-186589) The dataset includes a collection of hydrotreating reactions compiled from 41 peer-reviewed literature sources. These sources contain experimental data related to hydrotreating reactions. These reactions involve the reaction of chemical compounds with hydrogen gas in the presence of a catalyst to remove heteroatoms or to convert specific functional groups. The dataset contains reactions both with and without reaction conditions. Reaction conditions refer to the specific parameters under which the reaction takes place, such as temperature and pressure. For each reaction, the dataset includes both SMILES and SELFIES representations. SMILES (Simplified Molecular Input Line Entry System) and SELFIES (SELF-referencIng Embedded Strings) are two popular notations used to represent chemical structures in a compact and standardized format. The dataset was created with the aim of training a predictive model, specifically using the Molecular Transformer architecture.

bioprocessing, hydrotreating, deep learning algori↗

Estimating the Subsolar Magnetopause Position from Soft X-Ray Images Using a Low-Pass Image Filter

The Lunar Environment heliospheric X-ray Imager (LEXI) and Solar wind Magnetosphere Ionosphere Link Explorer (SMILE) missions will image the Earth’s dayside magnetopause and cusps in soft X-rays after their respective launches in the near future, to specify global magnetic reconnection modes for varying solar wind conditions. To support the success of these scientific missions, it is critical to develop techniques that extract the magnetopause locations from the observed soft X-ray images. In this research, we introduce a new geometric equation that calculates the subsolar magnetopause position ( R s ) from a satellite position, the look direction of the instrument, and the angle at which the X-ray emission is maximized. Two assumptions are used in this method: (1) The look direction where soft X-ray emissions are maximized lies tangent to the magnetopause, and (2) the magnetopause surface near the subsolar point is almost spherical and thus R s is nearly equal to the radius of the magnetopause curvature. We create synthetic soft X-ray images by using the Open Geospace General Circulation Model (OpenGGCM) global magnetohydrodynamic model, the galactic background, the instrument point spread function, and Poisson noise. We then apply the fast Fourier transform and Gaussian low-pass filters to the synthetic images to remove noise and obtain accurate look angles for the soft X-ray peaks. From the filtered images, we calculate R 2 and its accuracy for different LEXI locations, look directions, and solar wind densities by using the OpenGGCM subsolar magnetopause location as ground truth. Our method estimates R s with an accuracy of <0.3 R E when the solar wind density exceeds >10 cm -3 . The accuracy improves for greater solar wind densities and during southward interplanetary magnetic fields. The method captures the magnetopause motion during southward interplanetary magnetic field turnings. Consequently, the technique will enable quantitative analysis of the magnetopause motion and help reveal the dayside reconnection modes for dynamic solar wind conditions. This technique will support the LEXI and SMILE missions in achieving their scientific objectives.

Hyangpyo Kim↗

Bloom filters for molecules

Abstract Ultra-large chemical libraries are reaching 10s to 100s of billions of molecules. A challenge for these libraries is to efficiently check if a proposed molecule is present. Here we propose and study Bloom filters for testing if a molecule is present in a set using either string or fingerprint representations. Bloom filters are small enough to hold billions of molecules in just a few GB of memory and check membership in sub milliseconds. We found string representations can have a false positive rate below 1% and require significantly less storage than using fingerprints. Canonical SMILES with Bloom filters with the simple FNV (Fowler-Noll-Voll) hashing function provide fast and accurate membership tests with small memory requirements. We provide a general implementation and specific filters for detecting if a molecule is purchasable, patented, or a natural product according to existing databases at https://github.com/whitead/molbloom .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Creation of Polymer Datasets with Targeted Backbones for Screening of High-Performance Membranes for Gas Separation

A simple approach was developed to computationally construct a polymer dataset by combining simplified molecular-input line-entry system (SMILES) strings of a targeted polymer backbone and a variety of molecular fragments. This method was used to create 14 polymer datasets by combining seven polymer backbones and molecules from two large molecular datasets (MOSES and QM9). Polymer backbones that were studied include four polydimethylsiloxane (PDMS) based backbones, poly(ethylene oxide) (PEO), poly(allyl glycidyl ether) (PAGE), and polyphosphazene (PPZ). The generated polymer datasets can be used for various cheminformatics tasks, including high-throughput screening for gas permeability and selectivity. This study utilized machine learning (ML) models to screen the polymers for CO2/CH4 and CO2/N2 gas separation using membranes. Several polymers of interest were identified. Here the results highlight that employing an ML model fitted to polymer selectivities leads to higher accuracy in predicting polymer selectivity compared to using the ratio of predicted permeabilities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗