Search NASA⌕ Search

SEARCH · Search NASA

Results for “information entropy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A cosine-based correlation information entropy approach for building automatic fault detection baseline construction

Building automatic fault detection and diagnosis (AFDD) technologies have shown great potential for energy savings. To enable AFDD, a baseline depicting the normal operation mode is needed to detect whether the building operation deviates from normality. Existing research using physics-based knowledge and models for AFDD has mainly taken a trial-and-error approach to determine if a given baseline is sufficient via empirical experiments. A mechanism to support decisions such as how many samples and what samples should be included in the baseline is currently lacking. In this study, a data-driven method for AFDD baseline construction based on information entropy is developed. The entropy is derived based on cosine similarity among typical building automation system measurements in conjunction with outdoor weather information. The performance of the proposed method is evaluated using real building data. Evaluation results indicate that the fault detection strategy adopting the proposed method has similar or better accuracy in detecting faults compared to the same fault detection strategy using the baseline construction method from the literature. Additionally, the use of entropy enables the proposed method to automatically construct and assess the baseline consisting of information-rich samples.

42 ENGINEERING↗

Method of information entropy for convergence assessment of molecular dynamics simulations

The lack of a reliable method to evaluate the convergence of molecular dynamics simulations has contributed to discrepancies in different areas of molecular dynamics. Here, the method of information entropy is introduced to molecular dynamics for stationarity assessment. The Shannon information entropy formalism is used to monitor the convergence of the atom motion to a steady state in a continuous spatial domain and is also used to assess the stationarity of calculated multidimensional fields such as the temperature field in a discrete spatial domain. It is demonstrated in this work that monitoring the information entropy of the atom position matrix provides a clear indicator of reaching steady state in radiation damage simulations, non-equilibrium molecular dynamics thermal conductivity computations, and simulations of Poiseuille and Couette flow in nanochannels. A main advantage of the present technique is that it is non-local and relies on fundamental quantities available in all molecular dynamics simulations. Unlike monitoring average temperature, the technique is applicable to simulations that conserve total energy such as reverse non-equilibrium molecular dynamics thermal conductivity computations and to simulations where energy dissipates through a boundary as in radiation damage simulations. The method is applied to simulations of iron using the Tersoff/ZBL splined potential, silicon using the Stillinger–Weber potential, and to Lennard–Jones fluid. Its applicability to both solids and fluids shows that the technique has potential for generalization to other areas in molecular dynamics.

74 ATOMIC AND MOLECULAR PHYSICS↗

Information Entropy as Quantifier of Potential Predictability in the Tropical Indo-Pacific Basin

Global warming is posed to modify the modes of variability that control much of the climate predictability at seasonal to interannual scales. The quantification of changes in climate predictability over any given amount of time, however, remains challenging. Here we build upon recent advances in non-linear dynamical systems theory and introduce the climate community to an information entropy quantifier based on recurrence. The entropy, or complexity of a system is associated with microstates that recur over time in the time-series that define the system, and therefore to its predictability potential. A computationally fast method to evaluate the entropy is applied to the investigation of the information entropy of sea surface temperature in the tropical Pacific and Indian Oceans, focusing on boreal fall. In this season the predictability of the basins is controlled by two regularly varying non-linear oscillations, the El Niño-Southern Oscillation and the Indian Ocean Dipole. We compute and compare the entropy in simulations from the CMIP5 catalog from the historical period and RCP8.5 scenario, and in reanalysis datasets. Discrepancies are found between the models and the reanalysis, and no robust changes in predictability can be identified in future projections. The Indian Ocean and the equatorial Pacific emerge as troublesome areas where the modeled entropy differs the most from that of the reanalysis in many models. A brief investigation of the source of the bias points to a poor representation of the ocean mean state and basins' connectivity at the Indonesian Throughflow.

54 ENVIRONMENTAL SCIENCES↗

Influence of initial conditions on data-driven model identification and information entropy for ideal mhd problems

Data-driven methods of model identification are able to discern governing dynamics of a system from data. Such methods are well suited to help us learn about systems with unpredictable evolution or systems with ambiguous governing dynamics given our current understanding. Many plasma problems of interest fall into these categories as there are a wide range of models that exist, however each model is only useful in a certain regime and often limited by computational complexity. To ensure data-driven methods align with theory, they must be consistent and predictable when acting on data whose governing dynamics are known. Weak Sparse Identification of Nonlinear Dynamics (WSINDy) is a recently developed data-driven method that has shown promise in learning governing dynamics from data with high noise levels [1]. This work examines how WSINDy acts on ideal MHD test problems as the initial conditions are varied and specifies limiting requirements for successful equation identification. Furthermore, it is hard to recover the governing dynamics from data that emphasize a single dominant behavior. In these low information cases, Shannon information entropy is able to pick up on the redundancies in the data that affect recoverability.

97 MATHEMATICS AND COMPUTING↗

EE-SMOTE: An oversampling method in conjunction with information entropy for imbalanced learning

Imbalanced learning attracts great attention in various research fields. Existing literature-reported methodologies in imbalanced learning have shown drawbacks including over-generation or noisy/wrong samples generations. This paper presents EE-SMOTE, an oversampling technique based on information entropy, to support the imbalance classifications. Specifically, we propose a metric, Eigen-Entropy (EE), to identify homogenous samples from minority classes for oversampling technique, specifically, SMOTE to reach data balances for classification. Experiments on public dataset and real-world datasets demonstrate the efficacy and effectiveness of the proposed EE-SMOTE in imbalanced learning.

Huang, Jiajing↗

Information-entropy-driven generation of material-agnostic datasets for machine-learning interatomic potentials

In contrast to their empirical counterparts, machine-learning interatomic potentials (MLIAPs) promise to deliver near-quantum accuracy over broad regions of configuration space. However, due to their generic functional forms and extreme flexibility, they can catastrophically fail to capture the properties of novel, out-of-sample configurations, making the quality of the training set a determining factor, especially when investigating materials under extreme conditions. We propose a novel automated dataset generation method based on the maximization of the information entropy of the feature distribution, aiming at an extremely broad coverage of the configuration space in a way that is agnostic to the properties of specific target materials. The ability of the dataset to capture unique material properties is demonstrated on a range of unary materials, including elements with the FCC (Al), BCC (W), HCP (Be, Re and Os), graphite (C), and trigonal (Sb, Te) ground states. MLIAPs trained to this dataset are shown to be accurate over a range of application-relevant metrics, as well as extremely robust over very broad swaths of configurations space, even without dataset fine-tuning or hyper-parameter optimization, making the approach extremely attractive to rapidly and autonomously develop general-purpose MLIAPs suitable for simulations in extreme conditions.

36 MATERIALS SCIENCE↗

Nonequilibrium information entropy approach to ternary fission of actinides

Ternary fission of actinides probes the state of the nucleus at scission. Light clusters are produced in space and time very close to the scission point. Within the nonequilibrium statistical operator method, a generalized Gibbs distribution is constructed from the information given by the observed yields of isotopes. Using this relevant statistical operator, yields are calculated taking excited states and continuum correlations into account, in accordance with the virial expansion of the equation of state. Furthermore, clusters with mass number A≤10 are well described using the nonequilibrium generalizations of temperature and chemical potentials. Improving the virial expansion, in-medium effects may become of importance in determining the contribution of weakly bound states and continuum correlations to the intrinsic partition function. Yields of larger clusters, which fail to reach this quasiequilibrium form of the relevant distribution, are described by nucleation kinetics, and a saddle-to-scission relaxation time of about 7000 fm/c is inferred. Light-charged particle emission, described by reaction kinetics and virial expansions, may therefore be regarded as a very important tool to probe the nonequilibrium time evolution of actinide nuclei during fission.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Nowcasting Earthquakes With Stochastic Simulations: Information Entropy of Earthquake Catalogs

Earthquake nowcasting has been proposed as a means of tracking the change in large earthquake potential in a seismically active area. The method was developed using observable seismic data, in which probabilities of future large earthquakes can be computed using Receiver Operating Characteristic methods. Furthermore, analysis of the Shannon information content of the earthquake catalogs has been used to show that there is information contained in the catalogs, and that it can vary in time. So an important question remains, where does the information originate? In this paper, we examine this question using stochastic simulations of earthquake catalogs. Our catalog simulations are computed using an Earthquake Rescaled Aftershock Seismicity (“ERAS”) stochastic model. This model is similar in many ways to other stochastic seismicity simulations, but has the advantage that the model has only 2 free parameters to be set, one for the aftershock (Omori-Utsu) time decay, and one for the aftershock spatial migration away from the epicenter. Generating a simulation catalog and fitting the two parameters to the observed catalog such as California takes only a few minutes of wall clock time. While clustering can arise from random, Poisson statistics, we show that significant information in the simulation catalogs arises from the “non-Poisson” power-law aftershock clustering, implying that the practice of de-clustering observed catalogs may remove information that would otherwise be useful in forecasting and nowcasting. We also show that the nowcasting method provides similar results with the ERAS model as it does with observed seismicity.

58 GEOSCIENCES↗

Quantifying Information without Entropy: Identifying Intermittent Disturbances in Dynamical Systems

A system’s response to disturbances in an internal or external driving signal can be characterized as performing an implicit computation, where the dynamics of the system are a manifestation of its new state holding some memory about those disturbances. Identifying small disturbances in the response signal requires detailed information about the dynamics of the inputs, which can be challenging. This paper presents a new method called the Information Impulse Function (IIF) for detecting and time-localizing small disturbances in system response data. The novelty of IIF is its ability to measure relative information content without using Boltzmann’s equation by modeling signal transmission as a series of dissipative steps. Since a detailed expression of the informational structure in the signal is achieved with IIF, it is ideal for detecting disturbances in the response signal, i.e., the system dynamics. Those findings are based on numerical studies of the topological structure of the dynamics of a nonlinear system due to perturbated driving signals. The IIF is compared to both the Permutation entropy and Shannon entropy to demonstrate its entropy-like relationship with system state and its degree of sensitivity to perturbations in a driving signal.

42 ENGINEERING↗

Entropy and Boundary Based Adversarial Learning for Large Scale Unsupervised Domain Adaptation

Supervised semantic segmentation methods provide state-of-the-art performance, but their performance is limited by the amount of quality labeled data they need for training. Scarcity of labeled data and non-transferablity of models, due to cross-domain discrepancy makes it a bigger challenge for remote sensing imagery analysis. In this work, we approach this problem through adversarial learning, driven by entropy and boundary of region-of-interest for unsupervised domain adaptation. This concept helps with better boundary prediction and encourages target domain entropy maps (probability/uncertainty maps) to be similar to source domains. In particular, we showed that deriving informative entropy through the adversarial learning is essential to enable the adaptation. We used a large scale cross country building extraction dataset to validate the framework. The experimental results show the usefulness of considering boundary and entropy driven adversarial learning for adaptation.

Makkar, Nikhil↗

Non-equilibrium entropy production and information dissipation in a non-Markovian quantum dot

This study measures trajectory-level entropy production and information dissipation in a driven, non-Markovian quantum dot using time-resolved optical dynamics and machine-learning-based analysis. Although not a 2D-material system, it is relevant because it demonstrates quantitative extraction of nonequilibrium dynamics from nanoscale optical fluctuations, which is conceptually connected to the proposed studies of transient charge and spin dynamics at interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Model-free estimation of completeness, uncertainties, and outliers in atomistic machine learning using information theory

Abstract An accurate description of information is relevant for a range of problems in atomistic machine learning (ML), such as crafting training sets, performing uncertainty quantification (UQ), or extracting physical insights from large datasets. However, atomistic ML often relies on unsupervised learning or model predictions to analyze information contents from simulation or training data. Here, we introduce a theoretical framework that provides a rigorous, model-free tool to quantify information contents in atomistic simulations. We demonstrate that the information entropy of a distribution of atom-centered environments explains known heuristics in ML potential developments, from training set sizes to dataset optimality. Using this tool, we propose a model-free UQ method that reliably predicts epistemic uncertainty and detects out-of-distribution samples, including rare events in systems such as nucleation. This method provides a general tool for data-driven atomistic modeling and combines efforts in ML, simulations, and physical explainability.

36 MATERIALS SCIENCE↗

Protein Conformational States—A First Principles Bayesian Method

Automated identification of protein conformational states from simulation of an ensemble of structures is a hard problem because it requires teaching a computer to recognize shapes. We adapt the naïve Bayes classifier from the machine learning community for use on atom-to-atom pairwise contacts. The result is an unsupervised learning algorithm that samples a ‘distribution’ over potential classification schemes. We apply the classifier to a series of test structures and one real protein, showing that it identifies the conformational transition with >95% accuracy in most cases. A nontrivial feature of our adaptation is a new connection to information entropy that allows us to vary the level of structural detail without spoiling the categorization. This is confirmed by comparing results as the number of atoms and time-samples are varied over 1.5 orders of magnitude. Further, the method’s derivation from Bayesian analysis on the set of inter-atomic contacts makes it easy to understand and extend to more complex cases.

97 MATHEMATICS AND COMPUTING↗

Evolution of oxygen and stratification and their relationship in the North Pacific Ocean in CMIP6 Earth system models

Abstract. This study examines the linkages between the upper-ocean (0–200 m) oxygen (O2) content and stratification in the North Pacific Ocean using four Earth system models (ESMs), an ocean hindcast simulation, and an ocean reanalysis. The trends and variability in oceanic O2 content are driven by the imbalance between physical supply and biological demand. Physical supply is primarily controlled by ocean ventilation, which is responsible for the transport of O2-rich surface waters to the subsurface. Isopycnic potential vorticity (IPV), a quasi-conservative tracer proportional to density stratification that can be evaluated from temperature and salinity measurements, is used herein as a dynamical proxy for ocean ventilation. The predictability potential of the IPV field is evaluated through its information entropy. The results highlight a strong O2–IPV connection and somewhat higher (as compared to the rest of the basin) predictability potential for IPV across the tropical Pacific, where the El Niño–Southern Oscillation occurs. This pattern of higher predictability and strong anticorrelation between O2 and stratification is robust across multiple models and datasets. In contrast, IPV at mid-latitudes has low predictability potential and its center of action differs from that of O2. In addition, the locations of extreme events or hotspots may or may not differ between the two fields, with a strong model dependency, which persists in future projections. On the one hand, these results suggest that it may be possible to monitor ocean O2 in the tropical Pacific based on a few observational sites co-located with the more abundant IPV measurements; on the other, they lead us to question the robustness of the IPV–O2 relationship in the extratropics. The proposed framework helps to characterize and interpret O2 variability in relation to physical variability and may be especially useful in the analysis of new observation-based data products derived from the BGC-Argo float array in combination with the traditional but far more abundant Argo data.

Novi, Lyuba↗

ET-AL: Entropy-targeted active learning for bias mitigation in materials data

Growing materials data and data-driven informatics drastically promote the discovery and design of materials. While there are significant advancements in data-driven models, the quality of data resources is less studied despite its huge impact on model performance. In this work, we focus on data bias arising from uneven coverage of materials families in existing knowledge. Observing different diversities among crystal systems in common materials databases, we propose an information entropy-based metric for measuring this bias. To mitigate the bias, we develop an entropy-targeted active learning (ET-AL) framework, which guides the acquisition of new data to improve the diversity of underrepresented crystal systems. We demonstrate the capability of ET-AL for bias mitigation and the resulting improvement in downstream machine learning models. This approach is broadly applicable to data-driven materials discovery, including autonomous data acquisition and dataset trimming to reduce bias, as well as data-driven informatics in other scientific domains.

36 MATERIALS SCIENCE↗