Search NASA⌕ Search

Engineering topics

Glenski, Maria F.

Publications and source records attributed to Glenski, Maria F..

Foundation Models of Scientific Knowledge for Chemistry: Opportunities, Challenges and Lessons Learned

Foundation models pre-trained on large corpora demonstrate significant gains across many natural language processing tasks and domains e.g., law, healthcare, education, etc. However, only limited efforts have investigated the opportunities and limitations of applying these powerful models to science and security applications. In this work we develop foundation models of scientific knowledge for chemistry to augment scientists with the advanced ability to perceive and reason at scale previously unimagined. Specifically, we build large-scale (1.47B parameter) general-purpose models for chemistry that can be effectively used to perform a wide range of in-domain and out-of-domain tasks. Evaluating these models in a zero-shot setting, we analyze the effect of model and data scaling, knowledge depth, and temporality on model performance in context of model training efficiency. Our novel findings demonstrate that (1) model size significantly contributes to the task performance when evaluated in a zero-shot setting; (2) data quality (aka diversity) affects model performance more than data quantity; (3) similarly, unlike previous work (Luu et al., 2021) temporal order of the documents in the corpus boosts model performance only for specific tasks, e.g., SciQ; and (4) models pre-trained from scratch perform better on in-domain tasks than those tuned from general-purpose models like Open AI’s GPT-2.

Foundation Models, Chemistry↗

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information

Detecting and anticipating global proliferation expertise and capability evolution from unstructured, noisy, and incomplete public data streams is a highly desired, but extremely challenging task. Here, in this article, we present our pioneering data-driven approach to support the non-proliferation mission to detect and explain the evolution of proliferation expertise and capability development globally from terabytes of publicly available information (PAI), focusing on our knowledge extraction pipeline and descriptive analytics. We first discuss how we fuse nine open-source data streams, including multilingual data, to convert 4 TB of unstructured data to structured knowledge and encode dynamically evolving proliferation expertise representations—content and context graphs. For this, we rely on natural language processing (NLP) and deep learning (DL) models to perform information extraction, topic modeling, and distributed text representation (aka embedding) learning. We then present interactive, usable, and explainable descriptive analytics to refine domain knowledge and present it in a human-understandable form. Finally, we introduce future work avenues that will leverage our dynamic knowledge representations and descriptive analytics to enable predictive and prescriptive inferences to achieve real-time domain understanding and contextual reasoning about global proliferation expertise and capability evolution.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

VAINE: Visualization and AI for Natural Experiments

Natural experiments are observational studies where the assignment of treatment conditions to different populations occur by chance ``in the wild''. Researchers from fields such as economics, healthcare, and the social sciences leverage natural experiments to conduct hypothesis testing and causal effect estimation for treatment and outcome variables that would otherwise be costly, infeasible, or unethical. In this paper, we introduce VAINE (Visualization and AI for Natural Experiments), a visual analytics tool for identifying and understanding natural experiments from observational data. We then demonstrate how VAINE can be used to validate causal relationships, estimate average treatment effects, and identify statistical phenomena such as Simpson’s paradox through two use cases.

Guo, Grace↗

Explaining and predicting human behavior and social dynamics in simulated virtual worlds: reproducibility, generalizability, and robustness of causal discovery methods

Ground Truth program was designed to evaluate social science modeling approaches using simulation test beds with ground truth intentionally and systematically embedded to understand and model complex Human Domain systems and their dynamics Lazer et al. (Science 369:1060–1062, 2020). Our multidisciplinary team of data scientists, statisticians, experts in Artificial Intelligence (AI) and visual analytics had a unique role on the program to investigate accuracy, reproducibility, generalizability, and robustness of the state-of-the-art (SOTA) causal structure learning approaches applied to fully observed and sampled simulated data across virtual worlds. In addition, we analyzed the feasibility of using machine learning models to predict future social behavior with and without causal knowledge explicitly embedded. In this paper, we first present our causal modeling approach to discover the causal structure of four virtual worlds produced by the simulation teams—Urban Life, Financial Governance, Disaster and Geopolitical Conflict. Our approach adapts the state-of-the-art causal discovery (including ensemble models), machine learning, data analytics, and visualization techniques to allow a human-machine team to reverse-engineer the true causal relations from sampled and fully observed data. We next present our reproducibility analysis of two research methods team’s performance using a range of causal discovery models applied to both sampled and fully observed data, and analyze their effectiveness and limitations. We further investigate the generalizability and robustness to sampling of the SOTA causal discovery approaches on additional simulated datasets with known ground truth. Our results reveal the limitations of existing causal modeling approaches when applied to large-scale, noisy, high-dimensional data with unobserved variables and unknown relationships between them. We show that the SOTA causal models explored in our experiments are not designed to take advantage from vasts amounts of data and have difficulty recovering ground truth when latent confounders are present; they do not generalize well across simulation scenarios and are not robust to sampling; they are vulnerable to data and modeling assumptions, and therefore, the results are hard to reproduce. Finally, when we outline lessons learned and provide recommendations to improve models for causal discovery and prediction of human social behavior from observational data, we highlight the importance of learning data to knowledge representations or transformations to improve causal discovery and describe the benefit of causal feature selection for predictive and prescriptive modeling.

97 MATHEMATICS AND COMPUTING↗

Improving Synonym Recommendation Using Sentence Context

Traditional synonym recommendations often include ill-suited suggestions for writer's specific contexts. We propose a simple approach for contextual synonym recommendation by combining existing human-curated thesauri, e.g. WordNet, with pre-trained language models. We evaluate our technique by curating a set of word-sentence pairs balanced across corpora and parts of speech, then annotating each word-sentence pair with the contextually appropriate set of synonyms. We found that basic language model approaches have higher precision. Approaches leveraging sentence context have higher recall. Overall, the latter contextual approach had the highest F-score.

Glenski, Maria F.↗

November Mega-AI Newsletter

Mega AI, an internal investment at Pacific Northwest National Laboratory, aims to develop next-generation artificial intelligence (AI) capabilities unique to the Department of Energy national lab complex to address research gaps in large-scale multimodal representation learning, multitask inferences, and the need for increased generalizability, rapid adaptivity, and usability of AI technologies. In this newsletter, we highlight recent developments in the research community on next-gen AI technologies focusing on massive-scale model development, deployment and evaluation, data and code availability, model interactions, and new features and capabilities that are relevant to Mega AI’s goals and science and security applications.

Volkova, Svitlana↗

CrossCheck: Rapid, Reproducible, and Interpretable Model Evaluation

Evaluation beyond aggregate performance metrics, e.g. F1-score, is crucial to both establish an appropriate level of trust in machine learning models and identify future model improvements. In this paper we demonstrate CrossCheck, an interactive visualization tool for rapid crossmodel comparison and reproducible error analysis. We describe the tool and discuss design and implementation details. We then present three use cases (named entity recognition, reading comprehension, and clickbait detection) that show the benefits of using the tool for model evaluation. CrossCheck allows data scientists to make informed decisions to choose between multiple models, identify when the models are correct and for which examples, investigate whether the models are making the same mistakes as humans, evaluate models’ generalizability and highlight models’ limitations, strengths and weaknesses. Furthermore, CrossCheck is implemented as a Jupyter widget, which allows rapid and convenient integration into data scientists’ model development workflows.

Arendt, Dustin L.↗

Evaluating Deception Detection Model Robustness To Linguistic Variation

With the increasing use of automated, machine learning-driven tools and the downstream impact that algorithmic judgements can have, it is critical to develop models that are robust to evolving or manipulated inputs. Evaluating the reliability of multimodal models across linguistic variations to understand model susceptibility to intentional linguistic adversarial attacks as well as natural linguistic variations is essential in this pursuit. We present extensive analysis of model robustness and susceptibility to linguistic variations in the setting of deceptive news detection, a difficult classification task that is an increasingly important problem to solve with the impact of misinformation spread online. We evaluate the effectiveness of incorporating adversarial defense strategies and measure model susceptibility to state-of-the-art adversarial attacks using two types of linguistic attacks — character and word perturbations. We consider two multiclass prediction tasks — a 3-way classification of tweets as trustworthy, propaganda, or disinformation; and a 4-way classification as clickbait, hoax, satire, or conspiracy — and compare the performance of three embeddings that have been state-of-the-art for several NLP tasks — GloVe, ELMo, and BERT — to highlight consistent trends in susceptibility, high confidence misclassifications, and high impact failures. We find that character or mixed ensemble models are the most effective defense mechanisms and that character perturbations are a more effective attack than word perturbations for deception classification.

adversarial evaluation↗

Leveraging Community and Author Context to Explain the Performance and Bias of Text-Based Deception Detection Models

Deceptive news posts shared in online communities can be detected with NLP models, and much recent research has focused on the development of such models. In this work, we use characteristics of online communities and authors --- the context of how and where content is posted --- to explain the performance of a neural network deception detection model and identify sub-populations who are disproportionately affected by model accuracy or failure. We examine who is posting the content, and where the content is posted to. We find that while author characteristics are better predictors of deceptive content than community characteristics, both characteristics are strongly correlated with model performance. Traditional performance metrics such as F1 score may fail to capture poor model performance on isolated sub-populations such as specific authors, and as such, more nuanced evaluation of deception detection models is critical.

machine learning (ML), machine learning explanatio↗

Machine Intelligence to Detect, Characterise, and Defend against Influence Operations in the Information Environment

Social media has enabled a new era of manipulation in the information and cognitive domains. Deceptive content—misleading, falsified, and fabricated—is routinely created and spread in the modern social media environment with the intent to create confusion and widen political and social divides, and exploit the societal conflict exacerbated by these divides in the real-world (aka physical domain). Such disinformation campaigns demonstrate a threat to the integrity of economic, political, cultural, public health, and national security institutions around the world. In this work we overview our artificial intelligence (AI) capabilities to detect, describe, and defend against information operations on Twitter as an example social platform to understand the influence of misleading and falsified content diffusion and better enable those charged with defending against such manipulation to enable responsive parties to counter it. We first present novel linguistically-informed deep learning (DL) models for misinformation and disinformation detection, and present an in-depth linguistic analysis of psycho-linguistic markers across broad deception categories. We then demonstrate how our models perform in the multilingual and multimodal setting and categorize falsified and misleading content based on the intent to deceive. We also provide a large-scale analysis to describe user behavior and spread patterns while engaging with deceptive content and report novel findings about the immediate diffusion of deceptive content by characterizing the vulnerable sub-populations and their demographics, and explicitly measuring speed and scale of deception spread to uncover who shares deceptive content, how quickly, how much, and how evenly. In addition, we measure audience reactions to misinformation and disinformation at scale, distinguishing the reactions of users identified as bots versus humans. Finally, we take advantage of deep translation and generation models to create unique solutions for real-time defense against digital deception and discuss how to apply causal inference to prescribe and intervene into strategic communications jointly across information, cognitive, and physical domains.

artificial intelligence, deep learning, neural lan↗

Report on Next-Gen AI for Proliferation Detection Workshop: Domain-Aware Methods

The emergence of artificial intelligence (AI) and machine learning (ML) in the modern world has impacted nearly every application imaginable. This includes nuclear proliferation detection, which offers the potential to improve existing capabilities as well as create new ones. Proliferation detection seeks to detect and characterize attempts by state and non-state actors to acquire nuclear weapons or associated technology, materials, or knowledge. Such a mission is vitally important for global stability and security but is notoriously difficult. By leveraging advances in AI, exciting opportunities exist to enhance the proliferation detection regime. The Data Science and AI portfolio within the National Nuclear Security Administration’s Office of Defense Nuclear Nonproliferation Research and Development (DNN R&D) seeks to leverage the capabilities of the Department of Energy’s (DOE’s) national laboratories and other partners to develop AI systems that can accomplish otherwise impossible tasks in support of proliferation detection. As part of its efforts, the portfolio has created a series of workshops on Next-Gen AI for Proliferation Detection to help define the requirements for suitable AI systems, share successful research and best practices, and foster connection and understanding between the relevant parties including researchers and end-users. Each workshop in the series focuses on a specific and critical aspect of AI to enable it to accomplish proliferation detection objectives. The first workshop focused on explainability techniques; the second workshop and the topic of this report, covers methods for incorporating domain awareness into AI. The Next-Gen AI for Proliferation Detection Workshop: Domain-Aware Methods took place virtually over two days in February 2021 and included four keynote presentations, 22 technical presentations, and a concluding panel. The presentations, discussions, and workshop findings are summarized in this report.

97 MATHEMATICS AND COMPUTING↗