Search NASA⌕ Search

SEARCH · Search NASA

Results for “domain expertise”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency Analysis

Checkpoint/Restart (C/R) has been widely deployed in numerous HPC systems, Clouds, and industrial data centers, which are typically operated by system engineers. Nevertheless, there is no existing approach that helps system engineers without domain expertise and domain scientists without system fault tolerance knowledge identify those critical variables accounted for correct application execution restoration in a failure for C/R. To address this problem, we propose an analytical model and a tool (AutoCheck) that can automatically identify critical variables to checkpoint for C/R. AutoCheck relies on first, analytically tracking and optimizing data dependency between variables and other application execution state, and second, a set of heuristics that identify critical variables for checkpointing from the refined data dependency graph (DDG). AutoCheck allows programmers to pinpoint critical variables to checkpoint quickly within a few minutes. We evaluate AutoCheck on 13 representative HPC benchmarks, demonstrating that AutoCheck can efficiently identify correct critical variables to checkpoint.

HPC↗

XRF-ROI Finder: Machine Learning to Guide Region-of-Interest Scanning for X-ray Fluorescence Microscopy

The ROI-finder software is being developed for use by several Microscopy Group beamlines at Argonne National Laboratory, including 2-ID microprobes and 9-ID-B Bionanoprobe which use multi-scale scanning fluorescence microscopy to acquire elemental maps (multi-modal image data). Microscopy experiments require scan of samples at a coarse resolution followed by ROI identification using feature detection based on domain expertise. Finer resolution scans are then conducted based on identified ROI. The decision-making process based on domain expertise will be difficult to perform for faster data rates and much larger sampling volumes anticipated after APS-U necessitating the need for the ROI-finder software. The ROI- finder detects regions of interest through a continuous learning process, starting with a unsupervised representation learning and improving its recommendations through supervised learning and an interactive tool for user annotation. The scope of ongoing development efforts includes the integration of image registration module to correlate optical and X-ray images, extraction of feature morphology as well as elemental signatures in the image space and incorporation of beamtime streaming data by the scanning probe via EPICS.

CHOWDHURY, M. ARSHAD ZAHANGIR↗

An intelligent simulation training system

The Department of Industrial Engineering at the University of Central Florida, Embry-Riddle Aeronautical University and General Electric (SCSD) have been funded by the State of Florida to build an Intelligent Simulation Training System. The objective was and is to make the system generic except for the domain expertise. Researchers accomplished this objective in their prototype. The system is modularized and therefore it is easy to make any corrections, expansions or adaptations. The funding by the state of Florida has exceeded $3 million over the past three years and through the 1990 fiscal year. UCF has expended in excess of 15 work years on the project. The project effort has been broken into three major tasks. General Electric provides the simulation. Embry-Riddle Aeronautical University provides the domain expertise. The University of Central Florida has constructed the generic part of the system which is comprised of several modules that perform the tutoring, evaluation, communication, status, etc. The generic parts of the Intelligent Simulation Training Systems (ISTS) are described.

Biegel, John E.↗

Data for A Generalized Platform for Artificial Intelligence-powered Autonomous Protein Engineering

Proteins are the molecular machines of life with numerous applications in energy, health, and sustainability. However, engineering proteins with desired functions for practical applications remains slow, expensive, and specialist-dependent. Here we report a generally applicable platform for autonomous enzyme engineering that integrates machine learning and large language models with biofoundry automation to eliminate the need for human intervention, judgement, and domain expertise. Requiring only an input protein sequence and a quantifiable way to measure fitness, this automated platform can be applied to engineer a wide array of proteins. As a proof of concept, we engineer Arabidopsis thaliana halide methyltransferase (AtHMT) for a 90-foldimprovement in substrate preference and 16-fold improvement in ethyl-transferase activity, along with developing a Yersinia mollaretii phytase (YmPhytase) variant with 26-fold improvement in activity at neutral pH. This is accomplished in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme. This platform for autonomous experimentation paves the way for rapid advancements across diverse industries, from medicine and biotechnology to renewable energy and sustainable chemistry.

AI/ML↗

Modeling software systems by domains

The Software Architectures Engineering (SAE) Project at the Software Engineering Institute (SEI) has developed engineering modeling techniques that both reduce the complexity of software for domain-specific computer systems and result in systems that are easier to build and maintain. These techniques allow maximum freedom for system developers to apply their domain expertise to software. We have applied these techniques to several types of applications, including training simulators operating in real time, engineering simulators operating in non-real time, and real-time embedded computer systems. Our modeling techniques result in software that mirrors both the complexity of the application and the domain knowledge requirements. We submit that the proper measure of software complexity reflects neither the number of software component units nor the code count, but the locus of and amount of domain knowledge. As a result of using these techniques, domain knowledge is isolated by fields of engineering expertise and removed from the concern of the software engineer. In this paper, we will describe kinds of domain expertise, describe engineering by domains, and provide relevant examples of software developed for simulator applications using the techniques.

Dippolito, Richard↗

TCP/IP Interface for the Satellite Orbit Analysis Program (SOAP)

The Transmission Control Protocol/ Internet protocol (TCP/IP) interface for the Satellite Orbit Analysis Program (SOAP) provides the means for the software to establish real-time interfaces with other software. Such interfaces can operate between two programs, either on the same computer or on different computers joined by a network. The SOAP TCP/IP module employs a client/server interface where SOAP is the server and other applications can be clients. Real-time interfaces between software offer a number of advantages over embedding all of the common functionality within a single program. One advantage is that they allow each program to divide the computation labor between processors or computers running the separate applications. Secondly, each program can be allowed to provide its own expertise domain with other programs able to use this expertise.

Carnright, Robert↗

A Scientist-in-the-Loop Data Analytics Framework for Intelligent Simulation Model Tuning and Validation

This project developed a scientist-in-the-loop data analytics framework for intelligent simulation model tuning and validation, targeting the Weather Research and Forecasting (WRF) model and its solar energy variant, WRF-Solar-BNL. Domain experts, such as climate scientists, depend on large-scale numerical simulations for knowledge discovery and decision-making, yet the complexity of parameter tuning and the disconnect between automated optimization and domain expertise pose significant challenges. We extended an interactive visual analytics framework that enables domain experts to observe and intervene in the computational steering process by identifying disagreements between the simulation model, surrogate model, and the expert’s domain knowledge. Using Bayesian Optimization with Gaussian Process Regression as the surrogate model, our system allows users to probe parameter relationships, analyze correlation patterns, and adjust tuning parameters in real time. We developed use cases for solar irradiance forecasting through sustained collaboration with Brookhaven National Laboratory, resolving critical model configuration challenges and achieving meaningful reductions in prediction error. The project supported one PhD student, one MS student, and eight undergraduate students across three Data Science Capstone projects, resulting in one master’s thesis.

Dasgupta, Aritra [New Jersey Institute of Technolo↗

Leveraging Structured Biological Knowledge for Counterfactual Inference: A Case Study of Viral Pathogenesis

Counterfactual inference is a useful tool for comparing outcomes of interventions on complex systems. It requires us to represent the system in form of a structural causal model, complete with a causal diagram, probabilistic assumptions on exogenous variables, and functional assignments. Specifying such models can be extremely difficult in practice. The process requires substantial domain expertise, and does not scale easily to large systems, multiple systems, or novel system modifications. At the same time, many application domains, such as molecular biology, are rich in structured causal knowledge that is qualitative in nature. This manuscript proposes a general approach for querying a causal knowledge graph with a causal question and converting the qualitative result into a quantitative structural causal model that can learn from data to answer the question. Here, we demonstrate the feasibility, accuracy and versatility of this approach using two case studies in systems biology. The first demonstrates the appropriateness of the underlying assumptions and the accuracy of the results. The second demonstrates the versatility of the approach by querying a knowledge base for the molecular determinants of a severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)-induced cytokine storm and performing counterfactual inference to predict the causal effect of medical countermeasures for severely ill COVID-19 patients.

60 APPLIED LIFE SCIENCES↗

Neural networks for structural design - An integrated system implementation

The development of powerful automated procedures to aid the creative designer is becoming increasingly critical for complex design tasks. In the work described here Artificial Neural Nets are applied to acquire structural analysis and optimization domain expertise. Based on initial instructions from the user an automated procedure generates random instances of structural analysis and/or optimization 'experiences' that cover a desired domain. It extracts training patterns from the created instances, constructs and trains an appropriate network architecture and checks the accuracy of net predictions. The final product is a trained neural net that can estimate analysis and/or optimization results instantaneously.

Berke, Laszlo↗

Assurance of Fault Management: Risk-Significant Adverse Condition Awareness

Fault Management (FM) systems are ranked high in risk-based assessment of criticality within flight software, emphasizing the importance of establishing highly competent domain expertise to provide assurance for NASA projects, especially as spaceflight systems continue to increase in complexity. Insight into specific characteristics of FM architectures seen embedded within safety- and mission-critical software systems analyzed by the NASA Independent Verification Validation (IVV) Program has been enhanced with an FM Technical Reference (TR) suite. Benefits are aimed beyond the IVV community to those that seek ways to efficiently and effectively provide software assurance to reduce the FM risk posture of NASA and other space missions. The identification of particular FM architectures, visibility, and associated IVV techniques provides a TR suite that enables greater assurance that critical software systems will adequately protect against faults and respond to adverse conditions. The role FM has with regard to overall asset protection of flight software systems is being addressed with the development of an adverse condition (AC) database encompassing flight software vulnerabilities.Identification of potential off-nominal conditions and analysis to determine how a system responds to these conditions are important aspects of hazard analysis and fault management. Understanding what ACs the mission may face, and ensuring they are prevented or addressed is the responsibility of the assurance team, which necessarily should have insight into ACs beyond those defined by the project itself. Research efforts sponsored by NASAs Office of Safety and Mission Assurance defined terminology, categorized data fields, and designed a baseline repository that centralizes and compiles a comprehensive listing of ACs and correlated data relevant across many NASA missions. This prototype tool helps projects improve analysis by tracking ACs, and allowing queries based on project, mission type, domain component, causal fault, and other key characteristics. The repository has a firm structure, initial collection of data, and an interface established for informational queries, with plans for integration within the Enterprise Architecture at NASA IVV, enabling support and accessibility across the Agency. The development of an improved workflow process for adaptive, risk-informed FM assurance is currently underway.

Software Verification & Validation↗

Supporting Crew Autonomy in Deep Space Exploration: Preliminary Onboard Capability Requirements and Proposed Research Questions. Technical Report of the Autonomous Crew Operations Technical Interchange Meeting

Communication delays are a critical challenge posed by long duration deep space exploration. Space missions historically have relied on an ever-present Mission Control Center (MCC) to direct operations in near real-time. As unanticipated anomalies that defeat fault detection and resolution systems do arise, the lack of real-time communication will significantly weaken what the MCC support represents: a reliable safety net for the flight crew through its deep and diverse areas of expertise and investigative resources. As a consequence, future space vehicles and habitats need to be equipped with capabilities to support the flight crew to operate with little or no ground support. Considerations must be given to vehicle and mission designs that will fortify the traditionally ground-centered safety net and forge new support systems, when communication delays exist. In August 2018, NASA’s Human Research Program, through its Human Factors and Behavioral Performance Element, convened a Technical Interchange Meeting (TIM) on Autonomous Crew Operations at NASA Ames Research Center. The goal of the meeting was to gather input from NASA centers, industry, academia, and branches of the Department of Defense (DoD) to address how intelligent technologies can be applied to augment onboard capabilities to support crew anomaly response. The TIM featured 24 presentations by 29 speakers and hosted a total of 59 attendees, including 43 from 5 NASA centers (Ames, Johnson, Langley, Marshall, and Jet Propulsion Lab) and 4 from the DoD (3 from Army Research Lab and 1 from Naval Postgraduate School), with remaining attendees from academia (e.g., UC Davis, CMU) and industry (e.g., IBM, Siemens). Discussions were centered around three themes: standards and guidelines, lessons learned in analog environments, and technologies. To help provide a framework for discussion, a concept matrix describing anomaly response processes was created prior to the TIM (Figure 1, page 6). The matrix captures the steps involved (monitoring and detection, diagnosis, solution development and evaluation, solution implementation and verification, resolution documentation) as well as the resources and capabilities required to support these steps (data, knowledge, analysis, synthesis, resource management). A wallpaper size printout of the matrix was utilized at the TIM to solicit attendee inputs along the three themes; the activity garnered 108 submissions of ideas. Overall, what emerged from TIM discussions was a picture of mismatch between crew anomaly response needs and support that can be provided by existing intelligent technologies. The needs are broad, spanning multiple steps and processes/resources, with many of which lacking support from existing technologies, such as knowledge management throughout the steps of problem solving (especially in resolution documentation) and manpower management. The solutions provided by existing intelligent technologies are specific to the steps/processes that they are designed to support and constrained to solving only problems similar to those that have occurred before. What is lacking from technologies is typically made up by humans, specifically their complex critical thinking, creative problem solving, and domain expertise. In the end, the TIM highlighted the pressing need to support responses to onboard anomalies during autonomous crew operations, particularly those that have eluded the system tests, inspection, and other assurance processes. Such anomalies can potentially threaten crew and vehicle safety, as well as significantly impact overall operations with additional workload. These fairly rare events are difficult to anticipate and prepare for, given the state-of-the-art in intelligent technologies. This is true even for anomalies that stem from “unknown knowns”—cases in which there is sufficient external information to characterize the problem but the overall pattern fails to be recognized by the problem solver, or in which the internal knowledge needed to solve a problem is held tacitly and potentially accessible by the problem solver but not articulated. It follows that the ability to tackle anomalies lies not only with the availability of relevant information and knowledge but also their accessibility in times of need. To that end, we propose research questions along the following three broad themes: • How intelligent technologies can help make relevant knowledge and information available? • How intelligent technologies can help make relevant knowledge and information accessible? • How intelligent technologies can help support the crew operating as a team in anomaly response processes?

autonomous crew operations↗

Advances in Machine and Deep Learning for Modeling and Real-time Detection of Multi-Messenger Sources

We live in momentous times. The science community is empowered with an arsenal of cosmic messengers to study the universe in unprecedented detail. Gravitational waves, electromagnetic waves, neutrinos, and cosmic rays cover a wide range of wavelengths and timescales. Combining and processing these datasets that vary in volume, speed, and dimensionality requires new modes of instrument coordination, funding, and international collaboration with a specialized human and technological infrastructure. In tandem with the advent of large-scale scientific facilities, the last decade has experienced an unprecedented transformation in computing and signal-processing algorithms. The combination of graphics processing units, deep learning, and the availability of open source, high-quality datasets has powered the rise of artificial intelligence. This digital revolution now powers a multibillion dollar industry, with far-reaching implications in technology and society. In this chapter, we describe pioneering efforts to adapt artificial intelligence algorithms to address computational grand challenges in multi-messenger astrophysics. We review the rapid evolution of these disruptive algorithms, from the first class of algorithms introduced in early 2017 to the sophisticated algorithms that now incorporate domain expertise in their architectural design and optimization schemes. We discuss the importance of scientific visualization and extreme-scale computing in reducing time-to-insight and obtaining new knowledge from the interplay between models and data.

Artificial Intelligence↗

Advancing Fusion with Machine Learning Research Needs Workshop Report

Abstract Machine learning and artificial intelligence (ML/AI) methods have been used successfully in recent years to solve problems in many areas, including image recognition, unsupervised and supervised classification, game-playing, system identification and prediction, and autonomous vehicle control. Data-driven machine learning methods have also been applied to fusion energy research for over 2 decades, including significant advances in the areas of disruption prediction, surrogate model generation, and experimental planning. The advent of powerful and dedicated computers specialized for large-scale parallel computation, as well as advances in statistical inference algorithms, have greatly enhanced the capabilities of these computational approaches to extract scientific knowledge and bridge gaps between theoretical models and practical implementations. Large-scale commercial success of various ML/AI applications in recent years, including robotics, industrial processes, online image recognition, financial system prediction, and autonomous vehicles, have further demonstrated the potential for data-driven methods to produce dramatic transformations in many fields. These advances, along with the urgency of need to bridge key gaps in knowledge for design and operation of reactors such as ITER, have driven planned expansion of efforts in ML/AI within the US government and around the world. The Department of Energy (DOE) Office of Science programs in Fusion Energy Sciences (FES) and Advanced Scientific Computing Research (ASCR) have organized several activities to identify best strategies and approaches for applying ML/AI methods to fusion energy research. This paper describes the results of a joint FES/ASCR DOE-sponsored Research Needs Workshop on Advancing Fusion with Machine Learning, held April 30–May 2, 2019, in Gaithersburg, MD (full report available at https://science.osti.gov/-/media/fes/pdf/workshop-reports/FES_ASCR_Machine_Learning_Report.pdf ). The workshop drew on broad representation from both FES and ASCR scientific communities, and identified seven Priority Research Opportunities (PRO’s) with high potential for advancing fusion energy. In addition to the PRO topics themselves, the workshop identified research guidelines to maximize the effectiveness of ML/AI methods in fusion energy science, which include focusing on uncertainty quantification, methods for quantifying regions of validity of models and algorithms, and applying highly integrated teams of ML/AI mathematicians, computer scientists, and fusion energy scientists with domain expertise in the relevant areas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

EnergyPlus-MCP: A model-context-protocol server for ai-driven building energy modeling

Traditional building energy modeling with the EnergyPlus building performance simulation engine requires domain expertise, programming skills, and intensive manual efforts limiting its effective adoption. This paper introduces EnergyPlus-MCP, the first open-source Model Context Protocol (MCP) server specifically designed for EnergyPlus simulation workflows, establishing a new foundational infrastructure for AI-driven building energy modeling. The MCP server implements a layered architecture with 35 specialized tools spanning model management, editing and analysis, HVAC and other systems configuration inspection, and simulation execution, enabling Large Language Models to interact with EnergyPlus through conversational interfaces. The server addresses critical workflow barriers by automating model validation, streamlining energy efficiency measures modification, and providing intelligent output management with interactive visualization. Through practical demonstrations using a multi-zone building retrofit analysis, we show how the EnergyPlus-MCP server significantly reduces manual efforts while maintaining full simulation rigor. By providing accessible natural language interfaces to sophisticated building energy analysis, this approach enables scalable deployment of simulation expertise across public and private organizations, educational institutions, and research teams, fundamentally transforming traditional building energy modeling practices.

AI↗

Importance of Engineered and Learned Molecular Representations in Predicting Organic Reactivity, Selectivity, and Chemical Properties

Machine-readable chemical structure representations are foundational in all attempts to harness machine learning for the prediction of reactivities, selectivities, and chemical properties directly from molecular structure. The featurization of discrete chemical structures into a continuous vector space is a critical phase undertaken before model selection, and the development of new ways to quantitatively encode molecules is an active area of research. Here, we highlight the application and suitability of different representations, from expert-guided “engineered” descriptors to automatically “learned” features, in different prediction tasks relevant to organic and organometallic chemistry, where differing amounts of training data are available. These tasks include statistical models of stereo- and enantioselectivity, thermochemistry, and kinetics developed using experimental and quantum chemical data. The use of expert-guided molecular descriptors provides an opportunity to incorporate chemical knowledge, domain expertise, and physical constraints into statistical modeling. In applications to stereoselective organic and organometallic catalysis, where data sets may be relatively small and 3D-geometries and conformations play an important role, mechanistically informed features can be used successfully to obtain predictive statistical models that are also chemically interpretable. We provide an overview of several recent applications of this approach to obtain quantitative models for reactivity and selectivity, where topological descriptors, quantum mechanical calculations of electronic and steric properties, along with conformational ensembles, all feature as essential ingredients of the molecular representations used. Alternatively, more flexible, general-purpose molecular representations such as attributed molecular graphs can be used with machine learning approaches to learn the complex relationship between a structure and prediction target. This approach has the potential to out-perform more traditional representation methods such as “hand-crafted” molecular descriptors, particularly as data set sizes grow. One area where this is particularly relevant is in the use of large sets of quantum mechanical data to train quantitative structure–property relationships. A general approach toward curating useful data sets and training highly accurate graph neural network models is discussed in the context of organic bond dissociation enthalpies, where this strategy outperforms regression using precomputed descriptors. Finally, we describe how graph neural network predictions can be incorporated into mechanistically informed statistical models of chemical reactivity and selectivity. Once trained, this approach avoids the expensive computational overhead associated with quantum mechanical calculations, while maintaining chemical interpretability. We illustrate examples for which fast predictions of bond dissociation enthalpy and of the identities of radicals formed through cleavage of a molecule’s weakest bond are used in simple physical models of site-selectivity and reactivity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A generalized platform for artificial intelligence-powered autonomous enzyme engineering

Proteins are the molecular machines of life with numerous applications in energy, health, and sustainability. However, engineering proteins with desired functions for practical applications remains slow, expensive, and specialist-dependent. Here we report a generally applicable platform for autonomous enzyme engineering that integrates machine learning and large language models with biofoundry automation to eliminate the need for human intervention, judgement, and domain expertise. Requiring only an input protein sequence and a quantifiable way to measure fitness, this automated platform can be applied to engineer a wide array of proteins. As a proof of concept, we engineer Arabidopsis thaliana halide methyltransferase (AtHMT) for a 90-fold improvement in substrate preference and 16-fold improvement in ethyltransferase activity, along with developing a Yersinia mollaretii phytase (YmPhytase) variant with 26-fold improvement in activity at neutral pH. This is accomplished in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme. This platform for autonomous experimentation paves the way for rapid advancements across diverse industries, from medicine and biotechnology to renewable energy and sustainable chemistry.

59 BASIC BIOLOGICAL SCIENCES↗

Ontologizing health systems data at scale: making translational discovery a reality

Common data models solve many challenges of standardizing electronic health record (EHR) data but are unable to semantically integrate all of the resources needed for deep phenotyping. Open Biological and Biomedical Ontology (OBO) Foundry ontologies provide computable representations of biological knowledge and enable the integration of heterogeneous data. However, mapping EHR data to OBO ontologies requires significant manual curation and domain expertise. We introduce OMOP2OBO, an algorithm for mapping Observational Medical Outcomes Partnership (OMOP) vocabularies to OBO ontologies. Using OMOP2OBO, we produced mappings for 92,367 conditions, 8611 drug ingredients, and 10,673 measurement results, which covered 68–99% of concepts used in clinical practice when examined across 24 hospitals. When used to phenotype rare disease patients, the mappings helped systematically identify undiagnosed patients who might benefit from genetic testing. By aligning OMOP vocabularies to OBO ontologies our algorithm presents new opportunities to advance EHR-based deep phenotyping.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

ChemPix: automated recognition of hand-drawn hydrocarbon structures using deep learning

Inputting molecules into chemistry software, such as quantum chemistry packages, currently requires domain expertise, expensive software and/or cumbersome procedures. Leveraging recent breakthroughs in machine learning, we develop ChemPix: an offline, hand-drawn hydrocarbon structure recognition tool designed to remove these barriers. A neural image captioning approach consisting of a convolutional neural network (CNN) encoder and a long short-term memory (LSTM) decoder learned a mapping from photographs of hand-drawn hydrocarbon structures to machine-readable SMILES representations. We generated a large auxiliary training dataset, based on RDKit molecular images, by combining image augmentation, image degradation and background addition. Additionally, a small dataset of ~600 hand-drawn hydrocarbon chemical structures was crowd-sourced using a phone web application. These datasets were used to train the image-to-SMILES neural network with the goal of maximizing the hand-drawn hydrocarbon recognition accuracy. By forming a committee of the trained neural networks where each network casts one vote for the predicted molecule, we achieved a nearly 10 percentage point improvement of the molecule recognition accuracy and were able to assign a confidence value for the prediction based on the number of agreeing votes. The ensemble model achieved an accuracy of 76% on hand-drawn hydrocarbons, increasing to 86% if the top 3 predictions were considered.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗