Search NASA⌕ Search

SEARCH · Search NASA

Results for “Modeling Languages”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Activation Domain Hunter (ADhunter) v2.0

ADhunter is a software program that enables accurate identification and quantification of transcriptional activation domains. Unlike previous software, ADhunter uses protein representations from a pre-trained protein language model, model ensembling, and a training dataset from a diverse sampling of protein sequence space for state-of-the-art performance. These advantages enable improved perception of transcriptional activation domains across sequence space that can be used for mapping natural genetic circuits and engineering synthetic genetic circuits. In particular, ADhunter enables fine-tuned control of gene expression through synthetic transcription factors that can be used for complex control of cellular programs.

Waldburger, Lucas [Lawrence Berkeley National Labo↗

ellora-spack-gen

The project contains software to analyze the behavior of large language models at code generation tasks. Publicly available information is used to generate Spack package recipes for HPC developers. The software contains orchestration tooling, analysis, and plotting functionality.

Melone, CaetanoN↗

FAIR to WISE (F2W) v1.0.0

FAIR to WISE (F2W) is an iterative, large-language model (LLM) driven pipeline that turns unstructured research PDFs into structured, queryable knowledge graphs (KGs). Core features include schema-driven extraction to a LinkML model; full provenance capture; ontology-grounded enrichment (e.g., chemical validation and ChEBI lookup); graph construction to JSON-LD with stable IDs; and KG-RAG question answering with evidence-aware retrieval. The system is engineered for reproducibility and accessibility (open-source Ollama models, temperature=0, NVTX/Nsight profiling) with robust QA (relation verification, deduplication, and deterministic outputs). Primary uses are literature-to-KG automation, knowledge-grounded Q&A, and experimental steering support. We demonstrate the approach in organic photovoltaics, where the pipeline ingests papers, builds a domain KG, and evaluates answers against expert competency questions to guide experimental planning and interpretation. Compared with off-the-shelf LLMs and ad-hoc NLP tools, F2W addresses ontology gaps and reduces hallucination risk by grounding responses in extracted evidence and enforcing schema constraints; it also offers deterministic, provenance-linked outputs and open, cost-aware deployment. Evidence-aware ranking further improves answer quality over pure vector search.

Abramov, David [Lawrence Berkeley National Laborat↗

3_wise_bears

This is a small set of python scripts and HTML that uses OpenAI API to control large language models (LLMs) working agentically to solve a posed question / problem. There are 3 agents and they work to get it "just right" by taking on various "roles" of friendly and adversarial critics. It repeats a number of times specified by the user, and then writes a report.

DeBardeleben, Nathan Andrew [Los Alamos National L↗

Topological Signatures of Adversaries in Multimodal Alignments

Topological Data Analysis for Adversarial Detection (LANL O4937) - Detects adversarial examples in vision-language models using persistent homology and two-sample testing. Combines TDA features from CLIP embeddings with statistical methods (ME, SCF, SAMMD, C2ST) for robust detection across ImageNet, CIFAR-10/100.

Bhattarai, Manish↗

Atlas-UI-3

SAND2025-14754O Atlas-UI-3 is the third iteration of the LLM chat user interface called Atlas. The tool does not share code with previous versions. It supports chat with large language models (LLMs), connects to retrieval-augmented generation (RAG) servers, features a marketplace/store for model and content providers (MCP), and enables tool calling. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Melander, Darryl [Sandia National Lab. (SNL-CA), L↗

bibcheck

SAND2026-16981O Bibcheck is designed to extract bibliographies from research papers and perform metadata searches to identify errors. It assists authors in checking their bibliographies for metadata errors during the writing process and helps reviewers identify errors in bibliographies of papers under review. The software uses large language models (LLMs) to extract bibliography entries from PDF documents, classifies the type of bibliography entry, and verifies referenced works. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Pearson, Carl [Sandia National Lab. (SNL-CA), Live↗

TalkPipe Writing Assistant

SAND2025-14316O TalkPipe Writing Assistant offers AI assistance, providing help on a point-by-point basis. Authors can start with their own ideas—whether bullet points, partial paragraphs, or phrases—and specify the document type, desired tone, audience, and any other relevant context. As they write, the assistant provides tailored suggestions for each paragraph. They can request high-level concepts, draft a paragraph, or proofread existing text. The large language model (LLM) considers both preceding and following paragraphs to ensure coherence and flow. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Bauer, Travis [Sandia National Lab. (SNL-CA), Live↗

pas

PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models

Bhattarai, Manish↗

MTRE

Multi-Token Reliability Estimation (MTRE) is a lightweight, white-box hallucination detector for vision-language models. Instead of using only the first output token, MTRE aggregates logits from the first ~10 tokens and feeds them to a small attention-based reliability head; per-token scores are combined via a sequential log-likelihood-ratio test with early-stopping, and an MTRE-t variant calibrates thresholds via cross-fitting. MTRE reports average gains of +9.4% Accuracy and +14.8% AUROC over common baselines across MAD-Bench, MM-SafetyBench, MathVista, and arithmetic/counting tasks, while adding ~4.3M params and ~1% inference overhead (~26 MB VRAM, ~0.94 ms per detection). Key limitation: requires access to early token logits and is evaluated on a handful of open-source 7B VLMs.

Bhattarai, Manish [Los Alamos National Labs]↗

LLM Information Extraction Toolkit

A modular Python framework for information extraction using large language models with support for multiple backends and optional verification workflows.

Yoon, Hong-Jun [Oak Ridge National Laboratory (ORN↗

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat↗

MADA: Multi-Agent Design Assistant

MADA (Multi-Agent Design Assistant) is a Large Language Model (LLM) powered multi-agent framework that coordinates specialized agents for complex design workflows. The system was designed for HPC workflows with the following agents in mind: 1) A Job Management Agent (JMA) launches and manages ensemble simulations on HPC systems, 2) a Geometry Agent (GA) generates meshes, and 3) an Inverse Design Agent (IDA) proposes new designs informed by simulation outcomes. Our framework reduces cumbersome manual workflow setup, and enables automated design exploration at scale. However, the software also enables users to rapidly create new multi-agent systems. Simply define new agents in a configuration file, giving each their own set of tools (via MCP), and then chat and prompt your new multi-agent system. Is

Gunnarson, BrianS [Lawrence Livermore National Lab↗

The Artificial Intelligence Ontology: LLM-Assisted Construction of AI Concept Hierarchies

The Artificial Intelligence Ontology (AIO) is a systematization of artificial intelligence (AI) concepts, methodologies, and their interrelations. Developed via manual curation, with the additional assistance of large language models (LLMs), AIO aims to address the rapidly evolving landscape of AI by providing a comprehensive framework that encompasses both technical and ethical aspects of AI technologies. The primary audience for AIO includes AI researchers, developers, and educators seeking standardized terminology and concepts within the AI domain. We use the term “branches” for classes, and their subclasses, in our ontology that are subclasses of owl:Thing. AIO contains eight branches: Bias, Layer, Machine Learning Task, Mathematical Function, Model, Network, Preprocessing, and Training Strategy, each designed to support the modular composition of AI methods and facilitate a deeper understanding of deep learning architectures and ethical considerations in AI. AIO uses the Ontology Development Kit (ODK) for its creation and maintenance, with its content being more easily updated through AI-driven curation support. This approach not only ensures the ontology's relevance amidst the fast-paced advancements in AI but also significantly enhances its utility for researchers, developers, and educators by simplifying the integration of new AI concepts and methodologies. The ontology's utility is demonstrated through the annotation of AI methods data in a catalog of AI research publications and the integration into the BioPortal ontology resource, highlighting its potential for cross-disciplinary research. The AIO ontology is open source and is available on GitHub ( https://w3id.org/aio/ ) and BioPortal ( https://bioportal.bioontology.org/ontologies/AIO ).

Joachimiak, Marcin P. [Biosystems Data Science Dep↗

Data for A Generalized Platform for Artificial Intelligence-powered Autonomous Protein Engineering

Proteins are the molecular machines of life with numerous applications in energy, health, and sustainability. However, engineering proteins with desired functions for practical applications remains slow, expensive, and specialist-dependent. Here we report a generally applicable platform for autonomous enzyme engineering that integrates machine learning and large language models with biofoundry automation to eliminate the need for human intervention, judgement, and domain expertise. Requiring only an input protein sequence and a quantifiable way to measure fitness, this automated platform can be applied to engineer a wide array of proteins. As a proof of concept, we engineer Arabidopsis thaliana halide methyltransferase (AtHMT) for a 90-foldimprovement in substrate preference and 16-fold improvement in ethyl-transferase activity, along with developing a Yersinia mollaretii phytase (YmPhytase) variant with 26-fold improvement in activity at neutral pH. This is accomplished in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme. This platform for autonomous experimentation paves the way for rapid advancements across diverse industries, from medicine and biotechnology to renewable energy and sustainable chemistry.

AI/ML↗

Assessing the evolution of research topics in a biological field using plant science as an example

Scientific advances due to conceptual or technological innovations can be revealed by examining how research topics have evolved. But such topical evolution is difficult to uncover and quantify because of the large body of literature and the need for expert knowledge in a wide range of areas in a field. Using plant biology as an example, we used machine learning and language models to classify plant science citations into topics representing interconnected, evolving subfields. The changes in prevalence of topical records over the last 50 years reflect shifts in major research trends and recent radiation of new topics, as well as turnover of model species and vastly different plant science research trajectories among countries. Our approaches readily summarize the topical diversity and evolution of a scientific field with hundreds of thousands of relevant papers, and they can be applied broadly to other fields.

60 APPLIED LIFE SCIENCES↗

LLM integration into EPICS

The utilization of large language models (LLMs) such as ChatGPT has seen a remarkable increase in various fields over the past few years. These models have demonstrated their versatility and capability in understanding and generating human-like text, making them invaluable tools in numerous applications. In this project, we explore the integration of a LLM into the Experimental Physics and Industrial Control System (EPICS). The primary focus of this integration is to employ the LLM for advanced image processing and spatial analysis on images obtained from the beamlines. By leveraging the capabilities of the LLM, we aim to enhance the accuracy and efficiency of image interpretation, enabling more precise data analysis and decision-making within the EPICS framework. This integration not only showcases the potential of LLMs in scientific and industrial applications but also sets the stage for future advancements in automated control systems.

Adams, Ethan↗