Search NASASearch

SEARCH · Search NASA

Results for “Generative AI”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

On the Abuse and Detection of Polyglot Files

A polyglot is a file that is valid in two or more formats. Polyglot files pose a problem for file-upload and generative AI web interfaces that rely on format identification to determine how to securely handle incoming files. In this work we found that existing file-format and embedded-file detection tools, even those developed specifically for polyglot files, fail to reliably detect polyglot files used in the wild. To address this issue, we studied the use of polyglot files by malicious actors in the wild, finding 30 polyglot samples and 15 attack chains that leveraged polyglot files. Using knowledge from our survey of polyglot usage in the wild---the first of its kind---we created a novel data set based on adversary techniques. We then trained a machine learning detection solution, PolyConv, using this data set. PolyConv achieves a precision-recall area-under-curve score of 0.999 with an F1 score of 99.20% for polyglot detection and 99.47% for file-format identification, significantly outperforming all other tools tested. We developed a content disarmament and reconstruction tool, ImSan, that successfully sanitized 100% of the tested image-based polyglots, which were the most common type found via the survey. Our work provides concrete tools and suggestions to enable defenders to better defend themselves against polyglot files, as well as directions for future work to create more robust file specifications and methods of disarmament.

Oesch, T [ORNL] (ORCID:0000000269091022)

Geospatial Diffusion for Land Cover Imperviousness Change Forecasting

Land-use and land-cover (LULC) has a significant effect on several Earth system processes. For example, impervious surfaces reduce infiltration and speed water flow, impacting regional hydrology and flood risk. While Earth System models have improved forecasting hydrologic and atmospheric processes at higher resolutions, the ability to forecast LULC change has lagged behind. In this paper, we propose a new paradigm exploiting Generative AI (GenAI) for land cover change forecasting by framing it as a data synthesis problem conditioned on historical and auxiliary data-sources. To demonstrate the feasibility of our methodology, we perform experiments where a diffusion model is trained for decadal forecasting of imperviousness change across the entire United States. We find that our model yields MAE lower than a no-change baseline for resolutions ≥ 0.7 X 0.7km2 on average, demonstrating its ability to capture and project accurate spatiotemporal patterns. Finally, we discuss future research to incorporate Earth's physical properties and enabling scenario simulations via driver variables.

Varshney, Debvrat [ORNL] (ORCID:0000000188981736)

PTMdiscoverer

PTMdiscoverer is a new approach that leverages generative AI to accelerate the discovery of post-translational modifications (PTMs) in proteomics experiments.

Bilbao, Aivett [Pacific Northwest National Laborat

Bioenergy Research Centers Data Sharing Portal

The bioenergy.org website is the end product of the Data Sharing Shared Research objective for the Bioenergy Centers. The objective is to Enhance BRC data legacy through the development and use of shared software tools to make previously published datasets more findable and accessible through the Inter-BRC Data Products Portal, and to explore the use of generative AI to assist in the exploration of published datasets. The software repository is released per the license information below.

Thrower, Nicholas

WorkJournalMaker (WJMaker) v0.5

The software generates and maintains daily work journal entries in text format, via a web browser. The journal entries are saved in a structured directory file tree on the system running the software. The software also incorporates a database so that it can track the location of files in the file system and various other metadata. The software allows the users to access their journal entries either through the browser or as discrete text files, facilitating sharing and open science. Additionally, to assist with the yearly PMP process, this tool connects to LLM APIs to provide summarization of the journal entries on a month-by-month or weekly basis. The advantage over similar technologies such as Apple Notes (extremely popular for notetaking) is that the instant software does not force the user to stay inside the Apple ecosystem, since it allows for export of the user's text files. This facilitates open science, so that researchers who use the tool can easily transfer their research notes to any other system. The WebJournalMaker repository is here: https://github.com/lbnl-science-it/WorkJournalMaker The WebJournalMaker repository is forked from the JournalSummarizer: https://github.com/tyfong-lbl/JournalSummarizer and builds on its code. I wrote the code for both of these software repos, using generative AI.

Fong, Timothy [Lawrence Berkeley National Laborato

SEED: Semantic Energy Exploration and Discovery

The Bioenergy Knowledge Discovery Framework (KDF) hosts a vast repository of specialized data, yet traditional keyword-based search methods often struggle to provide direct answers, requiring significant domain expertise and manual effort to filter through raw documents. To overcome these barriers, this software introduces a semantic search engine that enables both specialists and non-specialists to query the KDF using natural language. By shifting from rigid keyword matching to intent-based retrieval, the tool automatically identifies and ranks the most relevant sources within the database. The system functions by processing natural language queries to extract the most pertinent information, delivering an AI-generated plain-language summary alongside exact supporting quotes from retrieved documents. This integrated approach provides users with immediate, evidence-based answers while eliminating the need for exhaustive manual review. By surfacing direct insights and contextual evidence, the software enhances the usability of existing KDF resources and democratizes access to complex bioenergy data. Ultimately, this semantic search solution accelerates the discovery process and supports faster, more informed decision-making across the bioenergy sector.

Pan, Meiyu (Melrose) [Oak Ridge National Laborator

alchemist-nn

The alchemist-nn code provides a framework for training and running generative AI models for molecular structures.

Rogers, David [Oak Ridge National Laboratory (ORNL

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics

Assessing the nature of large language models: A caution against anthropocentrism.

Generative AI models garnered a large amount of public attention and speculation with the release of OpenAI’s chatbot, ChatGPT in November of 2022. At least two opinion camps exist – one that is excited about the possibilities these models offer for fundamental changes to human tasks, and another that is highly concerned about the power these models seem to have – especially since the release of GPT-4, which was trained on multimodal data and has ~1.7 trillion (T) parameters. We evaluated some concerns regarding these models’ power by assessing GPT-3.5 using standard, normed, and validated cognitive and personality measures. These measures come from the tradition of psychometrics in experimental psychology and have a long history of providing valuable insights and predictive distinctions in humans. For this seedling project, we developed a battery of tests that allowed us to estimate the boundaries of some of these models’ capabilities, how stable those capabilities are over a short period of time, and how they compare to humans.

97 MATHEMATICS AND COMPUTING

AI for Nuclear Safeguards Verification

The International Atomic Energy Agency (IAEA) utilizes AI/ML to analyze open-source information, including satellite imagery and scientific publications, to verify the completeness of State declarations regarding nuclear activities. AI/ML already assist the IAEA with automating processes and analysis of large datasets, including satellite imagery and unstructured data, improving efficiency and effectiveness of safeguards implementation. AI/ML in nuclear safeguards come with its own challenges that include the need for large, unbiased datasets, the risk of AI-generated fake information, including the potential for manipulation of satellite imagery.

97 MATHEMATICS AND COMPUTING

Data Center Market Report

The data center market is poised to explode in the coming decade due to undeniable drivers such as continued adoption of generative AI, increased data storage needs, and enterprise integration of AI in numerous industries [1] [2] [3]. Scalable power and increased computational capacity are at the forefront of considerations for hyperscalers, the major cloud service providers in this space. Lawrence Livermore National Laboratory is uniquely poised to help with informed decision making for data center market leaders during this phase of explosive expansion. National grid modeling expertise and cutting edge innovations in computer cooling systems place LLNL in an enviable position for creating economic impact in the data center industry by leveraging its expertise in these areas which can help the data center market keep up with growing demand.

97 MATHEMATICS AND COMPUTING

U.S. Solar Siting Regulation and Zoning Ordinances (2025)

A machine readable collection of documented solar siting ordinances at the state and local (e.g., county, township) level throughout the United States. The data were compiled using the Infrastructure Continuous Ordinance Mapping for Planning and Siting Systems (INFRA-COMPASS) tool, which leverages Large Language Models (LLMs) to automate the collection of local codes and ordinances applicable to energy infrastructure. URLs for the ordinance source documents are included in the Solar Ordinances spreadsheet. The GeoPackage file included below contains the jurisdiction shapes for each ordinance. Note that the GeoPackage file is formatted for ingestion by NLR's reVX setbacks tool and therefore does not contain any of the state-level regulations. NOTE: This data was collected with the help of generative AI. The Large Language Models used for this effort make mistakes. Always validate the data for critical use cases. This data is an update to a previously developed database of wind ordinances found in OEDI Submission 5734: see the "U.S. Solar Siting Regulation and Zoning Ordinances 2022" link below. INFRA-COMPASS version used for collection: v0.11.3 LLMs used for collection: GPT-4.1, GPT-4.1 mini, GPT-4.1 nano

14 SOLAR ENERGY

U.S. Wind Siting Regulation and Zoning Ordinances (2025)

A machine readable collection of documented wind siting ordinances at the state and local (e.g., county, township) level throughout the United States. The data were compiled using the Infrastructure Continuous Ordinance Mapping for Planning and Siting Systems (INFRA-COMPASS) tool, which leverages Large Language Models (LLMs) to automate the collection of local codes and ordinances applicable to energy infrastructure. URLs for the ordinance source documents are included in the Wind Ordinances spreadsheet. The GeoPackage file included below contains the jurisdiction shapes for each ordinance. Note that the GeoPackage file is formatted for ingestion by NREL's reVX setbacks tool and therefore does not contain any of the state-level regulations. NOTE: This data was collected with the help of generative AI. The Large Language Models used for this effort make mistakes. Always validate the data for critical use cases. This data is an update to a previously developed database of wind ordinances found in OEDI Submission 5733: see the "U.S. Wind Siting Regulation and Zoning Ordinances 2022" link below. INFRA-COMPASS version used for collection: v0.8.2 LLMs used for collection: GPT-4.1, GPT-4.1 mini, GPT-4.1 nano, GPT-4o mini

17 WIND ENERGY

Prime Time for Model-Predictive Control? Assessing the Technical and Market Readiness of Advanced Controls in Buildings

Despite three decades of extensive research and field testing that have consistently validated the benefits of Model Predictive Control (MPC) in building applications, the technology has seen limited market adoption. This paper evaluates the readiness of MPC for widespread deployment, showcases recent demonstrations and field tests across diverse building types, including residential, small commercial, large commercial, and campus settings. Our results demonstrate that MPC can optimize system operations to achieve load shifting, minimize curtailment of on-site generation, and reduce energy costs by up to 80 %, while maintaining or improving occupant comfort. We also show that MPC can effectively control large assets, such as MW-sized thermal storage systems, and respond to dynamic pricing signals. However, achieving scale remains difficult due to labor-intensive workflows, reliance on a “PhD-in-the-loop” for MPC design and maintenance, susceptibility to fragile data infrastructure, and persistent workforce education and acceptance barriers. To bridge this gap, we outline a transition from bespoke, labor intensive prototypes toward streamlined, segment-targeted deployment strategies that leverage model templates, semantic tools, and generative AI. By automating control configuration and reducing engineering effort, these recommendations provide a pathway for transforming successful research demonstrations into scalable, market ready solutions for MPC-based controls.

Pritoni, Marco

AI-assisted rapid crystal structure generation towards a target local environment

In material design, traditional crystal structure prediction approaches are expensive as they require extensive structural sampling through expensive energy minimization methods. Emerging artificial intelligence (AI) generative models have shown great promise in rapidly generating realistic crystals, but they typically handle only a few tens of atoms per unit cell. To overcome this limitation, we introduce a symmetry-informed approach, the Local Environment Geometry-Oriented Crystal Generator (LEGO-xtal). Our method generates initial structures using AI models trained on an augmented dataset, and then optimizes them using structure descriptors rather than energy-based optimization. We demonstrate its effectiveness by expanding from 25 known low-energy sp2 carbon allotropes to over 1700, all within 0.5 eV/atom of the ground-state energy of graphite. This framework offers a generalizable strategy for the targeted design of materials with modular building blocks, such as metal-organic frameworks and battery materials.

Ridwan, Osman Goni [University of North Carolina a

TRIM: AI Guided Random Number Generation for Resource-Constrained IoT Systems

Random numbers often serve as the backbone for many security solutions in diverse domains such as cryptography, side channel leakage prevention, and moving target defense. However, generating true random numbers requires a physical source of entropy (e.g. hardware, quantum, environmental phenomenon) making it difficult to realize at a large scale and at a low cost. On the flip side, pseudorandom number generators (easy to implement) following a specific distribution (e.g. Gaussian) can be easily compromised given a sufficient amount of traces. In this work, we have developed a machine learning-guided generative approach that can be used to create portable, resource-efficient, and cost-effective random number generators with high throughput and true randomness characteristics. We implement the proposed approach as a highly parameterized framework and perform extensive evaluation for different settings. The framework was able to learn from true random sources such as irrational numbers and environmental audio noise and imitate those sources towards generating new good quality random numbers on demand. We have generated more than 1 billion bits and observed robust performance in terms of true randomness metrics obtained from NIST SP 800-22 and FIPS 140-1 randomness test suites achieving a throughput of up to 142.85 Mbps. Compared to the state-of-the-art (SOTA) technique, the iso-cost setup of our framework can achieve more than 500 Mbps in a distributed setting. We have evaluated the efficacy of running the true randomness imitation AI models on target edge devices such as Raspberry Pi 4 (Model B), Nvidia Jetson Nano, Nvidia Jetson Orin Nano and Nvidia Jetson Xavier. We have also looked at the security of the TRIM framework itself against different adversarial threat models.

Cybersecurity

AI-Assisted Conceptual Development of a Pre-Geometric Cosmological Model - An Exercise in AI-Assisted Conceptual Framework Generation, Paper III: Cosmological Structure and Predictions

This paper develops the cosmological consequences of the replication-driven cosmogenesis framework introduced in Paper I and the emergent geometric structure established in Paper II. After the replication epoch freezes out, the coherent sector occupies a finite spectral band and contains a population of excited states. The relaxation of these excited coherent configurations does not produce coherent radiation; instead, all released energy flows into the incoherent substrate, where the randomizer acts as a rapid phase-scrambling mechanism. This process generates an effectively thermal radiation bath, providing a natural reheating mechanism that requires neither inflaton oscillations nor scalar-field potentials, and can be contrasted with standard scenarios of nonperturbative reheating dynamics. Subsequent symmetry-breaking transitions in the coherent vacuum inject additional radiation, yielding a multi-stage thermal history with well-defined energy transfers. We derive the effective equations of state for each component—the cosmological vacuum, the coherent vacuum, and the radiation bath—and show how their interplay produces an FRW-like expansion. The discrete sequence of coherent-state relaxations imprints a distinctive multi-peaked stochastic gravitational-wave background, whose spectral structure reflects the underlying hierarchy of coherent frequencies. Potential observational signatures in the LISA and mid-band frequency ranges are highlighted, providing concrete avenues to test this replication-based cosmological framework in the context of standard cosmological gravitational-wave backgrounds and LISA-oriented forecasts.

79 ASTRONOMY AND ASTROPHYSICS

AI-Assisted Conceptual Development of a Pre-Geometric Cosmological Model - An Exercise in AI-Assisted Conceptual Framework Generation, Paper II: Local Geometry and Metric Structure

This paper develops the geometric sector of the replication-driven cosmogenesis framework introduced in Paper I. Starting from a pre-geometric spectral substrate and a minimal set of replication axioms, we show how coherent self-replicating units generate a spatial adjacency graph whose continuum limit acquires an effective Riemannian structure. The replication dynamics determines a characteristic correlation length that seeds the local metric, while overlap relations among coherent units produce an isotropic neighborhood geometry with an emergent dimensionality $d_{\rm eff}\simeq 3$ across a broad range of replication factors. As replication slows and causal order stabilizes, a limiting signal speed $c_\ast$ appears, providing the basis for the Lorentzian structure of spacetime without assuming a pre-existing light cone. We derive conditions under which the adjacency graph converges to a smooth three-dimensional manifold, describe the transition from Euclidean to Lorentzian propagation, and identify geometric invariants controlled by the replication parameters. This work establishes the geometric and causal layer of the replication cosmogenesis program, bridging the spectral axioms of Paper I to the cosmological dynamics explored in Paper III.

79 ASTRONOMY AND ASTROPHYSICS