Search NASA⌕ Search

SEARCH · Search NASA

Results for “Embedding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Gradient Coding With Iterative Block Leverage Score Sampling

Gradient coding is a method for mitigating straggling servers in a centralized computing network that uses erasure-coding techniques to distributively carry out first-order optimization methods. Randomized numerical linear algebra uses randomization to develop improved algorithms for large-scale linear algebra computations. In this study, we propose a method for distributed optimization that combines gradient coding and randomized numerical linear algebra. The proposed method uses a randomized ℓ 2 -subspace embedding and a gradient coding technique to distribute blocks of data to the computational nodes of a centralized network, and at each iteration the central server only requires a small number of computations to obtain the steepest descent update. The novelty of our approach is that the data is replicated according to importance scores, called block leverage scores, in contrast to most gradient coding approaches that uniformly replicate the data blocks. Furthermore, we do not require a decoding step at each iteration, avoiding a bottleneck in previous gradient coding schemes. We show that our approach results in a valid ℓ 2 -subspace embedding, and that our resulting approximation converges to the optimal solution.

97 MATHEMATICS AND COMPUTING↗

Deep Koopman Neural Network for Analyzing High-Energy-Density Simulations of Electrical Wire Explosions

Megaampere-scale electrical wire experiments (EWEs) provide a platform for studying magnetohydrodynamic (MHD) instability growth in magneto-inertial fusion (MIF) devices. Even when nonlinear simulations of these experiments can digitally reproduce much of the experimentally observed instability growth, interpreting the results and understanding mode growth and evolution can be non-trivial. As a first step toward providing better interpretation of these simulation features, this work investigates the use of a deep neural network that uses Koopman operator theory to analyze the dynamics of pulsed-power-driven explosions of EWEs. This deep neural network is trained on 1-D resistive MHD simulations of EWEs. This neural network learns to transform the nonlinear data into a lower-dimensional representation where the time dynamics are linear. Layers of this neural network are shown to learn features of the simulations, including the locations of shock waves and different physical regimes of the simulation. Using the learned features, the network can compress a time state of the simulation consisting of 5120 data point into a 36-parameter lower-dimensional latent space embedding. Furthermore, these embeddings are shown to be clustered in the latent space by initial radius and time state.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

Detecting Living-off-the-land Attacks Using K-means And Graph Convolutional Networks

The code ingests Zeek logs derived from network packet captures and goes through data preprocessing before it gets passed into a K-Means model that labels each device as either a client or server. Graph Convolutional Network (GCN) model is used to obtain the embeddings to represent the features in lower dimension. Last, K-means cluster analysis is used to cluster the embeddings for each class.

Quach, Anna [Idaho National Laboratory (INL), Idah↗

An Open-source Llm Enhanced-tool Specialized In Helping Moose Related Problems And Tasks

MOOSEenger is an open-source, terminal-first chat application for the MOOSE ecosystem that couples specialized parsing of MOOSE documentation and “.i” input files with retrieval-augmented generation to deliver grounded answers about multiphysics modeling and workflows. It includes dedicated readers for MOOSE-style HTML and a pyhit-based parser that uses the MOOSE syntax tree to preserve block structure and attach retrieval metadata. A data-ingestion pipeline performs semantic chunking into atomic facts and stores them hierarchically in a local Chroma vector database that maintains parent–child relationships across documents; the system can ingest directories, individual files, and single-page web content, and it provides CRUD operations (insert, update, delete) to manage the corpus. At query time, relevant chunks are embedded, retrieved, and fused into the model context, with interactive features such as token streaming, persistent chat history, and dynamic RAG (retrieval triggered by user input or intermediate model output). Deployment is flexible: MOOSEenger runs with local Ollama models or remote Hugging Face/OpenAI backends—typically coordinating generation, lightweight tagging/summarization, and embeddings across three models—and it also supports a server mode and integration with the VS Code Continue interface.

Li, Mengnan [Idaho National Laboratory (INL), Idah↗

TEMPEST

This repository solves the problem of driver identification through vehicular and biometric data. Through an embedding-based approach and a novel loss function, we're able to distinguish between different drivers' behaviors. This also provides preprocessing for reproducibility of results.The code preprocesses vehicular data, trains neural networks, and outputs predictions.This code introduces a novel embedding-based neural network with a 91% rank-1 accuracy, as well as all code to reproduce training and results.

Musgrove, Kyle↗

An Improved Convection Parameterization with Detailed Aerosol–Cloud Microphysics for a Global Model

Abstract A new microphysical treatment that includes aerosol–cloud interactions and secondary ice production (SIP) mechanisms is implemented in the convection scheme of the Community Atmosphere Model, version 6 (CAM6). The approach is to embed a 1D Lagrangian parcel model in the bulk convective plume of the existing deep convection parameterization. Aerosol activation, growth processes including collision/coalescence, and three processes of SIP mechanisms, two of which are normally overlooked in atmospheric models, are represented in this embedded parcel model. These microphysical processes are treated with a hybrid bin/bulk scheme and a high spatial and temporal resolution for the integration of the embedded parcel in 1D, allowing vertical velocity to determine the microphysical evolution following the in-cloud motion during ascent. Simulations of an observed case (Midlatitude Continental Convective Clouds Experiment) of a mesoscale convective system in Oklahoma, United States, with a single-column model (SCAM) version of CAM, are compared with aircraft in situ and ground-based observations of microphysical properties from the convection and precipitation. Results from the validation show the new microphysical scheme has a good representation of the ice initiation in the bulk convective plume, including the known and empirically quantified pathways of primary and secondary initiation, with benefits for the accuracy of properties of its supercooled cloud liquid. The sensitivity simulations and use of tagging tracers for the validated simulation confirm that the newly included SIP mechanisms are of paramount importance for convective microphysics and can be successfully treated in the global model.

54 ENVIRONMENTAL SCIENCES↗

Subsurface Spectroscopy of Thermal Degradation Inside an Inert Plastic Bonded Explosive (PBX) Simulant Using Feedback-Assisted Wavefront Shaping

We characterize the subsurface thermal degradation of an inert analog of high-explosive molecular crystals (Eu:Y(acac) 3 (DPEPO)) (EYAD) embedded inside of a plastic bonded explosive simulant using feedback-assisted wavefront shaping-based fluorescence and Raman spectroscopies. This technique utilizes wavefront shaping to focus pump light inside a heterogeneous material onto a target particle, which significantly improves its spectroscopic signature. We find that embedding the EYAD crystals in the heterogeneous polymer results in improved thermal stability, relative to bare crystal measurements, with the crystal remaining fluorescent to >612 K inside of the heterogeneous material, while the bare crystal’s fluorescence is fully quenched by 500 K. We hypothesize that this improvement is due to the polymer restricting the effects of EYAD melting, which occurs at 400 K and is the primary mechanism for spectroscopic changes in the temperature range explored.

Anderson, Benjamin R.↗

Finite deformation implementation of a mixed-mode single-integral type cohesive zone with reorienting surfaces of separation

To model material ductile failure and crack propagation, cohesive zone elements can be embedded along potential fracture paths in a finite element simulation. When damage criteria are met, elements in the mesh decohere, simulating the formation and propagation of a crack. In this paper, we present a novel computational algorithm based on finite deformation theory, essential to modeling crack initiation and growth in solids undergoing large deformations. This new algorithm was formulated within a Lagrangian frame of reference to extend previous cohesive zone algorithms to include modeling crack growth in finite deformation contexts. The local coordinate system, necessary for defining an embedded cohesive zone, is constructed based upon the current configuration and is updated within the nonlinear iteration process, thereby resulting in the convergence of the solution for a growing crack in a large deformation quasi-static setting. The model’s accuracy was demonstrated by comparing finite element model simulation results with the analytic case of a constant surface separation, as shown in the verification examples. The power and efficacy of the algorithm to capture large deformations during crack growth were then demonstrated with a double cantilever beam example case. It indicates that the model can be applied to a variety of physical circumstances for predicting crack initiation and growth with delamination and fracture.

42 ENGINEERING↗

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Design of digital acquisition for beam current monitor

As a part of the Proton Improvement Plan – II (PIP-II) at Fermilab, instrumentation systems are being modernized to take advantage of the higher speeds and ease of use offered by standardized embedded systems like MicroTCA. A rear-transition module (RTM) is being designed to interface with said embedded systems. In each of the four identical channels on the RTM, the differential signal from an alternating-current current transformer (ACCT) transimpedance amplifier will again be amplified by a differential operation-amplifier, then filtered by a low-pass topology. The conditioned signal is then digitized at a maximum of 10MS/s by an analog to digital converter (ADC) integrated circuit. After digitization, the ADC passes the data to an off the shelf AdvancedMC (AMC) Xilinx FPGA module using low voltage differential signals. This paper will describe the simulation of analog circuitry for signal conditioning, simulation of digital signal integrity based on physical design as well as verification of design characteristics critical to signal integrity. This work aims to create a methodology that can be applied to future RTMs requiring application of high-speed digital design principles.

White, R.Turner [Fermilab]↗

Nonsteady Load Responses of Wind Turbines to Atmospheric and Mountain-Generated Turbulence Eddies, With Impacts on the Main Bearing: A Validation Study

Previous computational and field experiments identify three characteristic time scales in the aerodynamic responses of utility-scale wind turbine loads to atmospheric boundary layer (ABL) turbulence: a 30-90 second time scale for the passage of high/low speed "streaks" through the rotor plane, the blade and rotor rotation time scales (approximately 1 to 5 seconds), and a sub-second time scale created by blade rotation through gradients within eddy coherent structure. In the current study we compare aerodynamic load responses from daytime ABL turbulence quantified with large-eddy simulation and a actuator line model of the NREL 5 MW wind turbine with analysis of field data from the NREL/GE 1.5 MW wind turbine 5 kilometers east of the Rocky Mountain Front Range in Colorado. In addition, we contrast the responses to the passage of the mountain-generated eddies embedded within the westerly winds with the ABL eddies embedded within northerly/southerly winds. These analyses are in context with the nonsteady forcing of the main bearing by the aerodynamic generation of nontorque bending moments on the main shaft. Potentially relevant to main bearing failure mechanisms, both computational and field data show that the magnitudes of turbulence-generated nontorque bending moments, that we show generate nonsteady force on the main bearing, are of order, and often larger than, torque (which underlies power). However, the temporal variations in these two responses are uncorrelated, implying that the aerodynamic mechanisms that drive power and main bearing response are fundamentally different. We find this to be the case in the field with both mountain-generated eddies (westerly winds) and ABL-generated eddies (northerly/southerly winds). Whereas the time and length scales are comparable, the mountain eddies were somewhat more energetic than the northerly/southerly ABL eddies. Interestingly, however, the fluctuations in nontorque bending moment that force the main bearing were found to be stronger when forced by the ABL eddies than the mountain eddies. The field studies validate the key results from the computational study and show even stronger response in the nontorque bending moment than in the computer simulations. In all cases, the torque and nontorque bending moments are temporally uncorrelated, torque and power are driven by time variations in rotor-averaged horizontal wind velocity and nontorque bending moments are driven by time changes in the degree of nonuniformity in the distribution of velocity over the rotor plane. Thus the results generalize the mechanisms underlying nonsteady aerodynamic forcing to classes of turbulence eddy types with strength of order or stronger than ABL eddies with transverse scale of order the wind turbine rotor. These include atmospheric turbulence eddies, topography-generated turbulence eddies and, by extension, impacts of turbine-wake-scale turbulence eddies on downstream wind turbine rotors.

17 WIND ENERGY↗

Internship Work Report

I worked on two projects during my summer internship at Sandia. My official title was “Intern - Mission Tech Electrical Eng./Computer Eng.- R&D Undergraduate Summer.” I worked at the central location, which is Albuquerque, New Mexico. The department you are placed in at Sandia doesn’t always correspond to the people you will be working with. For example, I only directly worked with one person from my department this summer. On one of my projects, I worked with a diverse team of engineers from many different departments. On my other project, I mainly worked with two departments, as the project had two distinct parts. As mentioned earlier, I worked on two projects during my summer at Sandia. The first project focused on a lightweight embedded controller in an advanced FPGA System-on-Chip for radar signal processing applications. The term “controller” refers to a hardware device that directs the flow of data between two entities. An FPGA is a reprogrammable integrated circuit (as opposed to an integrated circuit with one purpose). An FPGA was used on this project so in order to protype various ideas for our System-on-Chip. My role on the project was to implement designs on the fabric of the FPGA and design a state machine (written in C) for the processor. My second project was also heavily involved with embedded systems but had a different application. It focused on using a Newton-Raphson control algorithm to stabilize an inverted pendulum using a novel microcontroller. The pendulum dynamics were derived, and it was successfully simulated in MATLAB. I worked on integrating the microcontroller with the inverted pendulum machinery, and converting the Newton-Raphson control algorithm from MATLAB into C. The inverted pendulum was successfully stabilized using a simple PID controller and industry-standard microcontroller. The project is still ongoing, and the team is gearing up for more tests using the novel Newton-Raphson control algorithm and novel microcontroller

42 ENGINEERING↗

Evaluating Iodine Immobilization Technologies: Cermets, Polycermets, and Polyhalmets

The work in this report documents the efforts conducted to assess the feasibility of some of the ideas documented in Pacific Northwest National Laboratory invention disclosure reports (IDRs) including: 1) Iodine capture in polyacrylonitrile (PAN)-containing composite sorbents (32451-E). In this work, the composites evaluated included Ag0, Bi0, Cu0, Bi2S3, and Cu2S embedded in PAN. 2) Metal iodide removal from these sorbents through dissolution in dimethyl sulfoxide (DMSO) (32729-E). In this work, PAN dissolution was evaluated for multiple types of sorbents including Ag-Pan, Bi-PAN, Cu-PAN, Bi2S3-PAN, and Cu2S-PAN. 3) Using metal-sulfide sorbents for iodine capture (32647-E). In this work, the composites evaluated under this IDR included Ag2S, Bi2S3, and Cu2S embedded in PAN. 4) Using low-melting metals to immobilize (encapsulate) iodine-loaded and polymer-containing sorbents into polymer-ceramic-metal (called polycermet) or polymer-halide-metal (called polyhalmet) composite waste forms (32625-E). In this work, the iodine-loaded PAN composites included AgI-PAN, BiI-PAN, and CuI-PAN. 5) Ceramic-metal composite waste form synthesis of polymer-containing materials using low-melting metals like bismuth, tin, or bismuth-tin alloys (32537-E). In this work, the metals evaluated included Bi, 58Bi-42Sn eutectic. 6) Cermets for immobilizing commercial sorbents loaded with radioiodine (32806-E). In this work, AgIX (iodine-loaded silver faujasite zeolite) was evaluated in cermet form.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗

Facial Named Entity Recognition by Attention-Based Graph Convolutional Neural Network

In the realm of facial recognition and analysis, the ability to accurately cluster large datasets of facial images stands as a cornerstone for various applications, ranging from security surveillance to user biometric identification. This project evolves a novel approach to facial data clustering by embedding facial images into a high-dimensional vector space using an advanced embedding model trained on separate data and assumes a graph-like structure on the high-dimensional vectors. We find our method works significantly better than common shallow methods.

97 MATHEMATICS AND COMPUTING↗

Non-destructive structural characterization of graphite components using mechanical resonance and deep learning

As compared to conventional nuclear reactors, microreactors have the potential to significantly reduce construction timelines and capital costs, decreasing the barriers for advanced nuclear reactor technologies. However, the lower power output of these microreactors (typically < 20 MWe) creates challenging economics if operation and maintenance costs cannot be sufficiently reduced. The compact size of these designs presents an opportunity for comprehensive in-situ structural health monitoring to provide real-time feedback in order to reduce operational costs associated with maintenance and downtime. Many microreactor concepts use graphite for both in-core neutron moderation and as a structural material, which has typically required some form of periodic and laborious inspection. This report provides a description and assessment of recent work with graphite to couple acoustic-based experimental measurements and characterization with machine learning models to mature structural health monitoring capabilities and generate benefits for the nuclear microreactor industry. With resilient embedded sensors in development in other programs funded by the US Department of Energy’s Office of Nuclear Energy and elsewhere, the work described herein builds upon previously funded efforts to mature non-destructive testing technology that relates measured vibrational signatures to structural changes, using a combination of new experimental measurements and machine learning processing. Building on past successful demonstrations of predictive workflows to identify structural changes in a hexagonal stainless steel test article with excellent acoustic propagation, we first performed baseline characterization on graphite samples with canonical geometries to ensure compatibility and confidence in the applied techniques for a material with distinctly different mechanical properties. In contrast to efforts in previous years, we worked exclusively with unidirectional vibration data that is more comparable to those expected from the existing embedded sensor technologies which are suitable for deployment in a reactor setting. Established acoustic and modern machine-learning-based characterization approaches were applied to the resulting datasets from these simple geometries. Both approaches were found to be highly capable of detecting even small geometric irregularities amongst nominally identical samples. As such, we then moved to testing these approaches for detection of artificial local stress perturbations introduced into a more complex geometry: a hexagonal block with drilled holes. A main outcome of this work is that a generalizable ML workflow can be used to detect and predict the characteristics of small artificial anomalies in a graphite component with a relevant geometry. While this work was performed using surficial vibration data, we expect the approach to be flexible and viable for other monitoring scenarios, such as those with different arrangements or types of sensor arrays. As compared to previously funded efforts, an existing ML workflow based on neural networks was enhanced through the addition of recently developed Fourier neural operators. As applied to previously collected and new vibration datasets, prediction accuracies of anomaly characterizations were greatly improved with minimal added computational cost. As trained on small durations of vibration data (tens of seconds) collected over a realistic number of locations, the model was able to reliably determine the presence of a subtle stress anomaly and begin to provide location estimates. Such an approach is likely to be viable for more relevant reactor damage scenarios for graphite components, such as progressive crack growth or creep.

36 MATERIALS SCIENCE↗