Search NASA⌕ Search

SEARCH · Search NASA

Results for “embedded”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Hybrid Additive Manufacturing and Electric Field-Assisted Sintering of High-temperature Heat Exchangers via Sacrificial Channel Molds

A hybrid manufacturing approach, integrating additive manufacturing (AM) with powder methodology via electric field-assisted sintering (EFAS), was developed for the fabrication of high-temperature compact heat exchangers (CHX) from refractory metals. The methodology employed additively manufactured sacrificial channel molds (SCMs) as shapeholders for CHX channels, which were embedded in metal powders using EFAS. Following embedding, the SCMs were chemically dissolved to form the internal channel network. SCMs were fabricated using both digital light processing (DLP) and direct ink writing (DIW) from chemically reactive, calcium-based ceramic feedstocks with varying ratios of Al2O3 reinforcement. The microstructure, phase composition, and dissolution behavior of both as-printed and embedded SCMs were investigated. The shrinkage behavior of the SCMs embedded in refractory metals, as well as the interfacial characteristics between the SCMs and metal matrix, were studied. The results showed that the SCMs containing sufficient chemical reactive ceramics dissolved effectively before and after embedding. The as-printed SCMs retained the phase composition of their feedstocks, but the embedded SCMs containing calcium-based ceramics and Al2O3 exhibited the formation of calcium aluminates due to high temperature exposure during embedding. Most SCMs exhibited a cellular Al2O3 network filled with Ca-rich ceramics. Shrinkage after embedding was strongly dependent on SCM density, with lower density SCMs exhibiting greater shrinkage. A thin SCM-affected zone was observed at the metal matrix surface, characterized by increased porosity compared to the bulk matrix. This effect was attributed to infiltration of the SCM materials into powder particle boundaries under pressure, followed by their removal during dissolution. This study demonstrates the feasibility of manufacturing CHXs from hard-to-process refractory metals for use in harsh environments.

36 - MATERIALS SCIENCE↗

Efficient and flexible multirate temporal adaptivity

In this work we present two new families of multirate time step adaptivity controllers, that are designed to work with embedded multirate infinitesimal (MRI) time integration methods for adapting time steps when solving problems with multiple time scales. We compare these controllers against competing approaches on two benchmark problems, showing that the proposed methods offer dramatically improved performance and flexibility. The combination of embedded MRI methods and the proposed controllers enable adaptive simulations of problems with a potentially arbitrary number of time scales, achieving high accuracy while maintaining low computational cost. Additionally, we introduce a new set of embeddings for the family of explicit multirate exponential Runge–Kutta (MERK) methods of orders 2 through 5, resulting in the first-ever fifth-order embedded MRI method. Finally, we compare the performance of a wide range of embedded MRI methods on our benchmark problems to provide guidance on how to select an appropriate MRI method and multirate controller.

97 MATHEMATICS AND COMPUTING↗

Fusion Model for Metagenomics

This work highlights the use of an embeddings approach that can encode multiple features and create efficient contextualization of profiled metagenomes derived from microbiome samples using computer vision models and image representations of the abundance profiles. The model's embeddings can be used to cluster existing samples based on multiple conditions and interpretations, and new embeddings can be quickly created for new samples and fitted to existing clusters to characterize them. This has practical applications for unknown, unlabeled microbiome samples. The model's embeddings can be used to cluster existing samples based on multiple conditions and interpretations, and new embeddings can be quickly created for new samples and fitted to existing clusters to characterize them. This has practical applications for unknown, unlabeled microbiome samples.

Valdes, CamiloA [Lawrence Livermore National Labor↗

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

Strongly Correlated States of Transition Metal Spin Defects: The Case of an Iron Impurity in Aluminum Nitride

We investigate the electronic properties of an exemplar transition metal impurity in an insulator, with the goal of accurately describing strongly correlated defect states. Here, we consider iron in aluminum nitride, a material of interest for hybrid quantum technologies, and we carry out calculations with quantum embedding methods, density matrix embedding theory (DMET) and quantum defect embedding theory (QDET), and with spin-flip time-dependent density functional theory (TDDFT). We show that both DMET and QDET accurately describe the ground state and low-lying excited states of the defect and that TDDFT yields photoluminescence spectra in agreement with experiments. In addition, we provide a detailed discussion of the convergence of our results as a function of the active space used in the embedding methods, thus defining a protocol to obtain converged data directly comparable with experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data‐driven variational method for discrepancy modeling: Dynamics with small‐strain nonlinear elasticity and viscoelasticity

Abstract The effective inclusion of a priori knowledge when embedding known data in physics‐based models of dynamical systems can ensure that the reconstructed model respects physical principles, while simultaneously improving the accuracy of the solution in the previously unseen regions of state space. This paper presents a physics‐constrained data‐driven discrepancy modeling method that variationally embeds known data in the modeling framework. The hierarchical structure of the method yields fine scale variational equations that facilitate the derivation of residuals which are comprised of the first‐principles theory and sensor‐based data from the dynamical system. The embedding of the sensor data via residual terms leads to discrepancy‐informed closure models that yield a method which is driven not only by boundary and initial conditions, but also by measurements that are taken at only a few observation points in the target system. Specifically, the data‐embedding term serves as residual‐based least‐squares loss function, thus retaining variational consistency. Another important relation arises from the interpretation of the stabilization tensor as a kernel function, thereby incorporating a priori knowledge of the problem and adding computational intelligence to the modeling framework. Numerical test cases show that when known data is taken into account, the data driven variational (DDV) method can correctly predict the system response in the presence of several types of discrepancies. Specifically, the damped solution and correct energy time histories are recovered by including known data in the undamped situation. Morlet wavelet analyses reveal that the surrogate problem with embedded data recovers the fundamental frequency band of the target system. The enhanced stability and accuracy of the DDV method is manifested via reconstructed displacement and velocity fields that yield time histories of strain and kinetic energies which match the target systems. The proposed DDV method also serves as a procedure for restoring eigenvalues and eigenvectors of a deficient dynamical system when known data is taken into account, as shown in the numerical test cases presented here.

Masud, Arif↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue Visualization

Domain-specific visualizations sometimes focus on narrow, albeit important, tasks for one group of users. This focus limits the utility of a visualization to other groups working with the same data. While tasks elicited from other groups can present a design pitfall if not disambiguated, they also present a design opportunity—namely, the development of visualizations that support multiple groups. This development choice presents a trade-off of broadening the scope but limiting support for the more narrow tasks of any one group, which in some cases can enhance the overall utility of the visualization. We investigate this scenario through a design study where we develop Guidepost, a notebook-embedded visualization of data that helps scientists assess compute wait times, machine learning researchers understand prediction accuracy, and system maintainers analyze usage trends. We adapt the use of personas for visualization design from existing literature in the HCI and design domains, applying them to categorize tasks based on their uniqueness across stakeholder personas. Under this model, tasks shared between all groups should be supported by interactive visualizations and tasks unique to each group can be deferred to scripting with notebook-embedded visualization design. We evaluate our visualization through real-world case studies and a task-focused evaluation with nine participants. We observe that together, Guidepost's visual encodings, interactions, and export capabilities support the tasks of our differing personas.

97 MATHEMATICS AND COMPUTING↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

DNA-PAINT Imaging with Hydrogel Imprinting and Clearing

Hydrogel-embedding is a versatile technique in fluorescence microscopy, offering stabilization, optical clearing, and the physical expansion of biological specimens. DNA-PAINT is a super-resolution microscopy approach based on the diffusion and transient binding of fluorescently labeled oligos, but its feasibility in hydrogels has not yet been explored. In this study, we demonstrate that polyacrylamide hydrogels support sufficient diffusion for effective DNA-PAINT imaging. Using acrydite-anchored oligonucleotides imprinted from patterned DNA origami nanostructures and microtubule filaments in fixed cells, we find that hydrogel embedding preserves docking strand positioning at the nanoscale. Sample clearing via protease treatment had minor structural effects on the microtubule structure and enhanced diffusion and accessibility to hydrogel-imprinted docking strands. Our work demonstrates promising potential for diffusion and binding-based fluorescence imaging applications in hydrogel-embedded samples.

DNA origami↗

Enhancing the Carbon Monoxide Oxidation Performance through Surface Defect Enrichment of Ceria-Based Supports for Platinum Catalyst

Effective synthesis and application of single-atom catalysts on supports lacking enough defects remain a significant challenge in environmental catalysis. Herein, we present a universal defect-enrichment strategy to increase the surface defects of CeO 2 -based supports through H 2 reduction pretreatment. The Pt catalysts supported by defective CeO 2 -based supports, including CeO 2 , CeZrO x , and CeO 2 /Al 2 O 3 (CA), exhibit much higher Pt dispersion and CO oxidation activity upon reduction activation compared to their counterpart catalysts without defect enrichment. Specifically, Pt is present as embedded single atoms on the CA support with enriched surface defects (CA-HD) based on which the highly active catalyst showing embedded Pt clusters (Pt C ) with the bottom layer of Pt atoms substituting the Ce cations in the CeO 2 surface lattice can be obtained through reduction activation. Embedded PtC can better facilitate CO adsorption and promote O 2 activation at Pt C –CeO 2 interfaces, thereby contributing to the superior low-temperature CO oxidation activity of the Pt/CA-HD catalyst after activation.

36 MATERIALS SCIENCE↗

Hierarchical Conditioning of Diffusion Models Using Tree-of-Life for Studying Species Evolution

A central problem in biology is to understand how organisms evolve and adapt to their environment by acquiring variations in the observable characteristics or traits of species across the tree of life. With the growing availability of large-scale image repositories in biology and recent advances in generative modeling, there is an opportunity to accelerate the discovery of evolutionary traits automatically from images. Toward this goal, we introduce Phylo-Diffusion, a novel framework for conditioning diffusion models with phylogenetic knowledge represented in the form of HIERarchical Embeddings (HIER-Embeds). We also propose two new experiments for perturbing the embedding space of Phylo-Diffusion: trait masking and trait swapping, inspired by counterpart experiments of gene knockout and gene editing/swapping. Our work represents a novel methodological advance in generative modeling to structure the embedding space of diffusion models using tree-based knowledge. Our work also opens a new chapter of research in evolutionary biology by using generative models to visualize evolutionary changes directly from images. We empirically demonstrate the usefulness of Phylo-Diffusion in capturing meaningful trait variations for fishes and birds, revealing novel insights about the biological mechanisms of their evolution. (Model and code can be found at imageomics.github.io/phylo-diffusion)

Khurana, Mridul↗

Constraining Erosion Rates and Landscape Evolution With In Situ 10 Be and 26 Al Cosmogenic Nuclides at Table Mountain, Antarctica

Abstract This study investigates surface weathering and sediment preservation at Table Mountain, a high‐elevation, hyperarid, polar landscape in the Transantarctic Mountains. We report cosmogenic nuclide concentrations ( 10 Be and 26 Al) in quartz from bedrock surfaces, erratic boulder lag, and cobbles embedded within Sirius Group sediments to quantify erosion rates. In situ 10 Be and 26 Al depth profiles from a 2.95 m permafrost core in the Sirius Group further constrain surface erosion rates and elucidate landscape stability. Measured 10 Be and 26 Al concentrations from two sandstone bedrock surfaces adjacent to Sirius Group sediments give erosion rates of 0.18–0.28 m/Myr. An erratic sandstone boulder within the lag above the Sirius Group yields erosion rates of ∼0.42 ± 0.03 m/Myr, whereas two cobbles embedded within the Sirius Group yield higher rates of 0.81–1.12 m/Myr. Depth profiles of in situ 10 Be and 26 Al indicate no vertical mixing of Sirius Group permafrost since deposition. Depth profile models are best explained by erosion rates of 0.53 +0.13 / −0.12 m/Myr, and an exposure age of 0.78 +0.06 / −0.08 Ma. We view the model “age” to represent the ∼0.8‐million‐year time‐scale for surface lowering equivalent to one attenuation length of cosmic ray production to achieve steady‐state conditions. Continual exhumation of embedded clasts from within the Sirius Group results in an accumulation of clasts forming the observed erosional lag deposit covering the landscape. Our erosion rates of the Sirius Group surface based on in situ 10 Be and 26 Al depth profiles are an order‐of‐magnitude larger than those based on meteoric 10 Be infiltration and further clarification is required.

58 GEOSCIENCES↗

Active doping controls the mode of failure in dense colloidal gels

Mechanical properties of disordered materials are governed by their underlying free energy landscape. In contrast to external fields, embedding a small fraction of active particles within a disordered material generates nonequilibrium internal fields, which can help to circumvent kinetic barriers and modulate the free energy landscape. In this work, we investigate through computer simulations how the activity of active particles alters the mechanical response of deeply annealed polydisperse colloidal gels. We show that the “swim force” generated by the embedded active particles is responsible for determining the mode of mechanical failure, i.e., brittle vs. ductile. We find, and theoretically justify, that at a critical swim force the mechanical properties of the gel decrease abruptly, signaling a change in the mode of mechanical failure. The weakening of the elastic modulus above the critical swim force results from the change in gel porosity and distribution of attractive forces among gel particles, while below the critical swim force, the ductility enhancement is caused by an increase of gel structural disorder. Above the critical swim force, the gel develops a pronounced heterogeneous structure characterized by multiple pore spaces, and the mechanical response is controlled by dynamical heterogeneities. We contrast these results with those of a simulated monodisperse gel that exhibits a nonmonotonic trend of ductility modulation with increasing swim force, revealing a complex interplay between the gel energy landscape and embedded activity.

Zhou, Tingtao (ORCID:000000021766719X)↗

Learning nuclear cross sections across the chart of nuclides with graph neural networks

We explore the use of deep learning techniques to learn how nuclear cross sections change as we add or remove protons and neutrons. As a proof of principle, we focus on the neutron-induced reactions in the fast energy regime. Our approach follows a two-stage learning framework. First, we apply representation learning to encode cross section data into a latent space using either variational autoencoders (VAEs) or implicit neural representations (INRs). Then, we train graph neural networks (GNNs) on the resulting embeddings to predict missing values across the nuclear chart by leveraging the topological structure of neighboring isotopes. We demonstrate accurate cross section predictions within a 9 × 9 block of missing nuclei. We also find that the optimal GNN training strategy depends on the type of latent representation used, with VAE embeddings performing best under end-to-end optimization in the original space, while INR embeddings achieve better results when the GNN is trained only in the latent space. Furthermore, using clustering algorithms, we map groups of latent vectors into regions of the nuclear chart and show that VAEs and INRs can discover some of the neutron magic numbers. These findings suggest that deep-learning models based on the representation encoding of cross sections combined with graph neural networks hold significant potential in augmenting nuclear theory models, e.g., by providing reliable estimates of covariances of cross sections, including cross-material covariances.

Machine learning↗

Impacts of Spatial Resolution in a High-Fidelity Capacity Expansion Model: An ERCOT Case Study

Capacity expansion models are important tools in examining the evolution of the electric power sector. Embedded in these tools are many modeling choices with consequential impacts on computational burden and associated analysis. In this study, we adjust the spatial resolution of the Regional Energy Deployment System (ReEDS) to understand the implications of higher-fidelity modeling on energy system projections and model solve times. The native ReEDS regions capture the contiguous United States in 134 balancing areas whereas the regions in the higher-resolution version are defined by over 3,000 U.S. counties. Using both resolutions, we conduct a case study of the Texas Interconnection (The Electric Reliability Council of Texas [ERCOT]) to explore differences in model projections and to inform appropriate applications of high spatial resolution in a large-scale, applied capacity expansion model.

county↗

Position-Enhanced Gradient Attack (PEGA) on Medical Language Models

Federated Learning (FL) enables collaborative training of language models on sensitive clinical notes without sharing the data. However, this paradigm is vulnerable to gradient inversion attacks that can reconstruct private data from shared gradients. We find that state-of-the-art attacks are less effective in the medical domain, failing to overcome the unique challenges posed by its specialized vocabulary and unstructured format. To address this, we introduce the Position-Enhanced Gradient Attack (PEGA), a novel attack that makes gradients position-aware by optimizing token and position embeddings simultaneously. PEGA employs two key innovations: a periodic sorting of positional embeddings to resolve token order ambiguity and a late-stage embedding replacement strategy to correct hard-to-recover critical tokens. To evaluate the leakage of sensitive data more directly, we also propose the Unified PHI-Recall (UPHI), a new metric measuring the recovery of Protected Health Information. Experiments on the MIMIC-III dataset show that PEGA significantly outperforms leading attacks like TAG and LAMP, particularly in its ability to reconstruct identifiable patient information, exposing a more severe and nuanced privacy risk in federated medical NLP.

Xu, Nuo [University of Minnesota]↗

ChemEcho v1.0

ChemEcho is a tool that converts tandem mass spectra into embeddings used to build machine learning (ML) models with fully explainable predictions. It provides an API for transforming raw tandem mass spectral data into embeddings, along with functions for training and validating ML models. Additionally, it includes utilities for retrieving and cleaning training data. ChemEcho is broadly applicable in ML pipelines that use tandem mass spectra for a variety of tasks, such as chemical classification or bioactivity mining. While there are existing methods to generate embeddings from fragmentation data, ChemEcho's approach ensures that predictions remain interpretable, enabling experts to evaluate results and generate hypotheses about the underlying data.

Harwood, Thomas [Lawrence Berkeley National Labora↗