Search NASA⌕ Search

SEARCH · Search NASA

Results for “Base Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

A Preferences Corpus and Annotation Scheme for Human-Guided Alignment of Time-Series GPTs

The process of time-series forecasting such as predicting trajectories of silicon content in blast furnaces is a difficult task. Most time-series approaches today focus on scalar-type MSE loss optimization. This optimization approach, while widely common, could benefit from the use of human expert or process-level preferences. In this paper, we introduce a novel alignment and fine-tuning approach that involves learning from a corpus of preferred and dis-preferred time-series prediction trajectories. Our contributions include (1) a preference annotation pipeline for time-series forecasts, (2) the application of Score-based Preference Optimization (SPO) to train decoder-only transformers from preferences, and (3) results showing improvements in forecast quality. The approach is validated on both proprietary blast furnace data and the UCI Appliances Energy dataset. The proposed preference corpus and training strategy offer a new option for fine-tuning sequence models in industrial settings.

DPO↗

Using Calibrated Sodium Data for Preliminary Validation of the SRT Code for Advanced Reactors

Various types of non-light water reactors are currently engaged in the U.S. licensing process. Because of inherent differences compared with well-established large light water reactors, appropriate assessment tools are needed. Specifically, source term analysis, which determines environmental dose impacts from potential accident scenarios, is a crucial part of design and licensing. The U.S. Nuclear Regulatory Commission has emphasized the importance of mechanistic source term analysis for advanced reactor deployments. To align with these needs, Argonne National Laboratory has developed the Simplified Radionuclide Transport (SRT) source term analysis code for metal fuel Sodium-cooled Fast Reactors (SFRs) and microreactors. SRT conducts time-dependent radionuclide transport and retention in SFRs for core and ex-core radionuclide source accident sequences. The main objective of SRT is to provide rapid sensitivity and uncertainty analyses, incorporating parametric uncertainties and summarizing probabilistic results. As part of the code validation process, a study focused on the bubble scrubbing module was performed using an experiment recently carried out by the University of Wisconsin-Madison. Based on the analysis, the modeling approach in SRT provides accurate results for small and large aerosols, while slight underprediction of radionuclide aerosol removal are observed for medium sized aerosols. However, the deviation is minor, considering the highly uncertain phenomenon and range of results, and is in the conservative direction. In addition, uncertainty information derived from the experiments is further implemented, reflecting the actual span of parameters, which leads to enhanced agreement with code predictions. The results demonstrate that SRT provides reasonable predictions for the bubble scrubbing process in sodium pool.

Kam, Dong Hoon↗

Overview of the SCEC/USGS Community Stress Drop Validation Study Using the 2019 Ridgecrest Earthquake Sequence

We present initial findings from the ongoing Community Stress Drop Validation Study to compare spectral stress-drop estimates for earthquakes in the 2019 Ridgecrest, California, sequence. This study uses a unified dataset to independently estimate earthquake source parameters through various methods. Stress drop, which denotes the change in average shear stress along a fault during earthquake rupture, is a critical parameter in earthquake science, impacting ground motion, rupture simulation, and source physics. Spectral stress drop is commonly derived by fitting the amplitude-spectrum shape, but estimates can vary substantially across studies for individual earthquakes. Sponsored jointly by the U.S. Geological Survey and the Statewide (previously, Southern) California Earthquake Center our community study aims to elucidate sources of variability and uncertainty in earthquake spectral stress-drop estimates through quantitative comparison of submitted results from independent analyses. The dataset includes nearly 13,000 earthquakes ranging from M 1 to 7 during a two-week period of the 2019 Ridgecrest sequence, recorded within a 1° radius. Here, in this article, we report on 56 unique submissions received from 20 different groups, detailing spectral corner frequencies (or source durations), moment magnitudes, and estimated spectral stress drops. Methods employed encompass spectral ratio analysis, spectral decomposition and inversion, finite-fault modeling, ground-motion-based approaches, and combined methods. Initial analysis reveals significant scatter across submitted spectral stress drops spanning over six orders of magnitude. However, we can identify between-method trends and offsets within the data to mitigate this variability. Averaging submissions for a prioritized subset of 56 events shows reduced variability of spectral stress drop, indicating overall consistency in recovered spectral stress-drop values.

58 GEOSCIENCES↗

Nickel Binding to the c-Src SH3 Domain Facilitates Crystallization

Introduction: Numerous X-ray crystal structures of the c-Src SH3 domain have provideda large sampling of atomic-level information for this important signaling domain. Multiple crystalforms have been reported, with variable crystal lattice contacts and chemical crystallizationconditions. Materials and Methods: We crystallized the c-Src SH3 domain in a crystallization buffercontaining NiCl2. Results: A unique crystal structure of the Src SH3 domain in the trigonal space group H32 isdetermined to 1.45 Å resolution. Crystal packing and anomalous scattering reveal that this crystalform is mediated by two ordered nickel ions provided by the crystallization buffer. Nickelcoordination occurs in a 2:2 stoichiometry, which dimerizes two SH3 domain monomers across apseudo-twofold rotation axis and involves the native N-terminal c-Src SH3 amino acid sequence, asurface-exposed histidine residue, and ordered water molecules. Discussion: This study provides an example of metal-mediated crystallization and metal binding byN-terminal protein residues, contrasting with the Amino-Terminal Copper and Nickel Binding(ATCUN) motif. Conclusion: Alternative avenues help widen the potential for future crystallography-based studiesof the c-Src SH3 domain.

Biochemistry & Molecular Biology↗

SAGAbg. II. The Low-mass Star-forming Sequence Evolves Significantly between 0.05 < z < 0.21

The redshift-dependent relation between galaxy stellar mass and star formation rate (SFR), known as the star-forming sequence (SFS), is a key observational yardstick for galaxy assembly. We use the SAGAbg-A sample of background galaxies from the Satellites Around Galactic Analogs (SAGA) Survey to model the low-redshift evolution of the low-mass SFS. The sample is comprised of 23,258 galaxies with Hα-based SFRs spanning 6 < log 10 (M * /[M ⊙ ]) < 10 and z < 0.21 (t < 2.5 Gyr). Although it is common to bin or stack galaxies at z ≲ 0.2 for galaxy population studies, the difference in lookback time between z = 0 and z = 0.21 is comparable to the time between z = 1 and z = 2. We develop a model to account for both the physical evolution of low-mass SFS and the selection function of the SAGA Survey, allowing us to disentangle redshift evolution from redshift-dependent selection effects across the SAGAbg-A redshift range. Our findings indicate significant evolution in the SFS over the last ∼2.5 Gyr, with a rising normalization: $\langle$SRF(M * = 10 8.5 M ⊙ )$\rangle$ (z) = 1.24$^{+0.25}_{–0.23}$z – 1.47$^{+0.03}_{–0.03}$. We also identify the redshift limit at which a static SFS is ruled out at the 95% confidence level, which is z = 0.05 based on the precision of the SAGAbg-A sample. Comparison with cosmological hydrodynamic simulations reveals that some contemporary simulations underpredict the recent evolution of the low-mass SFS. This demonstrates that the recent evolution of the low-mass SFS can provide new constraints on the assembly of the low-mass Universe and highlights the need for improved models in this regime.

79 ASTRONOMY AND ASTROPHYSICS↗

Web Based Beamline Control System (Bluesky Web) v0.1.0

Bluesky Web is a web based interface that provides beam line controls to the end user. It allows users to issue commands to various physical devices at a beam line end station like motors and cameras. It utilizes an open source Python library (Bluesky) as the controller. It uses Bluesky to also allow for running "plans" or a sequence of device operations that can be used when running an experiment. This program is different from other controls technologies because it is intended to be open source and can be accessed from a web browser, as opposed to other paid software that is run as a stand-alone application on a computer.

De Leon, Seij↗

Automated Direct Perturbation Calculations with SCALE TSUNAMI [Abstract]

In nuclear criticality safety analysis, the sensitivity of the eigenvalue keff to uncertainties in nuclear data and its evaluation are crucial. The TSUNAMI sequences within the SCALE code system offer users various options with both multigroup (MG) and continuous-energy (CE) 3D Monte Carlo (MC) transport capabilities for calculating keff sensitivity coefficients and storing them in a sensitivity data file (SDF). Each methodology available in TSUNAMI offers distinct advantages and limitations, and its effectiveness can vary based on the specific problem being solved. As a best practice, practitioners typically use the direct perturbation (DP) method as a confirmatory step alongside their sensitivity calculations to verify the accuracy of the sensitivity data generated. In this process, DP calculations are usually performed on select nuclides, those considered most important for validating their total sensitivities. However, because of code limitations, analysts use a workaround method when conducting DP calculations for a single nuclide: rather than perturbing the nuclide's microscopic cross section, an equivalent number density for this nuclide is calculated to reflect the effect of a change in the macroscopic cross section due to a perturbation in the microscopic cross section. The current approach requires rerunning the CSAS criticality calculation several times with model changes. Although this method can yield results with acceptable accuracy, it is labor-intensive and prone to errors.

AZURE: SAMMY↗

Unlocking saponin biosynthesis in soapwort

Abstract Soapwort ( Saponaria officinalis ) is a flowering plant from the Caryophyllaceae family with a long history of human use as a traditional source of soap. Its detergent properties are because of the production of polar compounds (saponins), of which the oleanane-based triterpenoid saponins, saponariosides A and B, are the major components. Soapwort saponins have anticancer properties and are also of interest as endosomal escape enhancers for targeted tumor therapies. Intriguingly, these saponins share common structural features with the vaccine adjuvant QS-21 and, thus, represent a potential alternative supply of saponin adjuvant precursors. Here, we sequence the S . officinalis genome and, through genome mining and combinatorial expression, identify 14 enzymes that complete the biosynthetic pathway to saponarioside B. These enzymes include a noncanonical cytosolic GH1 (glycoside hydrolase family 1) transglycosidase required for the addition of d- quinovose. Our results open avenues for accessing and engineering natural and new-to-nature pharmaceuticals, drug delivery agents and potential immunostimulants.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improved deep learning prediction of antigen–antibody interactions

Identifying antibodies that neutralize specific antigens is crucial for developing effective immunotherapies, but this task remains challenging for many target antigens. The rise of deep learning–based computational approaches presents a promising avenue to address this challenge. Here, we assess the performance of a deep learning approach through two benchmark tests aimed at predicting antibodies for the receptor-binding domain of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike protein. Three different strategies for constructing input sequence alignments are employed for predicting structural models of antigen–antibody complexes. In our initial testing set, which comprises known experimental structures, these strategies collectively yield a significant top-ranked prediction for 61% of cases and a success rate of 47%. Notably, one strategy that utilizes the sequences of known antigen binders outperforms the other two, achieving a precision of 90% in a subsequent test set of ~1,000 antibodies, balanced between true and control antibodies for the antigen, albeit with a lower recall of 25%. Our results underscore the potential of integrating deep learning methods with single B cell sequencing techniques to enhance the prediction accuracy of antigen–antibody interactions.

Science & Technology - Other Topics↗

Dissecting neurofilament tail sequence-phosphorylation-structure relationships with multicomponent reconstituted protein brushes

Neurofilaments (NFs) are multisubunit, bottlebrush-shaped intermediate filaments abundant in the axonal cytoskeleton. Each NF subunit contains a long intrinsically disordered tail domain, which protrudes from the NF core to form a “brush” surrounding each NF. Precisely how the tails’ variable charge patterns and repetitive phosphorylation sites mediate their conformation within the brush remains an open question in axonal biology. We address this problem by grafting recombinant NF tail protein constructs NF-Light, -Medium, and -Heavy (NFL, NFM, and NFH) to surfaces, yielding protein brushes of defined stoichiometry that can be phosphorylated in vitro. Atomic force microscopy measurements reveal that brush height depends on composition monotonically but not always linearly for binary NFL:NFM or NFL:NFH systems, and that NFM-based brushes are highly extended, while brushes incorporating the much larger NFH are surprisingly compact even after multisite phosphorylation. Complementary self-consistent field theory (SCFT) predicts multilayer brush morphologies for NFM and phosphorylated NFH brushes. Further experiments and SCFT analysis with designed mutants reveal that N-terminal negative charges in the NFH tail repel phosphorylated residues to generate the multilayer morphology, while the C-terminal charge-neutral region contributes to multilayer brush morphology but not total brush height. Charge-shuffled NFM variants show that charge segregation promotes brush collapse near physiological ionic strengths. Collectively, this study supports a role for NFM in establishing a dynamic range for NF brush conformation, lending insight into previous in vitro and in vivo findings. More broadly, this work establishes a platform for dissecting contributions of disordered protein sequence to conformation at interfaces.

Science & Technology - Other Topics↗

Genesis: A Compiler Framework for Hamiltonian Simulation on Hybrid CV-DV Quantum Computers

We introduce Genesis, the first compiler designed to support Hamiltonian Simulation on hybrid continuous-variable (CV) and discrete-variable (DV) quantum computing systems. Genesis is a two-level compilation system. At the first level, it decomposes an input Hamiltonian into basis gates using the native instruction set of the target hybrid CV-DV quantum computer. At the second level, it tackles the mapping and routing of qumodes/qubits to implement long-range interactions for the gates decomposed from the first level. Rather than a typical implementation that relies on SWAP primitives similar to qubit-based (or DV-only) systems, we propose an integrated design of connectivity-aware gate synthesis and beamsplitter SWAP insertion tailored for hybrid CV-DV systems. We also introduce an OpenQASM-like domain-specific language (DSL) named CVDV-QASM to represent Hamiltonian in terms of Pauli-exponentials and basic gate sequences from the hybrid CVDV gate set. Genesis has successfully compiled several important Hamiltonians, including the Bose-Hubbard model, Z2−Higgs model, Hubbard-Holstein model, Heisenberg model and Electron-vibration coupling Hamiltonians, which are critical in domains like quantum field theory, condensed matter physics, and quantum chemistry. Our implementation is available at Genesis-CVDV-Compiler https://github.com/ruadapt/Genesis-CVDV-Compiler

Chen, Henry↗

Flow matching meets biology and life science: a survey

Over the past decade, advances in generative modeling, such as generative adversarial networks, masked autoencoders, and diffusion models, have significantly transformed biological research and discovery, enabling breakthroughs in molecule design, protein generation, catalysis discovery, drug discovery, and beyond. At the same time, biological applications have served as valuable testbeds for evaluating the capabilities of generative models. Recently, flow matching has emerged as a powerful and efficient alternative to diffusion-based generative modeling, with growing interest in its application to problems in biology and life sciences. This paper presents the first comprehensive survey of recent developments in flow matching and its applications in biological domains. We begin by systematically reviewing the foundations and variants of flow matching, and then categorize its applications into three major areas: biological sequence modeling, molecule generation and design, and peptide and protein generation. For each, we provide an in-depth review of recent progress. We also summarize commonly used datasets and software tools, and conclude with a discussion of potential future directions.

59 BASIC BIOLOGICAL SCIENCES↗

Computationally designed coiled coil ‘bundlemers’ as model colloidal nanoparticles for solution assembly and materials design (Final Report)

As a collaborative team at the University of Delaware and the University of Pennsylvania, Kloxin, Pochan and Saven designed new biomimetic nanomaterials de novo, leveraging a variety of complementary areas of expertise: computational design of biopolymers (Saven at the University of Pennsylvania), and synthesis and characterization (Kloxin and Pochan at the University of Delaware). Overall activities included: sequence-specific peptide synthesis; covalent crosslinking; noncovalent assembly; site-specific functionalization; and nanostructural characterization using electron microscopy and solution-phase (x-ray and neutron) scattering. Using natural and non-natural amino acids, the team created modular, functional peptide building blocks for elaboration of new nanostructured materials. Ultimately, the development of robust peptide-based, building blocks provides tools for researchers to readily produce complex nanomaterial structures in a wide range of applications. The project had three, interconnecting goals in an effort to provide the broader scientific community with a new peptide-based paradigm for materials design and characterization. First, we further developed the coiled-coil bundle-based toolbox (otherwise known as the ‘bundlemer’ toolbox) via computational design with experimental bundle assembly verification. Second, we developed new uses of covalent interactions, in addition to desired physical (noncovalent) interactions, to assemble bundlemers into 1-D polymer chains with targeted chain rigidity, length, and dispersity. Thirds, we used the above designs to experimentally realize (physical or covalent) polymers to target the creation of liquid crystals or to realize interparticle assembly into nanoporous lattices. The close integration of the three groups was instrumental in success of the biomolecular materials design, formation, and understanding for future designs.

36 MATERIALS SCIENCE↗

Microbial secondary metabolites: advancements to accelerate discovery towards application

Microbial secondary metabolites not only have key roles in microbial processes and relationships but are also valued in various sectors of today’s economy, especially in human health and agriculture. The advent of genome sequencing has revealed a previously untapped reservoir of biosynthetic capacity for secondary metabolites indicating that there are new biochemistries, roles and applications of these molecules to be discovered. New predictive tools for biosynthetic gene clusters (BGCs) and their associated pathways have provided insights into this new diversity. Advanced molecular and synthetic biology tools and workflows including cell-based and cell-free expression facilitate the study of previously uncharacterized BGCs, accelerating the discovery of new metabolites and broadening our understanding of biosynthetic enzymology and the regulation of BGCs. These are complemented by new developments in metabolite detection and identification technologies, all of which are important for unlocking new chemistries that are encoded by BGCs. This renaissance of secondary metabolite research and development is catalysing toolbox development to power the bioeconomy.

Dinglasan, Jaime Lorenzo N↗

Meld: A project for exploring how to meet DUNE's framework needs

Existing data-processing frameworks for HEP experiments are largely based on collider-physics concepts, which may be based on rigid, event-based data hierarchies. These data organizations are not always helpful for neutrino experiments, which must sometimes work around such restrictions by manually splitting apart events into constructs that are better suited for neutrino physics. The purpose of Meld is to explore more flexible data organizations by treating a frameworks job as: (1) A graph of data-product sequences connected by (2) User-defined functions that serve as operations to (3) Framework-provided higher-order functions.

Knoepfel, KyleJ. [Fermi National Accelerator Labor↗

Sodium azide mutagenesis induces a unique pattern of mutations

The nature and effect of mutations are of fundamental importance to the evolutionary process. The generation of mutations with mutagens has also played important roles in genetics. Applications of mutagens include dissecting the genetic basis of trait variation, inducing desirable traits in crops, and understanding the nature of genetic load. Previous studies of sodium azide-induced mutations have reported single nucleotide variants (SNVs) found in individual genes. To characterize the nature of mutations induced by sodium azide, we analyze whole-genome sequencing (WGS) of 11 barley lines derived from sodium azide mutagenesis, where all lines were selected for diminution of plant fitness owing to induced mutations. We contrast observed mutagen-induced variants with those found in standing variation in WGS of 13 barley landraces. Here, we report indels that are two orders of magnitude more abundant than expected based on nominal mutation rates. We found induced SNVs are very specific, with C → T changes occurring in a context followed by another C on the same strand (or the reverse complement). The codons most affected by the mutagen include the sodium azide-specific CC motif (or the reverse complement), resulting in a handful of amino acid changes and few stop codons. The specific nature of induced mutations suggests that mutagens could be chosen based on experimental goals. Sodium azide would not be ideal for gene knockouts but will create many missense mutations with more subtle effects on protein function.

Genetics & Heredity↗

Replace Human Intelligence with Fast and Smart Geometric Reasoning and Graph Neural Network to Accelerate Next Gen ModSim Workflows

We present an agent-guided approach to CAD geometry decomposition that automates hex/hybrid meshing with graph neural networks (GNNs) to accelerate next-generation ModSim workflows. Our end-to-end pipeline (i) reduces 3D boundary-representation (B-Rep) models to a 2D chordal axis skeleton (CAT) and then to a 1D bipartite graph of surface and curve nodes, (ii) assigns per node labels as Cubit® WebCut actions, (iii) trains a multi-action GNN under supervised learning, and (iv) predicts five surface-node and three curve-node actions on out-of-distribution test geometries. Each graph node carries geometric, topological, and meshing attributes drawn from the B-Rep “skin” and CAT “skeleton,” with two-way mappings across 3D↔2D↔1D representations to maintain traceability back to 3D CAD. The supervised learning model exhibits stable convergence of the binary cross-entropy loss and achieves 98.7% accuracy on unseen lattice models. To operationalize decision-making, we rank predicted commands by geometric significance and prototyped the agent-guided workflow through the Cubit® Meshing PowerTool GUI. As a stretch goal, we explore reinforcement learning (RL) to reduce or remove label requirements and to learn policies for action sequences that maximize total reward (e.g., size of hex-meshable regions and resulting hex mesh quality). When all-hex meshing is not feasible, the agent assists in producing hybrid meshes—prioritizing hex in critical regions and transitioning to tetrahedral elements (tets) elsewhere—maintaining fidelity while ensuring robustness. The overarching objective is to replace manual, heuristics-based decomposition with data-driven, reproducible automation, cutting meshing turnaround time by orders of magnitude. We anticipate direct impact on simulation workflows through intelligent, scalable decomposition of complex CAD models into hex-meshable subdomains.

97 MATHEMATICS AND COMPUTING↗

Closing the Loop between In Situ Stress Complexity and EGS Fracture Complexity

We present an agent-guided approach to CAD geometry decomposition that automates hex/hybrid meshing with graph neural networks (GNNs) to accelerate next-generation ModSim workflows. Our end-to-end pipeline (i) reduces 3D boundary-representation (B-Rep) models to a 2D chordal axis skeleton (CAT) and then to a 1D bipartite graph of surface and curve nodes, (ii) assigns per node labels as Cubit® WebCut actions, (iii) trains a multi-action GNN under supervised learning, and (iv) predicts five surface-node and three curve-node actions on out-of-distribution test geometries. Each graph node carries geometric, topological, and meshing attributes drawn from the B-Rep “skin” and CAT “skeleton,” with two-way mappings across 3D↔2D↔1D representations to maintain traceability back to 3D CAD. The supervised learning model exhibits stable convergence of the binary cross-entropy loss and achieves 98.7% accuracy on unseen lattice models. To operationalize decision-making, we rank predicted commands by geometric significance and prototyped the agent-guided workflow through the Cubit® Meshing PowerTool GUI. As a stretch goal, we explore reinforcement learning (RL) to reduce or remove label requirements and to learn policies for action sequences that maximize total reward (e.g., size of hex-meshable regions and resulting hex mesh quality). When all-hex meshing is not feasible, the agent assists in producing hybrid meshes—prioritizing hex in critical regions and transitioning to tetrahedral elements (tets) elsewhere—maintaining fidelity while ensuring robustness. The overarching objective is to replace manual, heuristics-based decomposition with data-driven, reproducible automation, cutting meshing turnaround time by orders of magnitude. We anticipate direct impact on simulation workflows through intelligent, scalable decomposition of complex CAD models into hex-meshable subdomains.

42 ENGINEERING↗