Search NASA⌕ Search

SEARCH · Search NASA

Results for “encoding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Comparative proteomics of a versatile, marine, iron-oxidizing chemolithoautotroph

This study conducted a comparative proteomic analysis to identify potential genetic markers for the biological function of chemolithoautotrophic iron oxidation in the marine bacterium Ghiorsea bivora. To date, this is the only characterized species in the class Zetaproteobacteria that is not an obligate iron-oxidizer, providing a unique opportunity to investigate differential protein expression to identify key genes involved in iron-oxidation at circumneutral pH. Over 1000 proteins were identified under both iron- and hydrogen-oxidizing conditions, with differentially expressed proteins found in both treatments. Notably, a gene cluster upregulated during iron oxidation was identified. This cluster contains genes encoding for cytochromes that share sequence similarity with the known iron-oxidase, Cyc2. Interestingly, these cytochromes, conserved in both Bacteria and Archaea, do not exhibit the typical β-barrel structure of Cyc2. This cluster potentially encodes a biological nanowire-like transmembrane complex containing multiple redox proteins spanning the inner membrane, periplasm, outer membrane, and extracellular space. The upregulation of key genes associated with this complex during iron-oxidizing conditions was confirmed by quantitative reverse transcription-PCR. These findings were further supported by electromicrobiological methods, which demonstrated negative current production by G. bivora in a three-electrode system poised at a cathodic potential. This research provides significant insights into the biological function of chemolithoautotrophic iron oxidation.

59 BASIC BIOLOGICAL SCIENCES↗

A recyclable self-healing composite with advanced sensing property

Polymer-based composites frequently encounter damage, often lurking beneath the surface and proving challenges to their early detection and repair. While material-based sensors show promise for encoding self-sensing properties within these composites, their in situ healing and reprocessability remain significant challenges. Therefore, the overarching goal of this study is the creation of a reprocessable polymeric composite encoded with self-healing attributes and the ability to autonomously sense damage. At the core of this innovation are vitrimers, a polymeric material characterized by a covalently adaptive dynamic network responsive to external factors such as heat. They combine thermoset-like resilience with thermoplastic-like flowability on demand under external stimuli. We nanoengineer a polyester-based vitrimeric polymer by incorporating piezoresistive carbon nanotubes (CNTs) as reinforcing elements that not only enhance its mechanical strength but also create a percolation network within the composite, thereby enabling piezoresistive self-sensing properties, all the while preserving the intrinsic self-healing capabilities offered by the vitrimeric matrix. The fabrication process of the composite involves a solvent-free in situ polymerization method that combines epoxy and anhydride-containing monomers with ~ 0.1 wt.% of CNTs. Once it was established that the introduction of CNTs into the polymeric matrix did not compromise the mechanical properties of the composite, their strain-sensing properties were characterized by applying cyclic loading while measuring their electrical resistance. Strikingly, CNT-enhanced vitrimer composite consistently retains its mechanical and sensing properties through repeated cycles of reshaping and reprocessing, underscoring its potential as a robust distributed strain sensor. This polyester-based vitrimeric composite is also easily recyclable without harsh chemical treatments. Preliminary findings from this study conclusively demonstrate that the bulk composite boasts both self-sensing capabilities and in situ detect healing properties, charting a promising course towards the development of a mechanically resilient multifunctional composite that seamlessly integrates selfsensing and healing capabilities.

Rohewal, Sargun Singh↗

Unlocking the distinctive enzymatic functions of the early plant biomass deconstructive genes in a brown rot fungus by cell-free protein expression

ABSTRACT Saprotrophic fungi that cause brown rot of woody biomass evolved a distinctive mechanism that relies on reactive oxygen species (ROS) to kick-start lignocellulosic polymers’ deconstruction. These ROS agents are generated at incipient decay stages through a series of redox relays that shuttle electrons from fungus’s central metabolism to extracellular Fenton chemistry. A list of genes has been suggested encoding the enzyme catalysts of the redox processes involved in ROS’s function. However, navigating the functions of the encoded enzymes has been challenging due to the lack of a rapid method for protein synthesis. Here, we employed cell-free expression system to synthesize four redox or degradative enzymes, which were identified, by transcriptomic data, as conserved players of the ROS oxidation phase across brown rot fungal species. All four enzymes were successfully expressed and showed activities that enable confident assignment of function, namely, benzoquinone reductase (BQR), ferric reductase, α-L-arabinofuranosidase (ABF), and heme-thiolate peroxidase (HTP). Detailed analysis of their catalytic features within the context of brown rot environments allowed us to interpret their roles during ROS-driven wood decomposition. Specifically, we validated the functions of BQR as the driver redox enzyme of Fenton cycles and reconstructed its interactions with the co-occurring HTP or laccase and ABF. Taken together, this research demonstrated that the cell-free expression platform is adequate for synthesizing functional fungal enzymes and provided an alternative route for the rapid characterization of fungal proteins, escalating our understanding of the distinctive biocatalyst system for plant biomass conversion. IMPORTANCE Brown rot fungi are efficient wood decomposers in nature, and their unique degradative systems harbor untapped catalysts pursued by the biorefinery and bioremediation industries. While the use of “omics” platforms has recently uncovered the key “oxidative-hydrolytic” mechanisms that allow these fungi to attack lignocellulose, individual protein characterization is lagging behind due to the lack of a robust method for rapid synthesis of crucial fungal enzymes. This work delves into the studies of biochemical functions of brown rot enzymes using a rapid, cell-free expression platform, which allowed the successful depictions of enzymes’ catalytic features, their interactions with Fenton chemistry, and their roles played during the incipient stage of brown rot when fungus sets off the reactive oxygen species for oxidative degradation. We expect this research could illuminate cell-free protein expression system’s use to fulfill the increasing need for functional studies of fungal enzymes, advancing the discoveries of novel biomass-converting catalysts.

60 APPLIED LIFE SCIENCES↗

Anaerobic benzene oxidation in Geotalea daltonii involves activation by methylation and is regulated by the transition state regulator AbrB

ABSTRACT Benzene is a widespread groundwater contaminant that persists under anoxic conditions. The aim of this study was to more accurately investigate anaerobic microbial degradation pathways to predict benzene fate and transport. Preliminary genomic analysis of Geotalea daltonii strain FRC-32, isolated from contaminated groundwater, revealed the presence of putative aromatic-degrading genes. G. daltonii was subsequently shown to conserve energy for growth on benzene as the sole electron donor and fumarate or nitrate as the electron acceptor. The hbs gene, encoding for 3-hydroxybenzylsuccinate synthase (Hbs), a homolog of the radical-forming, toluene-activating benzylsuccinate synthase (Bss), was upregulated during benzene oxidation in G. daltonii , while the bss gene was upregulated during toluene oxidation. Addition of benzene to the G. daltonii whole-cell lysate resulted in toluene formation, indicating that methylation of benzene was occurring. Complementation of σ 54 - (deficient) E. coli transformed with the bss operon restored its ability to grow in the presence of toluene, revealing bss to be regulated by σ 54 . Binding sites for σ 70 and the transition state regulator AbrB were identified in the promoter region of the σ 54 -encoding gene rpoN, and binding was confirmed. Induced expression of abrB during benzene and toluene degradation caused G. daltonii cultures to transition to the death phase. Our results suggested that G. daltonii can anaerobically oxidize benzene by methylation, which is regulated by σ 54 and AbrB. Our findings further indicated that the benzene, toluene, and benzoate degradation pathways converge into a single metabolic pathway, representing a uniquely efficient approach to anaerobic aromatic degradation in G. daltonii . IMPORTANCE The contamination of anaerobic subsurface environments including groundwater with toxic aromatic hydrocarbons, specifically benzene, toluene, ethylbenzene, and xylene, has become a global issue. Subsurface groundwater is largely anoxic, and further study is needed to understand the natural attenuation of these compounds. This study elucidated a metabolic pathway utilized by the bacterium Geotalea daltonii capable of anaerobically degrading the recalcitrant molecule benzene using a unique activation mechanism involving methylation. The identification of aromatic-degrading genes and AbrB as a regulator of the anaerobic benzene and toluene degradation pathways provides insights into the mechanisms employed by G. daltonii to modulate metabolic pathways as necessary to thrive in anoxic contaminated groundwater. Our findings contribute to the understanding of novel anaerobic benzene degradation pathways that could potentially be harnessed to develop improved strategies for bioremediation of groundwater contaminants.

Bullows, James E.↗

Host-specific adaptation in Fusarium oxysporum correlates with distinct accessory chromosome content in human and plant pathogenic strains

ABSTRACT Fusarium oxysporumis a cross-kingdom pathogen. While some strains cause disseminated fusariosis and blinding corneal infections in humans, others are responsible for devastating vascular wilt diseases in plants. To better understand the distinct adaptations ofF. oxysporumto animal or plant hosts, we conducted a comparative phenotypic and genetic analysis of two strains: MRL8996 (isolated from a keratitis patient) and Fol4287 (isolated from a wilted tomato [Solanum lycopersicum]). Infection of mouse corneas and tomato plants revealed that, while both strains cause symptoms in both hosts, MRL8996 caused more severe corneal disease in mice, whereas Fol4287 induced more pronounced wilting symptoms in tomato plants.In vitroassays using abiotic stress treatments revealed that the human pathogen MRL8996 was better adapted to elevated temperatures, whereas the plant pathogen Fol4287 was more tolerant to osmotic and cell wall stresses. Both strains displayed broad resistance to antifungal treatment, with MRL8996 exhibiting the paradoxical effect of increased tolerance to higher concentrations of the antifungal caspofungin. We identified a set of accessory chromosomes (ACs) that encode genes with different functions and have distinct transposon profiles between MRL8996 and Fol4287. Interestingly, ACs from both genomes also encode proteins with shared functions, such as chromatin remodeling and post-translational protein modifications. Our phenotypic assays and comparative genomics analyses lay the foundation for future studies correlating genotypes with phenotype and for developing targeted antifungals for agricultural and clinical uses. IMPORTANCE Fusarium oxysporumis a cross-kingdom fungal pathogen that infects both plants and animals. In addition to causing many devastating wilt diseases, this group of organisms was recently recognized by the World Health Organization as a high-priority threat to human health. Climate change has increased the risk ofFusariuminfections, asFusariumstrains are highly adaptable to changing environments. Deciphering fungal adaptation mechanisms is crucial to developing appropriate control strategies. We performed a comparative analysis ofFusariumstrains using an animal (mouse) and plant (tomato) host andin vitroconditions that mimic abiotic stress. We also performed comparative genomics analyses to highlight the genetic differences between human and plant pathogens and correlate their phenotypic and genotypic variations. We uncovered important functional hubs shared by plant and human pathogens, such as chromatin modification, transcriptional regulation, and signal transduction, which could be used to identify novel antifungal targets.

Microbiology↗

Diverse and unconventional methanogens, methanotrophs, and methylotrophs in metagenome-assembled genomes from subsurface sediments of the Slate River floodplain, Crested Butte, CO, USA

We use metagenome-assembled genomes (MAGs) to understand single-carbon (C1) compound-cycling—particularly methane-cycling—microorganisms in montane riparian floodplain sediments. We generated 1,233 MAGs (>50% completeness and <10% contamination) from 50- to 150-cm depth below the sediment surface capturing the transition between oxic, unsaturated sediments and anoxic, saturated sediments in the Slate River (SR) floodplain (Crested Butte, CO, USA). We recovered genomes of putative methanogens, methanotrophs, and methylotrophs (n = 57). Methanogens, found only in deep, anoxic depths at SR, originate from three different clades (Methanoregulaceae, Methanotrichaceae, and Methanomassiliicoccales), each with a different methanogenesis pathway; putative methanotrophic MAGs originate from within the Archaea (Candidatus Methanoperedens) in anoxic depths and uncultured bacteria (Ca. Binatia) in oxic depths. Genomes for canonical aerobic methanotrophs were not recovered. Ca. Methanoperedens were exceptionally abundant (~1,400× coverage, >50% abundance in the MAG library) in one sample that also contained aceticlastic methanogens, indicating a potential C1/methane-cycling hotspot. Ca. Methylomirabilis MAGs from SR encode pathways for methylotrophy but do not harbor methane monooxygenase or nitrogen reduction genes. Comparative genomic analysis supports that one clade within the Ca. Methylomirabilis genus is not methanotrophic. The genetic potential for methylotrophy was widespread, with over 10% and 19% of SR MAGs encoding a methanol dehydrogenase or substrate-specific methyltransferase, respectively. MAGs from uncultured Thermoplasmata archaea in the Ca. Gimiplasmatales (UBA10834) contain pathways that may allow for anaerobic methylotrophic acetogenesis. Overall, MAGs from SR floodplain sediments reveal a potential for methane production and consumption in the system and a robust potential for methylotrophy.

58 GEOSCIENCES↗

Molecular and structural characterization of a Bacillus cereus strain producing an anthrax-like capsule

Bacillus cereus is a ubiquitous Gram-positive, spore-forming, rod-shaped saprophytic bacterium, occasionally reported to cause food-borne illnesses. However, instances of B. cereus strains harboring anthrax toxin and capsule genes have elevated certain strains as formidable pathogens and biothreats. This study focuses on the genomic analysis and the structural characterization of capsular material produced by the virulent B. cereus PATH2418 strain, isolated from the wound of a traumatic open fracture patient. The genome was sequenced using Nanopore MinION sequencing, revealing a chromosome of 5,270,283 bp and three plasmids. One plasmid, pATH1, was found to encode an operon for the biosynthesis of a bacterial capsule. This operon had sequence homology to the Bacillus anthracis capBCADE operon, which encodes the poly-γ-D-glutamate (PDGA) capsule. The capsule production in B. cereus PATH2418 was influenced by temperature and CO 2 levels. Structural analysis of the capsular material using a combined approach of nuclear magnetic resonance (NMR) and high-performance liquid chromatography (HPLC) techniques confirmed the presence of a high-molecular-weight poly-γ-glutamate capsule, with an enantiomeric composition of approximately 67% D-glutamic acid and 33% L-glutamic acid, matching that of B. anthracis.

Bacillus cereus↗

Structural and functional analyses of SARS-CoV-2 Nsp3 and its specific interactions with the 5’ UTR of the viral genome

ABSTRACT Non-structural protein 3 (Nsp3) is the largest open reading frame encoded in the SARS-CoV-2 genome, essential for the formation of double-membrane vesicles (DMV) wherein viral RNA replication occurs. We conducted an extensive structure-function analysis of Nsp3 and determined the crystal structures of the ubiquitin-like 1 (Ubl1), nucleic acid binding (NAB), β-coronavirus-specific marker (βSM) domains, and a sub-region of the Y domain of this protein. We show that the Ubl1, ADP-ribose phosphatase (ADRP), human SARS Unique (HSUD), NAB, and Y domains of Nsp3 bind the 5’ UTR of the viral genome and that the Ubl1 and Y domains possess affinity for recognition of this region, suggesting high specificity. The Ubl1-Nucleocapsid (N) protein complex binds the 5’ UTR with greater affinity than the individual proteins alone. Our results suggest that multiple domains of Nsp3, particularly Ubl1 and Y, shepherd the 5’ UTR of the viral genome during translocation through the DMV membrane, priming the Ubl1 domain to load the genome onto N protein. IMPORTANCE The largest protein encoded by the SARS-CoV-2 genome is Nsp3. In infected cells, this multi-domain protein forms a pore structure in the virus-induced double-membrane vesicles (DMV). We have incomplete data on Nsp3 molecular structure, and here, we describe crystal structures for multiple domains of Nsp3. It is thought that newly replicated viral RNA transits through the DMV pore; however, we possess incomplete data on which regions of Nsp3 actually interact with RNA. Here, we present data showing that five domains of Nsp3 interact with the 5’ UTR of the SARS-CoV-2 RNA, including the Y domain for which no function has ever been discovered. These data suggest that the pore structure plays an active role in recognizing the terminal end of the genome, transiting and loading the viral RNA onto the cytoplasmic nucleocapsid protein. These data help expand our knowledge of Nsp3 structure and function and the SARS-CoV-2 replication cycle.

Microbiology↗

HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs

Scientific applications produce vast amounts of data, posing grand challenges in the underlying data management and analytic tasks. Progressive compression is a promising way to address this problem, as it allows for on-demand data retrieval with significantly reduced data movement cost. However, most existing progressive methods are designed for CPUs, leaving a gap for them to unleash the power of today’s heterogeneous computing systems with GPUs.In this work, we propose HP-MDR, a high-performance and portable data refactoring and progressive retrieval framework for GPUs. Our contributions are four-fold: (1) We carefully optimize the bitplane encoding and lossless encoding, two key stages in progressive methods, to achieve high performance on GPUs; (2) We propose pipeline optimization and incorporate it with data refactoring and progressive retrieval workflows to further enhance the performance for large data process; (3) We leverage our framework to enable high-performance data retrieval with guaranteed error control for common Quantities of Interest; (4) We evaluate HP-MDR and compare it with state of the arts using five real-world datasets. Experimental results demonstrate that HP-MDR delivers an average 13.68 × and 6.31 × throughput in data refactoring and progressive retrieval tasks, respectively. It also leads to 11.22 × throughput for recomposing required data representations under Quantity-of-Interest error control and 6.04 × performance for the corresponding end-to-end data retrieval, when compared with state-of-the-art solutions.

Li, Yanliang [University of Oregon]↗

QProR: An Efficient Framework for Quantity-of-Interest Based Progressive Retrieval with Guaranteed Error Control

Scientific applications generate an unprecedented volume of data, overwhelming the network and file systems’ bandwidth and posing challenges for efficient and scalable data retrieval and analysis. Progressive data compression offers a promising solution by enabling on-demand retrieval at reduced size. However, existing progressive methods either fail to bound the errors in essential quantities of interest (QoIs) derived from raw data or suffer from suboptimal retrieval efficiency. In this work, we propose QProR, an efficient QoI-based progressive framework that optimizes progressive retrieval for target QoIs. Our key contributions include: (1) a systematic framework that integrates error-controlled lossy compressors with bitplane encoding while decoupling the two processes for high flexibility and adaptability; (2) a novel weighted bitplane encoding method which incorperates QoI knowledge into data refactoring to enhance retrieval efficiency; (3) an optimized retrieval strategy that accounts for the varying impacts of different variables on multivariate QoIs; (4) comprehensive evaluations using six real-world datasets from multiple scientific applications and thorough comparisons against state of the arts. Experimental results demonstrate that QProR achieves up to 80.38% reduction in the retrieval size under the same requested QoI error tolerance, when compared with the best-performing existing methods. When transferring 384 GB of scientific data to remote sites, QProR delivers up to 1.68 × speedup in the end-to-end data transfer performance.

Li, Wenbo [University of Kentucky]↗

Emerging Flexible Designs for Geospatial Multimodal Foundation Models

Foundation models are rapidly transforming Earth observation by enabling scalable pretraining across diverse unlabeled geospatial modalities. However, their architectural diversity—ranging from encoder-only to encoder-decoder and masked autoencoding paradigms—makes it challenging to assess performance trade-offs in a consistent manner. In this work, we present an apples-to-apples comparison of leading FM architectures designed for geospatial multimodal reasoning, with a particular focus on flexibility across varied spectral band configurations. We standardize pretraining using identical self-supervised learning objectives and training datasets, and evaluate all models under consistent parameterization on the GEOBench benchmark across classification and segmentation tasks. Our results offer new insights into the design trade-offs between model flexibility, modality alignment, and downstream task performance. By highlighting architectural strengths and limitations under controlled conditions, this study provides practical guidance for building next-generation geospatial foundation models capable of robust multimodal reasoning.

Ambrozio Dias, Philipe [ORNL] (ORCID:0000000194277↗

Deep Learning-based Non-Stationary Bias Correction (NSBC)

This work develops the NSBC (non-stationary bias correction) methodology to correct temperature projection bias from E3SM. The NSBC deep learning framework consists of a three-part architecture: an auto-encoder for compressing the spatial information, an LSTM for predicting annual temperature mean, and a U-Net for capturing the residual bias in temperature. The non-stationary bias correction (NSBC) framework can correct the non-stationarity of the biases of the climate models, which significantly improves the accuracy of future temperature prediction and improves the overestimation of extreme high temperatures that many existing bias correction methods suffer from. Getting started 1. Obtain the historical climate simulation and observation data. The E3SM simulation data are available through https://aims2.llnl.gov/search/cmip6/. The pseudo observations, the Geophysical Fluid Dynamics Laboratory (GFDL)-ESM4 model (Krasting et al., 2018) are available through https://aims2.llnl.gov/search/cmip6/. The spatial resolution of E3SM and pseudo observation datasets are both regridded to a common 1° resolution grid using conservative interpolation. The regridded E3SM and pseudo observation with 1° resolution can be found throught ./data/. 2. Train the Auto-encoder model. Python 0-autoencoder.py 3. Train the LSTM Python 1-LSTM.py 4. Generate the annual mean temperature based on trained LSTM Python 2-generate_annual_mean_LSTM.py 5. Train the U-Net. Python 3-unet.py 6. Evaluation and compared with the baseline Python 4_evaluation.py Is there a deadline approaching that requires the release of yo

Lucas, Donald↗

Implementation of ISO 15118-202 messages within Everest EV Charging Open Source Framework [SWR-25-56]

This software implements the messages defined in the ISO 15118-202 standard within the Everest EV Charging open source framework. The protocol and messages defined in the ISO 15118-202 standard enable the exchange of additional information which is not available for exchange within the currently deployed EV/EVSE communications protocols. This information includes co-identification parameters, error message exchange and more. This fork of the everest-core repository adds a prototype of the Extensible Supply Equipment Communication Controller (SECC) Discovery Protocol (ESDP) implemented based on a draft of the ISO 15118-202 standard. This is achieved through additions and modifications to the EvseV2G module. The implementation provides a demonstration of the ESDP messages, encoding and decoding but does not include a full integration within the Everest framework. Much of the information being sent over ESDP in this implementation is set statically for the sake of demonstrating the protocol itself. This fork of the ext-switchev-iso15118 repository adds a prototype of the Extensible Supply Equipment Communication Controller (SECC) Discovery Protocol (ESDP) implemented based on a draft of the ISO 15118-202 standard. The implementation provides a demonstration of the ESDP messages, encoding and decoding but does not include a full integration within the Everest framework. Much of the information being sent over ESDP in this implementation is set statically for the sake of demonstrating the protocol itself. This fork adds the ESDP features for only the EVCC controller because that is the only portion that is utilized in the everest Software-in-the-Loop.

Watt, Ed [National Renewable Energy Laboratory (NR↗

genomeocean: a pretrained microbial genome foundational model (genomeoceanLLM) v1.0

We present Genomeocean, a foundational genome language model that represents the microbial genome sequences from complex environmental samples. By training on a large, diverse metagenomic dataset, Genomeocean learns species-specific sequence composition and can generate long, realistic open reading frames (ORFs). Our model employs a Byte-pair-encoding (BPE) tokenization strategy, allowing it to efficiently process large genomic datasets and generate long sequences up to 50kb. We demonstrate that fine-tuning Genomeocean can generate novel gene clusters encoding biosynthetic pathways, showcasing its ability to model both fundamental and complex biological processes. Our work establishes Genomeocean as a powerful tool for understanding microbial genome biology and paves the way for its application in a range of fields, from synthetic biology to microbiome research.

Wang, Zhong [Lawrence Berkeley National Laboratory↗

Multitask graph neural networks for elastoplastic response prediction in dual-phase polycrystals

Microstructure-sensitive prediction of elastoplastic response remains a recurring bottleneck in multiscale damage and fatigue modeling, where large ensembles of statistically distinct polycrystals are required to quantify variability and extreme-value behavior. In this work, we develop a multitask graph neural network (GNN) surrogate that maps dual-phase ferrite–martensite polycrystal microstructures to Statistical Volume Element (SVE)-level elastoplastic Quantities of Interest (QoIs). Each SVE is represented as a grain-adjacency graph, with node features encoding phase, geometry, and crystallographic orientation, and edge features encoding relative misorientation. A message-passing graph convolution generates node embeddings, which are pooled into a graph representation and passed to a multitask regression head that jointly predicts 10 scalar QoIs and vector-valued stress–strain responses in orthogonal loading directions across multiple martensite volume fractions and SVE sizes. Results show high accuracy for scalar QoIs and strong agreement for full stress–strain trajectories, with population envelopes reproducing both median behavior and finite-SVE variability across compositions and partition scales. A unified model trained on pooled volume-fraction data preserves most within-regime accuracy relative to regime-specific models while also capturing the broader cross-regime variation reflected in the pooled test set. Distributional comparisons further demonstrate that the surrogate preserves heterogeneity under SVE partitioning, enabling statistically consistent block-wise random-field construction for mesoscale analyses. Overall, the proposed grain-graph surrogate provides a practical pathway to accelerate ensemble-based studies of SVE-level constitutive variability in dual-phase polycrystals.

Crystal plasticity↗

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

Data for Promoter Deletion in the Soybean Compact Mutant Leads to Overexpression of a Gene with Homology to the C20-Gibberellin 2-Oxidase Family

Height is a critical component of plant architecture, significantly affecting crop yield. The genetic basis of this trait in soybean remains unclear. In this study, we report the characterization of the Compact mutant of soybean, which has short internodes. The candidate gene was mapped to chromosome 17, and the interval containing the causative mutation was further delineated using biparental mapping. Whole-genome sequencing of the mutant revealed an 8.7 kb deletion in the promoter of the Glyma.17g145200 gene, which encodes a member of the class III gibberellin (GA) 2-oxidases. The mutation has a dominant effect, likely via increased expression of the GA 2-oxidase transcript observed in green tissue, as a result of the deletion in the promoter of Glyma.17g145200. We further demonstrate that levels of GA precursors are altered in the Compact mutant, supporting a role in GA metabolism, and that the mutant phenotype can be rescued with exogenous GA3. We also determined that overexpression of Glyma.17g145200 in Arabidopsis results in dwarfed plants. Thus, gain of promoter activity in the Compact mutant leads to a short internode phenotype in soybean through altered metabolism of gibberellin precursors. These results provide an example of how structural variation can control an important crop trait and a role for Glyma.17g145200 in soybean architecture, with potential implications for increasing crop yield.

Biomass Analytics↗

Data for Expression of a Bacterial Trehalose 6-Phosphate Synthase Gene otsA in Camelina sativa Seeds Promotes the Channelling of Carbon Towards Oil Accumulation

Improving seed oil yield is essential for developing Camelina sativa as a sustainable biofuel crop. Fatty acid synthesis depends on the production of acetyl-CoA from photosynthetically derived sugars. Trehalose 6-phosphate (T6P), a proxy for sucrose availability, can link sugar status to plant growth and development. Synthesised by trehalose 6-phosphate synthase (TPS) from UDP-glucose and glucose-6-phosphate, T6P plays a regulatory role in metabolism. Our previous studies on Arabidopsis transgenic lines constitutively expressing the E. coli otsA (encoding TPS) showed increased T6P levels and seed triacylglycerol, along with stunted growth. In the present study we express otsA in camelina under the control of a seed-specific Phaseolin promoter. Seeds of the resulting transgenic lines accumulated high levels of T6P, and a 15%–20% increase in total fatty acids and triacylglycerol compared to wild-type. Molecular analysis showed the transgenic seeds had reduced SnRK1 activity, elevated WRI1 protein levels, and increased the levels of WRI1 and its target genes, along with enhanced rates of fatty acid synthesis that increased seed weights relative to wild type. Notably, the increase in oil did not affect seed protein levels but did reduce the soluble metabolite fraction. Crucially, seed-specific expression of otsA mitigated the growth defects associated with constitutive otsA expression, and the transgenic lines showed normal seed development and germination. These findings demonstrate that targeted T6P modulation via seed-specific otsA expression is an effective metabolic engineering strategy to boost oil production in camelina and potentially in other oilseed crops and bioenergy crops such as energycane, sorghum and miscanthus.

Lipids↗