Search NASA⌕ Search

SEARCH · Search NASA

Results for “ACID database”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Intern Poster: STIG Shouldn't Drop ACID

STIG (Structured Threat Intelligence Graph) is an open-source graph database tool from INL. It’s used to create and process cyber intelligence graphs, which are shared in the cyber threat intelligence community and used to train INL machine learning products like @DisCo. For quality machine learning and critical infrastructure defense, STIG’s database must be ACID: Atomic, Consistent, Isolated, Durable. Various ACID tests were designed and applied to STIG to ensure its behavior follows these properties.

99 - GENERAL AND MISCELLANEOUS↗

A global database of soil microbial phospholipid fatty acids and enzyme activities

Abstract Soil microbes drive ecosystem function and play a critical role in how ecosystems respond to global change. Research surrounding soil microbial communities has rapidly increased in recent decades, and substantial data relating to phospholipid fatty acids (PLFAs) and potential enzyme activity have been collected and analysed. However, studies have mostly been restricted to local and regional scales, and their accuracy and usefulness are limited by the extent of accessible data. Here we aim to improve data availability by collating a global database of soil PLFA and potential enzyme activity measurements from 12,258 georeferenced samples located across all continents, 5.1% of which have not previously been published. The database contains data relating to 113 PLFAs and 26 enzyme activities, and includes metadata such as sampling date, sample depth, and soil pH, total carbon, and total nitrogen. This database will help researchers in conducting both global- and local-scale studies to better understand soil microbial biomass and function.

Science & Technology - Other Topics↗

EnZymClass: Substrate specificity prediction tool of plant acyl-ACP thioesterases based on ensemble learning

Characterizing the functional properties of plant acyl-ACP thioesterases (TEs), a key enzyme class used in the production of renewable oleochemicals in microbial hosts, experimentally, can be an expensive and time consuming process since it requires manual screening of thousands of candidates in a database. Using amino acid sequence to computationally predict an enzyme’s function might accelerate this process; however obtaining the necessary amount of information on previously characterized enzymes and their respective sequences required by standard Machine Learning (ML) based approaches to accurately infer sequence-function relationships can be prohibitive, especially with a low-throughput testing cycle. Experimental noise, unbalanced dataset where high sequence similarity does not always imply identical functional properties will further prevent robust prediction performance. Herein we present a ML method, Ensemble method for enZyme Classification (EnZymClass), that is specifically designed to address these issues. We used EnZymClass to classify TEs into short, long and mixed free fatty acid substrate specificity categories. While general guidelines for inferring substrate specificity have been proposed before, prediction of chain-length preference from primary sequence has remained elusive for plant acyl-ACP TEs. By applying EnZymClass to a subset of TEs in the ThYme database, we identified two medium chain TEs, ClFatB3 and CwFatB2, with previously uncharacterized activity in E. coli fatty acid production hosts.

59 BASIC BIOLOGICAL SCIENCES↗

Isolation of genome-predicted Caldatribacterium ( Atribacterota ) reveals pervasive microbial cultivation problem due to folate precipitation

Most bacterial phyla have few or no pure cultures, including Atribacterota , comprised of ubiquitous anaerobes. Here, we report genome-guided enrichment and isolation of two Atribacterota species representing a new family, Caldatribacterium saccharofermentans from a hot spring, and Caldatribacterium inferamans from a deep aquifer. Both were co-enriched with sulfate-reducing bacteria and initially resisted isolation, which we link to inadvertent removal of precipitated folic acid by filter-sterilization of unbuffered Wolin’s vitamin solution. We then predict folate auxotrophy across the Atribacterota and ~29% of all bacteria, with extensive auxotrophy in 27% of phyla. Since ≥604 of 791 ( ≥ 76%) media with folic acid additions in the MediaDive database use unbuffered vitamin solutions in which folic acid is likely removed during filter-sterilization, we propose that folate auxotrophy limits culturability in defined media en masse. We also uncover unusual features of Caldatribacterium , including three lipid membrane-like layers (LMLs), with the inner LML surrounding the nucleoid, and a high percentage of secreted proteins, supporting a unique cell biology of Atribacterota .

Biological and medical sciences↗

Database of Nonaqueous Proton-Conducting Materials

This work presents the assembly of 48 papers, representing 74 different compounds and blends, into a machine-readable database of nonaqueous proton-conducting materials. SMILES was used to encode the chemical structures of the molecules, and we tabulated the reported proton conductivity, proton diffusion coefficient, and material composition for a total of 3152 data points. The data spans a broad range of temperatures ranging from -70 to 260 °C. To explore this landscape of nonaqueous proton conductors, DFT was used to calculate the proton affinity of 18 unique proton carriers. The results were then compared to the activation energy derived from fitting experimental data to the Arrhenius equation. It was found that while the widely recognized positive correlation between the activation energy and proton affinity may hold among closely related molecules, this correlation does not necessarily apply across a broader range of molecules. This work serves as an example of the potential analyses that can be conducted using literature data combined with emerging research tools in computation and data science to address specific materials design problems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Uncovering fast solid-acid proton conductors based on dynamics of polyanion groups and proton bonding strength

Achieving high proton conductivity in inorganic solids is key for advancing many electrochemical technologies, including low-energy nano-electronics and energy-efficient fuel cells and electrolyzers. A quantitative understanding of the physical traits of a material that regulate proton diffusion is necessary for accelerating the discovery of fast proton conductors. In this work, we have mapped the structural, chemical and dynamic properties of solid acids to the elementary steps of the Grotthuss mechanism of proton diffusion. Our approach combines ab initio molecular dynamics simulations, analysis of phonon spectra and atomic structure calculations. We have identified the donor–hydrogen bond lengths and the acidity of polyanion groups as key descriptors of local proton transfer and the vibrational frequencies of the cation framework as the key descriptor of lattice flexibility. The latter facilitates rotations of polyanion groups and long-range proton migration in solid acid proton conductors. The calculated lattice flexibility also correlates with the experimentally reported superprotonic transition temperatures. Using these descriptors, we have screened the Materials Project database and identified potential solid acid proton conductors with monovalent, divalent and trivalent cations, including Ag + , Sr 2+ , Ba 2+ and Er 3+ cations, which go beyond the traditionally considered monovalent alkali cations (Cs + , Rb + , K + , and NH 4 + ) in solid acids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS↗

In Silico Screening of CO 2 –Dipeptide Interactions for Bioinspired Carbon Capture

Carbon capture, sequestration and utilization offers a viable solution for reducing the total amount of atmospheric CO 2 concentrations. On an industrial scale, amine-based solvents are extensively employed for CO 2 capture through chemisorption. Nevertheless, this method is marked by the high cost associated with solvent regeneration, high vapor pressure, and the corrosive and toxic attributes of by-products, such as nitrosamines. An alternative approach is the biomimicry of sustainable materials that have strong affinity and selectivity for CO 2 . Bioinspired approaches, such as those based on naturally occurring amino acids, have been proposed for direct air capture methodologies. In this study, we present a database consisting of 960 dipeptide molecular structures, composed of the 20 naturally occurring amino acids. Furthermore, those structures were analyzed with a novel computational workflow presented in this work that considers certain interaction sites that determine CO 2 affinity. Density functional theory (DFT) and symmetry-adapted perturbation theory (SAPT) computations were performed for the calculation of CO 2 interaction energies, which allowed to limit our search space to 400 unique dipeptide structures. Using this computational workflow, we provide statistical insights into dipeptides and their affinity for CO 2 binding, as well as design principles that can further enhance CO 2 capture through cooperative binding.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multiplex detection and identification of viral, bacterial, and protozoan pathogens in human blood and plasma using an expanded high-density resequencing microarray platform

Introduction: Nucleic acid tests for blood donor screening have improved the safety of the blood supply; however, increasing numbers of emerging pathogen tests are burdensome. Multiplex testing platforms are a potential solution. Methods: The Blood Borne Pathogen Resequencing Microarray Expanded (BBP-RMAv.2) can perform multiplex detection and identification of 80 viruses, bacteria and parasites. This study evaluated pathogen detection in human blood or plasma. Samples spiked with selected pathogens, each with one of 6 viruses, 2 bacteria and 5 protozoans were tested on this platform. The nucleic acids were extracted, amplified using multiplexed sets of primers, and hybridized to a microarray. The reported sequences were aligned to a database to identify the pathogen. To directly compare the microarray to an emerging molecular approach, the amplified nucleic acids were also submitted to nanopore next generation sequencing (NGS). Results: The BBP-RMAv.2 detected viral pathogens at a concentration as low as 100 copies/ml and a range of concentrations from 1,000 to 100,000 copies/ml for all the spiked pathogens. Coded specimens were identified correctly demonstrating the effectiveness of the platform. The nanopore sequencing correctly identified most samples and the results of the two platforms were compared. Discussion: These results indicated that the BBP-RMAv.2 could be employed for multiplex detection with potential for use in blood safety or disease diagnosis. The NGS was nearly as effective at identifying pathogens in blood and performed better than BBP-RMAv.2 at identifying pathogen-negative samples.

59 BASIC BIOLOGICAL SCIENCES↗

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

Electronic structure theory on modeling short-range noncovalent interactions between amino acids

While short-range noncovalent interactions (NCIs) are proving to be of importance in many chemical and biological systems, these atypical bindings happen within the so-called van der Waals envelope and pose an enormous challenge for current computational methods. We introduce SNCIAA, a database of 723 benchmark interaction energies of short-range noncovalent interactions between neutral/charged amino acids originated from protein x-ray crystal structures at the “gold standard” coupled-cluster with singles, doubles, and perturbative triples/complete basis set [CCSD(T)/CBS] level of theory with a mean absolute binding uncertainty less than 0.1 kcal/mol. Subsequently, a systematic assessment of commonly used computational methods, such as the second-order Møller−Plesset theory (MP2), density functional theory (DFT), symmetry-adapted perturbation theory (SAPT), composite electronic-structure methods, semiempirical approaches, and the physical-based potentials with machine learning (IPML) on SNCIAA is carried out. It is shown that the inclusion of dispersion corrections is essential even though these dimers are dominated by electrostatics, such as hydrogen bondings and salt bridges. Overall, MP2, ωB97M-V, and B3LYP+D4 turned out to be the most reliable methods for the description of short-range NCIs even in strongly attractive/repulsive complexes. SAPT is also recommended in describing short-range NCIs only if the δMP2 correction has been included. The good performance of IPML for dimers at close-equilibrium and long-range conditions is not transferable to the short-range. We expect that SNCIAA will assist the development/improvement/validation of computational methods, such as DFT, force-fields, and ML models, in describing NCIs across entire potential energy surfaces (short-, intermediate-, and long-range NCIs) on the same footing.

Chemistry↗

High-resolution rovibrational spectroscopy of trans -formic acid in the v 1 OH stretching fundamental: Dark state coupling and evidence for charge delocalization dynamics

High-resolution infrared (IR) reduced-Doppler absorption spectra of jet-cooled gas phase trans-formic acid in the v 1 OH stretching fundamental region are reported for the first time, obtained by supersonically expanding trans- formic acid/Ar mixtures through a slit jet nozzle source and rotationally cooling to T rot ≈ 10.9(5) K, with ab- sorption signals recorded by high-resolution difference-frequency IR absorption spectroscopy. Two a/b-type rovibrational bands of comparable intensity, one ~10-fold weaker b-type band, and one ~6-fold weaker a-type band are observed, with vibrational band origins at 3570.493(5), 3566.793(5), 3560.032(9), and 3534.6869(2) cm –1 , respectively. Based on previous Raman jet spectroscopic work by Nejad and Sibert [A. Nejad, E.L. Sibert III, The Raman jet spectrum of trans-formic acid and its deuterated isotopologs: Combining theory and experi- ment to extend the vibrational database, J. Chem. Phys. 154(6) (2021) 064301.], these four rovibrational bands have been assigned to v 1 , (v 2 + v 7 ), (v 6 + 2v 7 + 2v 9 ), and 2v 3 , respectively. Specifically, two of the three upper dark states (2 1 7 1 (a') and 6 1 7 2 9 2 (a')) are close enough to the “bright” 1 1 (a') state to facilitate strong anharmonic resonance interactions, which results in intensity mixing into the two zero-order bands that would otherwise be “dark”. Furthermore, our high-resolution spectral analysis reveals that there are local rotational crossings be- tween these zero-order 1 1 and 2 1 7 1 states resulting in extra lines (i.e., some upper levels in the nominally v 1 band have majority zero-order 2 1 7 1 state character). This motivates development of a 3 coupled state (1 1 , 2 1 7 1 , and 6 1 7 2 9 2 ) picture to aid in the spectral analysis, which is able to match all 3 observed band origins and relative band intensities, as well as indicate the necessity of multistate (> 2) coupling. Though limited by range of J and Ka levels (J’ ≤ 9 and K a ’ ≤ 3) populated at supersonic jet temperatures, this work offers first precision spec- troscopic analysis of trans-formic acid in the v 1 OH stretch region, which should aid in assignment of the more complete yet highly congested room temperature FTIR spectra. Lastly, and in sharp contrast to the spectral complexity in the three predominantly b-type bands, the lone a-type 2v 3 rovibrational band at 3534.6869(2) cm –1 is well described by a simple, rigid asymmetric top Hamiltonian.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Benzyl trichloroacetimidates as derivatizing agents for phosphonic acids related to nerve agents by EI-GC-MS during OPCW proficiency test scenarios

Abstract The use of benzyl trichloroacetimidates for the benzylation of phosphonic acid nerve agent markers under neutral, basic, and slightly acidic conditions is presented. The benzyl-derived phosphonic acids were detected and analyzed by Electron Ionization Gas Chromatography–Mass Spectrometry (EI-GC–MS). The phosphonic acids used in this work included ethyl-, cyclohexyl- and pinacolyl methylphosphonic acid, first pass hydrolysis products from the nerve agents ethyl N -2-diisopropylaminoethyl methylphosphonothiolate (VX), cyclosarin (GF) and soman (GD) respectively. Optimization of reaction parameters for the benzylation included reaction time and solvent, temperature and the effect of the absence or presence of catalytic acid. The optimized conditions for the derivatization of the phosphonic acids specifically for their benzylation, included neutral as well as catalytic acid (< 5 mol%) and benzyl 2,2,2-trichloroacetimidate in excess coupled to heating the mixture to 60 °C in acetonitrile for 4 h. While the neutral conditions for the method proved to be efficient for the preparation of the p -methoxybenzyl esters of the phosphonic acids, the acid-catalyzed process appeared to provide much lower yields of the products relative to its benzyl counterpart. The method’s efficiency was tested in the successful derivatization and identification of pinacolyl methylphosphonic acid (PMPA) as its benzyl ester when present at a concentration of ~ 5 μg/g in a soil matrix featured in the Organisation for the Prohibition of Chemical Weapons (OPCW) 44th proficiency test (PT). Additionally, the protocol was used in the detection and identification of PMPA when spiked at ~ 10 μg/mL concentration in a fatty acid-rich liquid matrix featured during the 38 th OPCW-PT. The benzyl derivative of PMPA was partially corroborated with the instrument's internal NIST spectral library and the OPCW central analytical database (OCAD v.21_2019) but unambiguously identified through comparison with a synthesized authentic standard. The method’s MDL (LOD) values for the benzyl and the p -methoxybenzyl pinacolyl methylphosphonic acids were determined to be 35 and 63 ng/mL respectively, while the method’s Limit of Quantitation (LOQ) was determined to be 104 and 189 ng/mL respectively in the OPCW-PT soil matrix evaluated.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Generative $β$-hairpin design using a residue-based physicochemical property landscape

De novo peptide design is a new frontier that has broad application potential in the biological and biomedical fields. Most existing models for de novo peptide design are largely based on sequence homology that can be restricted based on evolutionarily derived protein sequences and lack the physicochemical context essential in protein folding. Generative machine learning for de novo peptide design is a promising way to synthesize theoretical data that are based on, but unique from, the observable universe. In this study, we created and tested a custom peptide generative adversarial network intended to design peptide sequences that can fold into the -hairpin secondary structure. This deep neural network model is designed to establish a preliminary foundation of the generative approach based on physicochemical and conformational properties of 20 canonical amino acids, for example, hydrophobicity and residue volume, using extant structure-specific sequence data from the PDB. The beta generative adversarial network model robustly distinguishes secondary structures of hairpin from α helix and intrinsically disordered peptides with an accuracy of up to 96% and generates artificial -hairpin peptide sequences with minimum sequence identities around 31% and 50% when compared against the current NCBI PDB and nonredundant databases, respectively. These results highlight the potential of generative models specifically anchored by physicochemical and conformational property features of amino acids to expand the sequence-to-structure landscape of proteins beyond evolutionary limits.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ToF-SIMS spectral analysis of Shewanella oneidensis MR-1 biofilms

Analysis of bacterial biofilms is particularly challenging and important with diverse applications from systems biology to biotechnology. Among the variety of techniques that have been applied, time-of-flight secondary ion mass spectrometry (ToF-SIMS) has many powerful features in studying the surface characteristics of biofilms. ToF-SIMS offers high spatial resolution, mass resolution, and mass accuracy, which permit surface sensitive analysis of biofilm components. Thus, ToF-SIMS provides a powerful solution to addressing the challenge of bacterial biofilm analysis. This dataset covers ToF-SIMS analysis of Shewanella oneidensis MR-1 isolated from freshwater lake sediment in New York state. The MR-1 strain is known to have metal and sulfur reducing properties and it can be used for bioremediation and wastewater treatment. There is a current need to identify small molecules and fragments produced from bacterial biofilms, especially those from extracellular polymeric substance (EPS). Static ToF-SIMS spectra of MR-1 were obtained using an IONTOF TOF.SIMS V instrument equipped with a 25 keV Bi$^+_3$ metal ion gun. Identified molecules and molecular fragments are compared against known biological databases and the reported peaks have at least 65 ppm mass accuracy. These molecules range from lipids, fatty acids, flavonoids, and quinolones to other naturally occurring organic compounds. It is anticipated that the mass spectral identification of key peaks will assist detection of metabolites, EPS molecules like polysaccharides, and biologically relevant small organic molecules using ToF-SIMS in future surface and interface research.

59 BASIC BIOLOGICAL SCIENCES↗

ToF-SIMS spectral data analysis of Paenibacillus sp. 300A biofilms and planktonic cells

Analysis of bacterial biofilms is particularly challenging and important with diverse applications from systems biology to biotechnology. Among the variety of techniques that have been applied, time-of-flight secondary ion mass spectrometry (ToF-SIMS) has many promising features in studying the surface characteristics of biofilms. ToF-SIMS offers high spatial resolution and high mass accuracy, which permit surface sensitive analysis of biofilm components. Thus, ToF-SIMS provides a powerful solution to addressing the challenge of bacterial biofilm analysis. This dataset covers ToF-SIMS analysis of Paenibacillus sp. 300A (300A) isolated from the Hanford site in Richland, WA. The strain is known to have metal and sulfur reducing properties and can be used for bioremediation, wastewater treatment, bioengineering and technology development. There is a current need to identify small molecules and fragments produced from bacterial biofilms. Static ToF-SIMS spectra of 300A were obtained using an IONTOF TOF-SIMS V instrument equipped with a 25 keV Bi 3 + metal ion gun. Identified molecules and molecular fragments are compared against known biological databases and the reported peaks have at least 65 ppm mass accuracy. These molecules range from lipids and fatty acids to flavonoids, quinolones, and other naturally occurring organic compounds. It is anticipated that the spectral identification of key peaks will assist detection of metabolites, extracellular polymeric substance molecules like polysaccharides, and biologically relevant small molecules using ToF-SIMS in future surface and interface research of bacterial biofilms.

Biofilms↗

Colloidal Behavior of Plutonium Oxide in Concentrated Electrolyte Solutions

The Hanford Site in Washington State manages legacy high-level radioactive waste streams that display major chemistry and engineering challenges, including the high salt levels and pH values that correspond to conditions under which many classical concepts describing chemical reactivity cannot be applied. One particular challenge that needs to be tackled at Hanford is that Pu concentrations, [Pu], in the soluble phases of the tank wastes are higher than expected based on the solubility of crystalline PuO2, which is widely accepted to be caused by the formation of PuO2 colloid, consisting of nano- to submicron-sized particles (PuO2 NPs). Fundamental research underpinning the behavior of PuO2 NPs under conditions not only relevant to the Hanford tank waste but at high ionic strength in general is needed to reliably predict the chemical reactivity of PuO2 NPs and develop engineering solutions to safely and efficiently process high-level radioactive waste into forms suitable for long-term storage. In this work, we study the behavior of PuO2 NPs (particle size ~100 nm) under high ionic strength conditions by reacting it with highly concentrated (up to 5 M) salt solutions. We explore different electrolyte compositions to elucidate the impact of different anions (NO3-, Cl-, ClO4-, SO42-, C2O42-, CO32-) on the stability of PuO2 NP in the acidic and alkaline pH regime. PuO2 NP aggregation and precipitation as function electrolyte concentration is tracked by a combination of liquid scintillation counting, dynamic light scattering for determination of particle size distributions, and zeta potentials as a proxy for particle charge. At acidic pH, electrolytes containing non-coordinating anions, such as NaNO3, NaCl, and NaClO4 mostly stabilize PuO2 NPs over a large electrolyte concentration range, showing only subtle differences in their reactivity. Other electrolyte anions show a more pronounced effect on the PuO2 NP stability: SO42-, binds directly to the particles’ surface, reverses the particle charge, and precipitates the PuO2 NPs efficiently even at intermediate sulfate concentrations (>0.1 M). In contrast, C2O42- is found to lead to high [Pu] in solution, in the milli-molar range, even at mildly acidic pH (~4). Thermodynamic modeling of the dissolved Pu concentrations using PHREEQC is unable to predict the observed [Pu] in the acidic pH regime, supporting the influence of colloids in maintaining elevated [Pu]. It is noteworthy that the current thermodynamic databases do not include constants for colloidal Pu phases and cannot accurately predict many of the high ionic strength solutions relevant to this work. The mechanisms and models responsible for these observations will need further investigation in the future. At high pH values (~12), PuO2 NPs exhibits classical sol-gel chemistry, meaning that upon destabilization of the colloidal sol, for example by addition of concentrated NaOH, highly porous and viscid PuO2 coagulates are formed that consist of a three-dimensional network likely held together by physical interactions. The PuO2 NP coagulate shows no significant reversibility of the aggregation when contacted with concentrated brines; however, PuO2 NPs can be efficiently resuspended in solution by addition of diluted electrolytes, alkaline solutions containing high amounts of carbonate, or simple addition of water. Especially carbonate is shown to stabilize PuO2 NPs in solution at high pH, characterized by stable colloidal suspensions that are resistant against sedimentation during centrifugation. Thermodynamic modeling of the carbonate system was able to predict an increasing dissolved Pu concentration with increasing carbonate concentration. However, the model was profoundly sensitive to the fixed redox potential and does not include any thermodynamic constants for colloidal Pu species.

Neumann, Julia↗

An optimized FM-index library for nucleotide and amino acid search

Abstract Background Pattern matching is a key step in a variety of biological sequence analysis pipelines. The FM-index is a compressed data structure for pattern matching, with search run time that is independent of the length of the database text. Implementation of the FM-index is reasonably complicated, so that increased adoption will be aided by the availability of a fast and flexible FM-index library. Results We present AvxWindowedFMindex (AWFM-index), a lightweight, open-source, thread-parallel FM-index library written in C that is optimized for indexing nucleotide and amino acid sequences. AWFM-index introduces a new approach to storing FM-index data in a strided bit-vector format that enables extremely efficient computation of the FM-index occurrence function via AVX2 bitwise instructions, and combines this with optional on-disk storage of the index’s suffix array and a cache-efficient lookup table for partial k-mer searches. The AWFM-index performs exact match count and locate queries faster than SeqAn3’s FM-index implementation across a range of comparable memory footprints. When optimized for speed, AWFM-index is $$\sim $$ ∼ 2–4x faster than SeqAn3 for nucleotide search, and $$\sim $$ ∼ 2–6x faster for amino acid search; it is also $$\sim $$ ∼ 4x faster with similar memory footprint when storing the suffix array in on-disk SSD storage. Conclusions AWFM-index is easy to incorporate into bioinformatics software, offers run-time performance parameterization, and provides clients with FM-index functionality at both a high-level (count or locate all instances of a query string) and low-level (step-wise control of the FM-index backward-search process). The open-source library is available for download at https://github.com/TravisWheelerLab/AvxWindowFmIndex.

59 BASIC BIOLOGICAL SCIENCES↗