Search NASA⌕ Search

SEARCH · Search NASA

Results for “sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

SynBio QC Dual Barcode QC (DBC) v1.0

This software was designed as a sequence validation tool for the assembly of synthetic constructs, where the constructs have a high degree of similarity and thus are barcoded prior to the sequencing library prep. It demultiplexes each FASTQ file for each barcode, then analyzes the resulting FASTQ files against a list of reference sequences for that barcode/library, combining the results from eight sequencing libraries to generate a summary, and the files needed to view the results in the Integrative Genomics Viewer (IGV) application for manual verification. This was developed for FASTQ files generated by PacBio sequencing, but could be used on any FASTQ files that do not have paired end reads. It can be used to analyze one - eight libraries at a time. Each construct is independently analyzed with only the sequences with the same barcode, in the same pooled library. Then the results are combined into a user friendly summary. This is used to identify which libraries of pooled sequences contains a perfect match, or fixable match to the reference file. This pipeline uses many freely available open source libraries, the value added is that in our application the steps of the pipeline are defined in Workflow Description Language (WDL) and run through the Cromwell workflow engine in Docker containers, for easy distribution and set up, as well as the user friendly html summary that is generated.

Simirenko, Lisa↗

Multiplexing and Demultiplexing Signals for Radiography Application Using the Discrete Fourier Transform

Our goal is to develop an X-ray phase-contrast imaging system that can provide excellent soft tissue contrast of phase, attenuation, and small-angle scatter. We propose to replace the common system of G0, G1, and G2 gradings with a biprism array to replace the G1 grading and introduce a novel X-ray tube designed to replace the motion of the phase stepping grading G2. The proposed X-ray tube uses temporal multiplexing to provide simultaneous virtual “electronic phase stepping.” In this work the discrete Fourier transform is used to separate from the composite measurement individual X-ray phase contrast measurements sampled at different frequencies. The method performs a discrete Fourier transform of a composite refence sequence to obtain using the frequency amplitudes calibration factors needed to extract the X-ray phase contrast measurement amplitudes from the composite image. The composite reference sequence is the sum of the individual sequences, at different frequencies, with amplitudes of one. The method takes the discrete Fourier transform of this composite reference sequence; whereby, the amplitude of each frequency component is compared with the total sum of its stand-alone sequence amplitude. A calibration factor is determined so that the amplitude of this composite reference frequency times the calibration factor must equal the total sum of the sequence amplitude—the zero-frequency amplitude of the discrete Fourier transform of its stand-alone sequence. To demultiplex the composite measured signal these calibration factors are multiplied by the amplitudes of the frequency components of the discrete Fourier transform of the composite X-phase-contrast measurement to obtain the amplitude of each frequency encoded measurement. Using these calibration factors, we demonstrate with the discrete Fourier transform in Mathematica the extraction of individual images from a composite image that one would expect obtaining from our proposed new X-ray phase contrast imaging system. We then demonstrate as an example how using images from X-ray phase contrast data one can calculate phase, attenuation and the dark field images using grading phase step data supplied to use from Microworks, GmbH in Karlsruhe, Germany.

42 ENGINEERING↗

Populus VariantDB v3.2 facilitates CRISPR and functional genomics research

The success of CRISPR genome editing studies depends critically on the precision of guide RNA (gRNA) design. Sequence polymorphisms in outcrossing tree species pose design hazards that can render CRISPR genome editing ineffective. Despite recent advances in tree genome sequencing with haplotype resolution, sequence polymorphism information remains largely inaccessible to various functional genomics research efforts. The Populus VariantDB v3.2 addresses these challenges by providing a user-friendly search engine to query sequence polymorphisms of heterozygous genomes. The database accepts short sequences, such as gRNAs and primers, as input for searching against multiple poplar genomes, including hybrids, with customizable parameters. We provide examples to showcase the utilities of VariantDB in improving the precision of gRNA or primer design. The platform-agnostic nature of the probe search design makes Populus VariantDB v3.2 a versatile tool for the rapidly evolving CRISPR field and other sequence-sensitive functional genomics applications. The database schema is expandable and can accommodate additional tree genomes to broaden its user base.

59 BASIC BIOLOGICAL SCIENCES↗

Rapid Detection and Quick Characterization of African Swine Fever Virus Using the VolTRAX Automated Library Preparation Platform

African swine fever virus (ASFV) is the causative agent of a severe and highly contagious viral disease affecting domestic and wild swine. The current ASFV pandemic strain has a high mortality rate, severely impacting pig production and, for countries suffering outbreaks, preventing the export of their pig products for international trade. Early detection and diagnosis of ASFV is necessary to control new outbreaks before the disease spreads rapidly. One of the rate-limiting steps to identify ASFV by next-generation sequencing platforms is library preparation. Here, we investigated the capability of the Oxford Nanopore Technologies’ VolTRAX platform for automated DNA library preparation with downstream sequencing on Nanopore sequencing platforms as a proof-of-concept study to rapidly identify the strain of ASFV. Within minutes, DNA libraries prepared using VolTRAX generated near-full genome sequences of ASFV. Thus, our data highlight the use of the VolTRAX as a platform for automated library preparation, coupled with sequencing on the MinION Mk1C for field sequencing or GridION within a laboratory setting. These results suggest a proof-of-concept study that VolTRAX is an effective tool for library preparation that can be used for the rapid and real-time detection of ASFV.

60 APPLIED LIFE SCIENCES↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

A Route to Design Novel Functional Peptides by Applying a Denoising Diffusional Model to mRNA Display Libraries

In vitro directed evolution techniques, such as mRNA display, enable peptide ligand discovery and optimization. However, physical libraries that rely on a genetic code can only search a small fraction of sequence space due to inherent biases in the genetic code and experimental limitations. To address this challenge, denoising diffusion implicit models (DDIMs) are applied to generate novel peptide ligands against B‐cell lymphoma extra‐large (Bcl‐x L ), a key cancer target. Starting with high‐throughput sequencing data from previous selections, a DDIM is trained to produce novel sequences with high affinity binding. Experimental validation confirms that most generated sequences are functionally equivalent to the original library members for Bcl‐x L binding and demonstrated comparable binding kinetics and affinity relative to the wildtype and nearest original neighbors. Importantly, this approach generated rare sequences not easily accessible via mutation and directed evolution. These results indicate that DDIMs can complement and expand directed evolution data, efficiently exploring underrepresented regions of sequence space. This approach provides a broadly applicable framework for accelerating ligand discovery and optimizing molecular properties across diverse targets.

Qi, Pearl [Mork Family Department of Chemical Engi↗

Apatite geochemistry as a tool for understanding the petrogenesis of layered mafic-ultramafic rocks in the Bushveld Complex, South Africa

The sources of the magmas that formed the Rustenburg Layered Suite of the Bushveld Complex in South Africa remain debated, despite decades of research. Vertical and lateral variation in bulk rock and mineral separate Sr-Nd isotopic compositions, which generally indicate enriched sources, demonstrate that the layered sequence was formed by the emplacement of multiple batches of magma, crucially resulting in episodes of PGE-Cr-V mineralisation. The Lu-Hf isotope compositions of zircon are, however, at odds with the bulk rock Sr-Nd isotopic heterogeneity as they show near homogeneous compositions throughout the layered sequence (εHf (2.06 Ga) =−8). This lack of variation in Hf isotope composition has been attributed to deep, continental lithospheric mantle-related and/or crustal contamination of plume-derived Bushveld magmas. In this study, we analysed the major, trace element and Sr-Nd isotope geochemistry of apatite in the Rustenburg Layered Suite. Apatite occurs as an intercumulus mineral in the lowermost regions and a cumulus mineral in the uppermost regions of the layered sequence and can therefore be used to test existing models for the isotopic disequilibrium between bulk rock Sr-Nd and zircon Hf isotopic compositions. Apatite is largely chlorapatite in the lowermost regions and fluorapatite in the uppermost regions of the layered sequence. The Merensky Reef is unusual in that it comprises both chlorapatite and fluorapatite. Apatite throughout the layered sequence is generally unzoned and shows no evidence of late-stage alteration. Trace element data show that apatite is enriched in L/HREE, with common negative Eu-Sr anomalies. These trace element signatures are consistent with a magmatic origin for the apatite grains, with prior, or concurrent, plagioclase crystallization from the same melt. Variability in in situ Sr and Nd isotope compositions of apatite is recorded throughout the layered sequence with εNd (2.06 Ga) compositions varying between −2.5 and − 10.2 and initial 87 Sr/ 86 Sr compositions varying between 0.7079 and 0.7103 (for the Marikana dikes only). The variability in Sr-Nd isotope compositions of apatite is consistent with the bulk rock (and mineral separate) variation in Sr-Nd isotope compositions, suggesting apatite preserves primary magmatic compositions in the Rustenburg Layered Suite.

Apatite↗

Role of Ribosomal Protein bS1 in Orthogonal mRNA Start Codon Selection

In many bacteria, the location of the mRNA start codon is determined by a short ribosome binding site sequence that base pairs with the 3'-end of 16S rRNA (rRNA) in the 30S subunit. Many groups have changed these short sequences, termed the Shine-Dalgarno (SD) sequence in the mRNA and the anti-Shine-Dalgarno (ASD) sequence in 16S rRNA, to create "orthogonal" ribosomes to enable the synthesis of orthogonal polymers in the presence of the endogenous translation machinery. However, orthogonal ribosomes are prone to SD-independent translation. Ribosomal protein bS1, which binds to the 30S ribosomal subunit, is thought to promote translation initiation by shuttling the mRNA to the ribosome. Thus, a better understanding of how the SD and bS1 contribute to start codon selection could help efforts to improve the orthogonality of ribosomes. Here, we engineered the Escherichia coli ribosome to prevent binding of bS1 to the 30S subunit and separate the activity of bS1 binding to the ribosome from the role of the mRNA SD sequence in start codon selection. We find that ribosomes lacking bS1 are slightly less active than wild-type ribosomes in vitro. Furthermore, orthogonal 30S subunits lacking bS1 do not have an improved orthogonality. Our findings suggest that mRNA features outside the SD sequence and independent of binding of bS1 to the ribosome likely contribute to start codon selection and the lack of orthogonality of present orthogonal ribosomes.

59 BASIC BIOLOGICAL SCIENCES↗

Native Chemical Ligation of Peptoid Oligomers

Bioorganic chemists are inspired by natural biopolymers to design peptidomimetic oligomers that can exhibit sequence-structure-function relationships. Biomimetic polymers can be synthesized to incorporate a specific sequence of nonbiological monomer units using a variety of iterative solution-phase or solid-phase reaction schemes. These protocols generally provide access to a vast diversity of oligomeric compounds but are limited with respect to their ability to attain protein-like chain lengths. This constraint can preclude access to sequence-defined synthetic macromolecules with sufficient sizes required to exhibit tertiary structure and other protein-mimetic attributes. In contrast, peptide chemists have overcome this limitation by developing convergent synthetic methods, such as native chemical ligation, to join individual, smaller peptide chains together to make larger peptides or full proteins. A similar convergent approach is needed to establish efficient synthetic routes to non-natural sequence-defined macromolecules. Herein, we adapt the peptide native chemical ligation method to peptoid oligomers, demonstrating how short chains can be conjoined to create sequence-defined peptoid macromolecules. Nanosheet-forming peptoid polymers with distinct surface loop display domains were generated by sequential ligation of several discrete fragments. This method provides a reliable convergent ligation route for sequence-defined polypeptoids that results in a native amide bond joining the fragments. We envision that this strategy will be useful in synthesizing peptoid-based proteomimetics that incorporate diverse chemical features.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Toward Computation-Guided Design of Tunable Organic-Inorganic CdS Quantum Dot Binary Superlattices

Combining the advantages of structural programmability in sequence-defined biomimetic molecules and the controllable packing geometry in nanoparticle superlattices, we demonstrate a self-assembled organic-inorganic superlattice whose structure can be altered with the slightest change in the sequence of the organic counterpart. Here, oleate-coated CdS quantum dots (QDs) form a square-packed superlattice with a 1:1 molar equivalence of a di-block amphiphilic peptoid (Nbrpe6Dig) in chloroform. In contrast, no apparent structure is observed in the organic solvent alone. Based on theoretical evidence, we show that the assembly is a binary superlattice where both the CdS QDs and the peptoids serve as building blocks and further predict a correlation between the superlattice structure and the peptoid sequence. The computationally guided prediction is validated by experiments where superlattice transformation is observed with modified peptoids. The mechanism identified in our work inspires new ways to control and tune organic-inorganic hybrid nanomaterial self-assembly.

Qi, Xin↗

Whole-genome demography of COVID-19 virus during its pandemic period and on “panvalent” vaccine design

With over 16 million submitted genomic sequences, the SARS-CoV-2 (SC2) virus, the cause of the most recent worldwide COVID-19 pandemic, has become the most sequenced genome of all known viruses, revealing, for example, a vast number of expanding viral lineages. Since the pandemic phase appears to be over, we performed a retrospective re-examination of the demographic grouping pattern and their genomic characteristics during the entire pandemic period up to the peak of the last pandemic wave. For our study, we extracted from the NCBI only unique viral sequences and converted each sequence data to a relational vector, indicating the presence/absence of each variational event compared to a “reference” sequence. Our study revealed several genomic features that are unexpected or different from those of previous studies. For example, approximately 44,000 variants with unique sequences emerged during the pandemic period; they group into only four major viral-genomic groups and each has a set of mostly unique highly-conserved variant-genotypes (HCVGs); and a small set from the first (“ancestral”) group was inherited by the three (“descendant”) groups, suggesting that HCVGs in the next group may be predictable from the current group(s). Such a concept may be potentially important in designing “panvalent” vaccines against the current and future waves of viral infections.

60 APPLIED LIFE SCIENCES↗

Surfactant-like peptide gels are based on cross-β amyloid fibrils

Surfactant-like peptides, in which hydrophilic and hydrophobic residues are encoded within different domains in the peptide sequence, undergo facile self-assembly in aqueous solution to form supramolecular hydrogels. These peptides have been explored extensively as substrates for the creation of functional materials since a wide variety of amphipathic sequences can be prepared from commonly available amino acid precursors. The self-assembly behavior of surfactant-like peptides has been compared to that observed for small molecule amphiphiles in which nanoscale phase separation of the hydrophobic domains drives the self-assembly of supramolecular structures. Here, we investigate the relationship between sequence and supramolecular structure for a pair of bola-amphiphilic peptides, Ac-KLIIIK-NH 2 (L2) and Ac-KIIILK-NH 2 (L5). Despite similar length, composition, and polar sequence pattern, L2 and L5 form morphologically distinct assemblies, nanosheets and nanotubes, respectively. Cryo-EM helical reconstruction was employed to determine the structure of the L5 nanotube at near-atomic resolution. Rather than displaying self-assembly behavior analogous to conventional amphiphiles, the packing arrangement of peptides in the L5 nanotube displayed steric zipper interfaces that resembled those observed in the structures of β-amyloid fibrils. Like amyloids, the supramolecular structures of the L2 and L5 assemblies were sensitive to conservative amino acid substitutions within an otherwise identical amphipathic sequence pattern. This study highlights the need to better understand the relationship between sequence and supramolecular structure to facilitate the development of functional peptide-based materials for biomaterials applications.

Das, Abhinaba [Emory University, Atlanta, GA (Unit↗

Prevalence and diversity of TAL effector-like proteins in fungal endosymbiotic Mycetohabitans spp.

EndofungalMycetohabitans(formerlyBurkholderia) spp. rely on a type III secretion system to deliver mostly unidentified effector proteins when colonizing their host fungus,Rhizopus microsporus. The one known secreted effector family fromMycetohabitansconsists of homologues of transcription activator-like (TAL) effectors, which are used by plant pathogenicXanthomonasandRalstoniaspp. to activate host genes that promote disease. These ‘BurkholderiaTAL-like (Btl)’ proteins bind corresponding specific DNA sequences in a predictable manner, but their genomic target(s) and impact on transcription in the fungus are unknown. Recent phenotyping of Btl mutants of twoMycetohabitansstrains revealed that the single Btl in oneMycetohabitans endofungorumstrain enhances fungal membrane stress tolerance, while others in aMycetohabitans rhizoxinicastrain promote bacterial colonization of the fungus. The phenotypic diversity underscores the need to assess the sequence diversity and, given that sequence diversity translates to DNA targeting specificity, the functional diversity of Btl proteins. Using a dual approach to maximize capture of Btl protein sequences for our analysis, we sequenced and assembled nineMycetohabitansspp. genomes using long-read PacBio technology and also mined available short-read Illumina fungal–bacterial metagenomes. We show thatbtlgenes are present across diverseMycetohabitansstrains from Mucoromycota fungal hosts yet vary in sequences and predicted DNA binding specificity. Phylogenetic analysis revealed distinct clades of Btl proteins and suggested thatMycetohabitansmight contain more species than previously recognized. Within our data set, Btl proteins were more conserved acrossM. rhizoxinicastrains than acrossM. endofungorum, but there was also evidence of greater overall strain diversity within the latter clade. Overall, the results suggest that Btl proteins contribute to bacterial–fungal symbioses in myriad ways.

Genetics & Heredity↗

The promising role of proteomes and metabolomes in defining the single-cell landscapes of plants

The plant community has a strong track-record of RNA sequencing technology deployment, which combined with the recent advent of spatial platforms (e.g., 10x genomics), has resulted in an explosion of outstanding single cell and nuclei datasets that can be put in an in situ context within tissues (e.g., a cell atlas)1. In the genomics era, application of proteomics technologies in the plant sciences has always trailed behind that of RNA sequencing technologies, largely due to accessibility, ease-of-use and access to expertise along with depth of analysis benefits. On the other hand, the use of early analytical tools for characterizing small molecules (metabolites) from plant systems predates nucleic acid sequencing and proteomics analysis2, as the search for plant-based natural products has played a significant role in improving human health throughout history. However, the employment of proteomics and metabolomics assays for characterizing plant cell processes now remains significantly behind transcriptional approaches, even though both provide a direct functional readout of cell states and phenotypes.

Anderton, Christopher R. [BATTELLE (PACIFIC NW LAB↗

Bayesian estimation of HIV acquisition dates for prevention trials

Accurate timing estimates of when participants acquire HIV in HIV prevention trials are necessary for determining antibody levels at acquisition. The Antibody-Mediated Prevention (AMP) Studies showed that a passively administered broadly neutralizing antibody can prevent the acquisition of HIV from a neutralization-sensitive virus. We developed a pipeline for estimating the date of detectable HIV acquisition (DDA) in AMP Study participants using diagnostic and viral sequence data. Using a Bayesian strategy that combines three streams of data (REN [rev/vpu/env/Δnef] sequence, GP [gag/Δpol] sequence, and diagnostic) where their 95% credible intervals overlap based on pre-specified criteria and decision rules. We evaluated the performance of our AMP pipeline using PacBio viral sequence data from 41 participants across two prospective acute HIV acquisition cohort studies, FRESH and RV217, with twice-weekly sampling. These cohort studies enrolled young women in South Africa and men and women in Kenya and Thailand, respectively, with a high likelihood of HIV acquisition. In evaluating performance, “true DDA” was the center of bounds between last-negative and first-positive RNA diagnostic tests (median time 4 days, range 2–7 days); bias was the mean difference between estimated and true DDA. Using diagnostic data alone yielded timing estimates with a bias of 2.4 days and root mean square error (RMSE) of 7.9 days. These results were improved using sequence + diagnostic data (bias 1.5 days, RMSE 6.9 days), as well as by restricting sequence-based estimation to samples from ≤5 weeks post-DDA (bias 0.2 days, RMSE 7.8 days).

59 BASIC BIOLOGICAL SCIENCES↗

Synthetic Biology PacBio/JAWS QC Analysis (PBJ) v3.0

This software was designed as a sequence validation tool for the assembly of synthetic constructs. It analyzes FASTQ files against a list of reference sequences, combining the results from eight sequencing libraries to generate a summary, and the files needed to view the results in the Integrative Genomics Viewer (IGV) application for manual verification. This was developed for FASTQ files generated by PacBio sequencing, but could be used on any FASTQ files that do not have paired end reads. It can be used to analyze one - eight libraries at a time, and assumes that each construct sequence in the reference will be in each pool, however, this is not a requirement. This is used to identify which libraries of pooled sequences contains a perfect match, or fixable match to the reference file. This pipeline uses many freely available open source libraries, the value added is that in our application the steps of the pipeline are defined in Workflow Description Language (WDL) and run through the Cromwell workflow engine in Docker containers, for easy distribution and set up, as well as the user friendly html summary that is generated.

Simirenko, Lisa↗

Small-Signal Stability of Grid-Forming Converters Under Fault Conditions

Threshold virtual impedance (TVI)-based current limiting for grid-forming converters (GFMs) has gained great interest due to its ability to maintain voltage source behaviour during faults. However, sequence component extraction (SCE) and negative-sequence control (NSC) are often overlooked in small-signal stability assessments during faults. This paper develops small-signal sequence impedance models for GFMs under four well-known SCE methods based on TVI current limiting control during symmetrical fault conditions. Using the developed impedance models, the impacts of SCE and NSC, and the voltage and current control loop bandwidths, on system stability during faults are investigated. Additionally, since negative-sequence TVI (TVI-) is typically added along with its positive-sequence counterpart, which is often inductive, inductive and resistive TVI- are examined. The findings suggest that a higher voltage or current control loop bandwidth has a negative impact on system stability, while SCE and NSC largely reduce the stable range for voltage and current control loop bandwidth during faults, and that the severity of such impacts is determined by the particular SCE method. Furthermore, it is observed that inductive TVI- significantly degrades system stability, while resistive TVI- can enhance stability when suitable SCE methods are appropriately selected and designed. Matlab/Simulink electromagnetic transient simulations validate these analytical results.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces

These data are from Bandopadhyay et al., "Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces". This study aims to understand the soil microbial ecology along terrestrial-aquatic interfaces of a freshwater and estuarine region and how it relates to organic matter. We analyzed soil microbial (16S rRNA gene) and organic matter (Fourier-transform ion cyclotron resonance mass spectrometry, FTICR-MS) composition from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie (freshwater) and Chesapeake Bay (estuarine) regions. This dataset includes 16S rRNA gene amplicon data (only processed file types included here) and organic matter composition from FTICR-MS data (raw and processed files included here) from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie and Chesapeake Bay regions. These sites are part of the COMPASS-FME project (https://compass.pnnl.gov/FME/COMPASSFME). File formats and software needed to access files: 16S rRNA gene amplicon data: These files follow the format reported here https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format#updates-in-v1.0.1. As per this format, there are four file types reported: 1. Taxon tables (also called sequence-by-sample or OTU (operational taxonomic unit)/ESV (exact sequence variant) tables) : available in a .txt file format and accessible using TextEdit or MS Excel. 2. Representative sequences (also called consensus sequences) : available in a .fasta format and accessible using TextEdit. 3. Sequencing metadata : available in a MS Excel workbook file format and CSV file format 4. Bioinformatic metadata : available in a MS Excel workbook file format and CSV file format FTICR-MS data: 1. Raw data converted to a processed file with intensities of the peaks in the given samples : available in a MS Excel CSV file format 2. Processed file used in analyses and visualizations (appended as icr_long_) : available in a MS Excel CSV file format 3. Metadata file for ICR features (appended as icr_meta) : available in a MS Excel CSV file format

54 ENVIRONMENTAL SCIENCES↗