Search NASASearch

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

RB-TnSeq barcode abundance data sets for Novosphingobium aromaticivorans grown on the β-5-linked aromatic dimer dehydrodiconiferyl alcohol

ABSTRACT A randomly barcoded transposon insertion sequencing (RB-TnSeq) library of Novosphingobium aromaticivorans DSM12444 was grown in media containing either glucose or the β-5-linked aromatic dimer dehydrodiconiferyl alcohol (DC-A) as the sole carbon source. The cultures were grown to saturation and then sequenced, yielding the barcode abundance data sets presented here.

Metz, Fletcher

RolyPoly (rp) v0.1.0

The Rolypoly pipeline is designed to process raw RNA-seq data and identify potential RNA viral sequences. It is split into several self contained steps: 1. input data filtering and QC, 2. Genome assembly and refinement, 3. Assembly filtering, 4. Mapping to known RNA viral genomes, 5. Searching for RNA viral marker genes. 6. Genome functional and structural annotation. 6. Report preparation and potential downstream analysis The last module, may include taxonomic assignment, host range estimation, and phenotypic prediction. There are many similar software, but they focus on human related viruses, and lack the downstream applications or differ in their sensitivity. The initial user base are non-computational microbial ecologists who wish to better understand the potential RNA viruses in their own generated samples.

Neri, Uri

Methane-cycling microbial communities from Amazon floodplains and upland forests respond differently to simulated climate change scenarios

Seasonal floodplains in the Amazon basin are important sources of methane (CH 4 ), while upland forests are known for their sink capacity. Climate change effects, including shifts in rainfall patterns and rising temperatures, may alter the functionality of soil microbial communities, leading to uncertain changes in CH 4 cycling dynamics. To investigate the microbial feedback under climate change scenarios, we performed a microcosm experiment using soils from two floodplains (i.e., Amazonas and Tapajós rivers) and one upland forest. We employed a two-factorial experimental design comprising flooding (with non-flooded control) and temperature (at 27 °C and 30 °C, representing a 3 °C increase) as variables. We assessed prokaryotic community dynamics over 30 days using 16S rRNA gene sequencing and qPCR. These data were integrated with chemical properties, CH 4 fluxes, and isotopic values and signatures. In the floodplains, temperature changes did not significantly affect the overall microbial composition and CH 4 fluxes. CH 4 emissions and uptake in response to flooding and non-flooding conditions, respectively, were observed in the floodplain soils. By contrast, in the upland forest, the higher temperature caused a sink-to-source shift under flooding conditions and reduced CH 4 sink capability under dry conditions. The upland soil microbial communities also changed in response to increased temperature, with a higher percentage of specialist microbes observed. Floodplains showed higher total and relative abundances of methanogenic and methanotrophic microbes compared to forest soils. Isotopic data from some flooded samples from the Amazonas river floodplain indicated CH 4 oxidation metabolism. This floodplain also showed a high relative abundance of aerobic and anaerobic CH 4 oxidizing Bacteria and Archaea. Taken together, our data indicate that CH 4 cycle dynamics and microbial communities in Amazonian floodplain and upland forest soils may respond differently to climate change effects. We also highlight the potential role of CH 4 oxidation pathways in mitigating CH 4 emissions in Amazonian floodplains.

16S rRNA sequencing

Bleach Rescues Nannochloropsis from an Obligate Parasite and Alters Microbial and Metabolite Signatures of Outdoor Cultures

Chemical agents are commonly used to protect algal crops. Yet, few studies have characterized the effects of these agents on associated microbial communities to understand effects on microbial functions relevant to algal crop production and protection. Here, we used shotgun metagenomic sequencing and untargeted exometabolite profiling to link the application of bleach, a -cidal agent used to protect algae from pests, to changes in community composition, metabolic pathways, and exometabolies - at a whole community level. Bleach protected the algal crop from crashing but altered bacterial diversity. Analysis of metagenome-assembled genomes (MAGs) revealed a classic predator-prey cycle between Oligoflexus and our target alga Nannochloropsis. Olifoflexus genomes from our study were notably similar to a previously identified BALO (Bdellovibrio and like organism), FD111, known to kill Nannochloropsis cultures, providing strong evidence that an FD111-like organism was responsible for the crash. Metabolic pathway composition differed between bleached and unbleached ponds, with abundance of twelve pathways related to stress tolerance, including the superpathway of methylglyoxal degradation, lipid IVA biosynthesis, and ectoine biosynthesis, greater in bleached ponds compared to unbleached ponds. Virulence factors related to adherence, biofilm formation, motility, and pathogenicity increased dramatically in bleached ponds with time, although this increase was not coupled with an increase in pathogens - algal or otherwise - or a decline in algal health. Our study highlights the importance of coupling 16S rRNA gene sequencing with whole genome data and other -omics tools to sketch a larger picture of community structure and function in crop systems. Moreover, our results highlight that continued long-term bleaching may lead to negative effects to crop health or downstream adverse health effects to humans or animals, depending on the algal product (i.e. human supplements or animal feedstocks). Future work on alternative treatment methods that would reduce resistance is necessary in the field.

09 BIOMASS FUELS

Model of metabolism and gene expression predicts proteome allocation in Pseudomonas putida

Abstract The genome-scale model of metabolism and gene expression (ME-model) forPseudomonas putidaKT2440,iPpu1676-ME, provides a comprehensive representation of biosynthetic costs and proteome allocation. Compared to a metabolic-only model,iPpu1676-ME significantly expands on gene expression, macromolecular assembly, and cofactor utilization, enabling accurate growth predictions without additional constraints. Multi-omics analysis using RNA sequencing and ribosomal profiling data revealed translational prioritization inP. putida, with core pathways, such as nicotinamide biosynthesis and queuosine metabolism, exhibiting higher translational efficiency, while secondary pathways displayed lower priority. Notably, the ME-model significantly outperformed the M-model in alignment with multi-omics data, thereby validating its predictive capacity. Thus,iPpu1676-ME offers valuable insights intoP. putida’s proteome allocation and presents a powerful tool for understanding resource allocation in this industrially relevant microorganism.

Mathematical & Computational Biology

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity

RhizoGrid Indexed Sorghum Rhizosphere Multi-Omics

PerCon SFA project data dentification of spatially resolved biomarkers of drought in Sorghum bicolor rhizosphere molecular-microbe interactions using a novel root cartography "RhizoGrid" system for sampling plants under drought and control conditions across 10 equally sized root zone environments (4 quadrants each). Each quadrant was sampled and processed for 16S amplicon, metabolomics, and X-ray computed tomography (XCT). Data download includes experimental metadata and results files for 16S rRNA sequence analysis of microbial community assembly (processed data files), liquid chromatography mass spectrometry (LC-MS) metabolomics analysis of microbial community root exudates (processed data files), X-ray computed tomography (XCT) spatial gradient analysis (raw and processed data files) of microbial community composition, and related computational modeling outputs.

59 BASIC BIOLOGICAL SCIENCES

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 2. Evaluating Controls on Flow Persistence in an Urbanized Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in an urbanized catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, distributed temperature sensing (DTS), continuous self-potential (SP) monitoring, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Field_Application subfolder contains the ATS XML input scripts, data files, output data for the SP site. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. The flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.m can only be used with COMSOL with MATLAB) is executed using the ATS output data to simulate the potential field. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) DTS Contains collated DTS data including raw Stokes and anti-Stokes measurement (provided as .h5 file). It also includes DTS processing.ipynb, a Jupyter notebook for calibrating the DTS data using dts_calibration Python package. cooler_calibration.csv is the DTS calibration CSV used in the calibration sequence. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion. 6) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 7) SP Contains the SP data collected in field at the SP sites (provided as CSV files). 8) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). Note: Code files (.ipynb, .py, .xml) can be opened in any standard code editor, .exo file can be viewed using Paraview, .h5 files can be opened using HDFView software and h5py Python package, and .resipy file can be opened with the open-source ResIPy software.

ATS

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 1. Evaluating Above- and Below-ground Controls of Flow Persistence in a Forested Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in a forested catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, ground penetrating radar (GPR), continuous self-potential (SP) monitoring, electromagnetic (EM) imaging, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Contains two subfolders: Synthetic and Field_Application subfolder. Synthetic subfolder contains the ATS XML input script (can be opened using any code editor) for the four synthetic hydrological cases tested (Connected and gaining, Connected and losing, Disconnected and losing, and dry stream). It also includes other experimental cases to test the influence of precipitation and concentration gradient. For each synthetic case, the flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.mph can be opened with the commercial software COMSOL and requires a license) is executed using the ATS output data to simulate the potential field. It also includes the Synthetic_model_plot.ipynb (can be opened using any code editor) to visualize the SP result and generate manuscript figures. The data subfolder contains mesh files to run both the ATS (.exo and .stl files can be viewed using Paraview; .h5 files can be opened using HDFView software and h5py Python package) and COMSOL models. Field_Application subfolder contains two subfolders: ES_MDA_inversion and Final_Model. ES_MDA_inversion contains the Python script (.py can be opened using any code editor) and SP observation data used to run the Ensemble Smoother with Multiple Data Assimilation (ES-MDA) inversion sequence to get the optimal model parameters. The Final_model subfolder contains the ATS XML input scripts, data files, output data for the two SP sites. The same workflow steps outlined for the Synthetic subfolder apply here. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) EM Contains the CSV file of the EM data from the DUALEM-42, including spatial coordinates (x, y, z), apparent conductivity, and in-phase measurements at 2 m coil separations for horizontal coplanar (HCP) and perpendicular (PRP) geometries. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion (.resipy can be opened with the open-source ResIPy software). 6) GPR Includes GPR field datasets collected at 100 MHz and 250 MHz antenna frequencies, along with the processing/interpretation project file (GPR_process.gpz can be viewed using EKKO_Project 6, a commercial software by Sensors & Software that requires a license). 7) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 8) SP Contains the SP data collected in field at the two SP sites (one in the perennial reach and the other in the intermittent reach), provided as DAT files. 9) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). It also includes DTW.ipynb, a Jupyter notebook containing the code for the dynamic time warping (DTW) with sliding window to evaluate SP signal synchronicity.

ATS

Recent Progress on Surface Water Quality Models Utilizing Machine Learning Techniques

Surface waterbodies are heavily exposed to pollutants caused by natural disasters and human activities. Empowering sensor technologies in water quality monitoring, sufficient measurements have become available to develop machine learning (ML) models. Numerous ML models have quickly been adopted to predict water quality indicators in various surface waterbodies. This paper reviews 78 recent articles from 2022 to October 2024, categorizing water quality models utilizing ML into three groups: Point-to-Point (P2P), which estimates the current target value based on other measurements at the same time point; Sequence-to-Point (S2P), which utilizes previous time series data to predict the target value at one time point ahead; and Sequence-to-Sequence (S2S), which uses previous time series data to forecast sequential target values in the future. The ML models used in each group are classified and compared according to water quality indicators, data availability, and model performance. Widely used strategies for improving performance, including feature engineering, hyperparameter tuning, and transfer learning, are recognized and described to enhance model effectiveness. The interpretability limitations of ML applications are discussed. This review provides a perspective on emerging ML for surface water quality models.

machine learning (ML)

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES

Assessment of Accelerated Stress Testing Data for Silicon Photovoltaics Using Tensor Decomposition Methods

In this work, we examine the use of high-order tensor decompositions to analyze degradation pathways emerging from accelerated stress testing of silicon photovoltaic (PV) modules. Matrix-based decompositions are powerful tools for studying two-dimensional data arrays and form the foundation of a host of classical data analysis techniques. Tensors are high-order extrapolations of matrices that are able to account for more parameter dimensions, and a variety of tensor decomposition methods have been developed that similarly seek to extend insights from matrix decompositions to higher dimensions. Applying and interpreting tensor decomposition methods to sequences of PV module image data, we seek to uncover and isolate different degradation modes occurring from accelerated stress testing procedures. Further, we consider the contributions of different modes to PV module performance degradations.

data analysis

Short-period post-common envelope binaries with Balmer emission from SDSS and LAMOST based on ZTF photometric data

ABSTRACT We present here 55 short-period post-common envelope binaries (PCEBs) containing a hot white dwarf (WD) and a low-mass main sequence (MS). Based on the photometric data from Zwicky Transient Facility survey data Release 19 (ZTF DR19), the light curves are analysed for about 200 WDMS binaries with emission line(s) identified from the Sloan Digital Sky Survey (SDSS) or the Large Sky Area Multi-Object Fibre Spectroscopic Telescope (LAMOST) spectra, in which 55 WDMS binaries are found to exhibit variability in their luminosities with a short period and are thus short-period binaries (i.e. PCEBs). In addition, it is found that the orbital periods of these PCEBs locate in a range from 2.2643 to 81.1526 h. However, only six short-period PCEBs are newly discovered and the orbital periods of 19 PCEBs are improved in this work. Meanwhile, it is found that three objects are newly discovered eclipsing PCEBs, and a object (i.e. SDSS J1541) might be the short-period PCEB with a late M-type star or a brown dwarf companion based on the analysis of its spectral energy distribution. At last, the mechanism(s) being responsible for the emission features in the spectra of these PCEBs are discussed, the emission features arising in their optical spectra might be caused by the stellar activity or an irradiated component owing to a hot WD companion because most of them contain a WD with an effective temperature higher than $\sim$10 000 K.

Li, Lifang

Integration of ultra-low coverage whole-genome sequences for reconstructing the evolutionary history of Galapagos giant tortoises

Genomic data from contemporary and historical samples often need to be coupled for evolutionary reconstructions of multitaxon complexes. However, the genetic data recovered from historical samples may result only in ultra-low coverage whole-genome sequences (ulcWGS; <0.15× depth), leading to inaccurate evolutionary inferences given a preponderance of missing data. Using the Galapagos giant tortoise radiation as a study system (Chelonoidis spp., composed of 13 extant and four extinct lineages), we assembled a novel methodological pipeline that removes potential noise introduced by the missing data and enhances the evolutionary signal from ulcWGS samples. We leveraged existing tools for phylogenomic placement (EPA-ng), population genomic structure (smartsnp) and admixture (Admixfrog, NGSadmix) to demonstrate that the evolutionary history of samples can be uncovered with sequencing depths as low as 0.008–0.139×. Importantly, these approaches do not use genotype imputation of the ulcWGS samples, which would require extensive reference datasets. Our application to two cases of extinct lineages of Galapagos giant tortoises, with and without references from the same lineage, demonstrates the general value of the approach. We confirm where the extinct lineages from San Cristóbal and Santa Fe islands fit into the Galapagos giant tortoise radiation, and that these lineages were evolutionarily distinct entities.

ancient DNA

Untargeted, tandem mass spectrometry (LC/MS-MS) metaproteomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization. This package contains soil metaproteomics data in the context of site specific metagenomes from soil depth profiles in three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. These metaproteomes were collected in 2018 after 4.5 years of warming from five depth intervals (0-10 cm, 10-30 cm, 30-45 cm, 45-60 cm, 60-80 cm). For protein identification, the collected spectra were searched following a target-decoy search strategy against a database of metagenome predicted proteins (covering 96 samples from 2014 to 2021) representing the complete sequence diversity at the site. Data was searched with mass spectrometry database search tool (MS-GF+) using Pacific Northwest National Laboratory (PNNL)'s Data Management System (DMS) Processing pipeline. The metagenomes are published as part of another data package. Raw metaproteomic data and the data products from MS-GF+ are deposited in the Mass Spectrometry Interactive Virtual Environment (MassIVE) database under accession no. MSV000097826. Here we present a dataset that includes spectral counts for the detected proteins across samples (EMSL50964_BrodieAllMAGs_Globals_SC.txt), the sequences of the detected proteins, and sample metadata file that contains site information for the soil metaproteome samples.

Belowground Biogeochemistry Science Focus Area

A practical approach to using the Genomic Standards Consortium MIxS reporting standard for comparative genomics and metagenomics

Comparative analysis of (meta)genomes necessitates aggregation, integration, and synthesis of well-annotated data using standards. The Genomic Standards Consortium (GSC) collaborates with the research community to develop and maintain the Minimal Information about any (x) Sequence (MIxS) reporting standard for genomic data. To facilitate use of the GSC’s MIxS reporting standard, we provide a description of the structure and terminology, how to navigate ontologies for required terms in MIxS, and demonstrate practical usage through a soil metagenome example.

standards, metadata, genome, metagenome, schema, v

Impedance Scan of Inverter-Based Resources and Diesel Generator for Stability Analysis: Preprint

Impedance-based methods are widely used for power system stability analysis with inverter-based resources (IBRs), e.g., assessing dynamic interactions between the power grid and an IBR, control interactions between multiple IBRs, and the sub-synchronous oscillation and damping phenomenon. Since it is difficult to get a numerical model 100% matching with the hardware IBR, using the hardware inverter directly to obtain its output impedance has become a prominent approach nowadays. Therefore, this article presents the impedance scan using hardware IBRs, and also a hardware diesel generator as it still stays with the grid before the grid completely goes to renewable. The devices under test (DuTs) for the impedance scan includes two 3-..phi.., 480 V, 60 Hz commercial grid-forming IBRs (one of 250 kVA and another of 125 kVA rating) in series with ..delta..-Y transformers, one 3-..phi.., 480 V, 60 Hz commercial grid-following IBR (of 125 kVA rating), and a 3-..phi.., 480 V, 60 Hz commercial diesel generator (of 187.5 kVA rating). Using voltage signals perturbed with sub-, inter-, and higher harmonic components, and measuring the current response, the positive-sequence impedances are computed via an offline- based post-analysis. Moreover, best-fit transfer functions are estimated that closely resemble the measured data points of the positive-sequence impedances. Based on the observations from various outcomes of the hardware experiments, this article also provides some fundamental insights on the equivalent positive- sequence impedance of a combination of multiple hardware components by comparing the estimated and the empirically computed impedances. A comparative insight on the damping capability of the DuTs using the positive-sequence impedances of the hardware is also discussed.

grid following inverter

Machine learning for reactor power monitoring with limited labeled data

Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND