Search NASA⌕ Search

SEARCH · Search NASA

Results for “Software engineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Obstacles to Practical Digital Supply Chain Risk Management in the Energy Sector

Cyber supply chain risk management (C-SCRM) programs must consider operations that depend on the lifecycles of digital components such as hardware, firmware, software, and services. We integrate academic literature, historical incidents, and existing standards to identify obstacles faced by C-SCRM programs.

Business Process Management & Integration↗

WaterTAP 1.0 Release

The Water treatment Technoeconomic Assessment Platform (WaterTAP) is an open-source Python-based software package that supports the simulation and optimization of process-scale water treatment trains. WaterTAP seeks to provide the broader water research community with an integrated modeling capability to evaluate cost, energy, and environmental tradeoffs across water treatment options and identify high impact opportunities for innovation including novel materials, processes, and systems. An updated version of WaterTAP is released quarterly and each includes documentation and release notes.

AS↗

Carbon-13 NMR spectra of lignin isolated from field grown transgenic poplar

Here we present a curated dataset of a series of 13C nuclear magnetic resonance (NMR) spectra of lignin isolated from transgenic monolignol 4-O-methyltransferase (MOMT4) engineered poplar. The transgenic poplar was collected from a 3-year field trial experiment. The poplar was Soxhlet-extracted with toluene/ethanol and the extractives-free poplar was then ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h. The ball-milled materials were then subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96:4 (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover the lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide for NMR characterization. 13C experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence (zgpg) on a Prodigy platform cryoprobe. The NMR spectra were acquired under the following conditions: spectra width 229 ppm, 64k data points, 1s pulse delay, and 6k scans. All the data was processed using the Bruker’s TopSpin 3.6 software. Additional meta data is embedded in the raw spectra files.

13C NMR, lignin, poplar, field trial, MOMT4, CBI↗

Web-Based Tools for Data-Informed Remedy Optimization: Software Theory and User Guide

This report documents the development and application of two web-based decision-support tools for pump-and-treat (P&T) groundwater remediation systems: PTOLEMY (Pump-and-Treat Optimized Location Evaluation to Maximize Yields) and OPTIMA (Optimization for Pump-and-Treat Implementation, Management, & Assessment). These tools enhance remedy design and management by leveraging advanced computational methods – specifically deep learning and multi-objective optimization – within a user-friendly platform. By integrating data-driven models with established hydrogeological knowledge, PTOLEMY and OPTIMA enable more efficient evaluation of well placement and operational strategies, helping site managers balance multiple remediation objectives under complex conditions. Both tools are implemented as modules within the SOCRATES (Suite Of Comprehensive Rapid Analysis Tools for Environmental Sites) web platform, which provides data access, visualization, and analytics to support remedy optimization across sites in the U.S. Department of Energy Office of Environmental Management complex. PTOLEMY is a rapid screening module designed to identify promising locations for new extraction wells. It employs a multi-channel three-dimensional convolutional neural network (MC3D-CNN) trained on high-fidelity simulation data to predict the relative performance (in terms of contaminant mass recovery) of potential well sites. Through an interactive web interface, PTOLEMY visualizes the probability of high performance across a site, highlighting areas where an extraction well is likely to yield above-threshold contaminant removal over a multi-year period. PTOLEMY’s map-based displays and exportable results support transparent communication of screening analyses. By focusing attention on the most favorable candidate locations, the tool augments traditional engineering judgment and physics-based modeling, providing a data informed basis for subsequent detailed evaluations. OPTIMA is a multi objective optimization module designed to find wellfield layouts and operating schedules that meet various cleanup goals. It quickly evaluates thousands of candidate setups – combinations of well locations, timing, and rates – and returns a small set of best trade-off options for comparison. At its core, OPTIMA uses a U-Net-based surrogate model – a deep-learning emulator of a groundwater flow and transport simulator – to dramatically accelerate scenario evaluations. Coupling this fast surrogate with the NSGA-II (Non-dominated Sorting Genetic Algorithm II) evolutionary algorithm, OPTIMA explores a wide decision space of well locations and schedules to identify Pareto-optimal solutions that trade off key objectives (e.g., minimizing cleanup time, maximizing contaminant mass removal, and minimizing plume extent). The tool outputs a family of optimal configurations and visualizes their trade-offs (Pareto frontiers of cleanup metrics and maps of optimized well placements). Site managers can use these results to understand the range of viable strategies and to select candidate designs for more detailed verification. OPTIMA is currently under active development and not yet fully released; this guide provides early documentation to support planning and gather user feedback.

54 ENVIRONMENTAL SCIENCES↗

High-performance finite elements with MFEM

The MFEM (Modular Finite Element Methods) library is a high-performance C++ library for finite element discretizations. MFEM supports numerous types of finite element methods and is the discretization engine powering many computational physics and engineering applications across a number of domains. Furthermore, this paper describes some of the recent research and development in MFEM, focusing on performance portability across leadership-class supercomputing facilities, including exascale supercomputers, as well as new capabilities and functionality, enabling a wider range of applications. Much of this work was undertaken as part of the Department of Energy’s Exascale Computing Project (ECP) in collaboration with the Center for Efficient Exascale Discretizations (CEED).

97 MATHEMATICS AND COMPUTING↗

Comparative analysis of thermal management systems in electric vehicles at extreme weather conditions: Case study on Nissan Leaf 2019 Plus, Chevrolet Bolt 2020 and Tesla Model 3 2020

With the surge in electric vehicle (EV) adoption and the need for extended driving ranges, optimizing energy efficiency, particularly through thermal management, is critical, especially in extreme weather. Managing the substantial energy needed for cabin climate control and battery temperature regulation can increase energy demands by over 50 %, severely limiting range. This study conducts a comparative analysis of thermal management systems (TMS) in three popular EV vehicles, 2020 Chevrolet Bolt, 2019 Nissan Leaf Plus, and 2020 Tesla Model 3, evaluating their distinct TMS configurations and performance under varied weather conditions. Using both numerical simulations and experimental data collected on a controlled test bench at Argonne National Laboratory, we assess how TMS architecture and operational modes influence energy consumption and range. A comprehensive TMS model was developed, integrating cabin and battery thermal sub-models in the Autonomie software platform, to simulate temperature fluctuations and range impacts. Cabin climate was modeled using a mono-zonal approach, while battery cell temperature distribution was estimated through a 2D nodal structure. Each vehicle's distinct TMS setup was evaluated: the Chevrolet Bolt and Tesla Model 3 use a dual evaporator vapor compression cycle with a PTC heater for the cabin and a coolant loop for battery thermal management; the Nissan Leaf Plus employs a heat pump with a PTC heater for the cabin and air-cooling for the battery. Tests conducted at ambient temperatures of 35°C, 22°C, -7°C, and -18°C reveal significant differences in energy use and range reduction across both configurations and conditions. At 35°C, the Tesla Model 3, Chevrolet Bolt, and Nissan Leaf Plus have a range reduction of 8%, 9%, and 13%, respectively, due to air conditioning. In winter, heating technology is paramount; at -7°C, the Nissan Leaf's heat pump configuration achieves a lower range reduction (19.3%) compared to the Tesla and Chevrolet Bolt PTC heaters, which reduce range by 28.3% and 31%, respectively. Further, this study provides valuable insights for automotive engineers, EV technology researchers, and thermal management system designers aiming to enhance electric vehicle performance by understanding how different weather conditions and TMS architectures impact energy consumption and driving range.

33 ADVANCED PROPULSION SYSTEMS↗

Spin-Controllable Dynamics in Defect-Engineered Carbon Nanotubes as Single Photon Emitters: Data-Driven Modeling and Computations

Quantum technologies, such as quantum computing and sensing, require efficient single-photon emission (SPE) sources that operate at room temperature in telecom wavelengths. While several materials can serve as SPE sources, no single platform meets all the criteria for efficiency, ambient operation, and scalability. Single-walled carbon nanotubes (SWCNTs) with covalently attached molecules offer a promising solution. Their SPE can be easily tuned via modifications of the SWCNT's diameter, chirality, and bonded molecules, enabling emission across near-IR to telecom wavelengths at ambient conditions. However, to fully realize the potential of SWCNTs and unlock their quantum capabilities, a deeper understanding of how structural defects from molecular adducts affect their emission and competing photoexcited processes is essential. To address this gap in our knowledge, this project combined quantum chemistry calculations with data-driven methods of cheminformatics (QSAR) and machine learning (ML). The developed computational approaches have provided several design strategies for covalent functionalization of SWCNTs to improve their optical response. The collaboration with Los Alamos National Lab (LANL) enabled direct comparison of computational and experimental data, facilitating method validation. This partnership was enhanced through access to LANL's Center for Integrated Nanotechnologies (CINT) utilizing User Facility Program and summer internships, which provided three NDSU graduate students with hands-on experience at LANL. The outcomes of this project included (1) Advancing the current stage of computational methods in accurate modeling of non-adiabatic spin-dependent photoexcited dynamics and its applicability to nanosystems consisting of thousands of atoms, realized as open-access codes linked to existing DFT-based software; (2) Establishing the relationship between the structure of adducts and SWCNTs and intrinsic excitonic and spin properties of defect states for guiding novel synthetic strategies and experimental probes of chemically functionalized SWCNTs as near-IR emitting materials; (3) Generating virtual libraries of hypothetical functionalized SWCNTs for virtual screening of their chemical structures and optical properties, leveraging new functionalities of SWCNTs; (4) Offering a unique experience for NDSU graduate students that prepared them for future scientific careers related to materials modeling and big data processing. These results were summarized in 12 published journal papers and 3 recently submitted papers. One of a key finding is that the position of defect sites on the SWCNT surface primarily drives the emission redshift (up to 100 meV), while the polarity of the defect-inducing molecules has a much smaller effect (~10 meV). However, the electron-donating or withdrawing properties of a molecule influence selecting reactivity of defect sites. These insights important for optimizing synthetic protocols for desired emissions in SWCNTs. We also revealed that the interaction between two defects at various positions on the SWCNT enhances the redshift and optical activity of states, favoring strong near-IR emission. This suggests that manipulations in defect concentrations is a promising strategy for controlling efficient emission. Mostly important, the defect position was found controllable by the spin states of photoexcited intermediates: Excited aromatic molecules form ortho defects with SWCNTs at their singlet states in the presence of oxygen, while oxygen-free conditions favor para defects via the triplet-state mechanism. Additionally, a heat-activated [2+2] cycloaddition reaction facilitates divalent defect formation with fewer bonding positions that narrows emission bands. These groundbreaking findings have been experimentally validated and significantly advance our understanding of defect chemistry in SWCNTs. Using a novel encoding technique and 3D-MoRSE descriptors, we developed highly accurate ML/QSAR models to predict both the 3D structure and optical properties of SWCNTs with chemical defects. This model enabled the creation of a virtual library of 125,556 structures, providing new insights into the relationship between SWCNT-defect structure and emission.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Electric Drive Technologies Research: ELT223 Component Modeling, Co-Optimization, and Trade-Space Evaluation Annual Report

This project is intended to support the development of new traction drive systems that meet the targets of 100 kW/L for power electronics and 50 kW/L for electric machines with reliable operation to 300,000 miles. To meet these goals, new designs must be identified that make use of state-of-the-art and next-generation electronic materials and design methods. Designs must exploit synergies between components, for example converters designed for high-frequency switching using wide band gap (WBG) devices and ceramic capacitors. This project included: (1) a survey of available technologies; (2) investigating new technologies, that for example, reduce volume of thermal management or magnetic components; (3) the development of computer aided design tools that consider the converter volume, reliability, and electrical performance; (4) exercising the design software to evaluate performance gaps and predict the impact of certain technologies and design approaches, i.e. GaN semiconductors, ceramic capacitors, ceramic thermal management components, and select topologies; (5) building and testing hardware prototypes to validate models and concepts. The design tools enable co-optimization of the power module and passive elements and provide some design guidance. At the end of the project, new advanced computing methods, such as machine learning approaches, were considered.

33 ADVANCED PROPULSION SYSTEMS↗

SynBio QC Dual Barcode QC (DBC) v1.0

This software was designed as a sequence validation tool for the assembly of synthetic constructs, where the constructs have a high degree of similarity and thus are barcoded prior to the sequencing library prep. It demultiplexes each FASTQ file for each barcode, then analyzes the resulting FASTQ files against a list of reference sequences for that barcode/library, combining the results from eight sequencing libraries to generate a summary, and the files needed to view the results in the Integrative Genomics Viewer (IGV) application for manual verification. This was developed for FASTQ files generated by PacBio sequencing, but could be used on any FASTQ files that do not have paired end reads. It can be used to analyze one - eight libraries at a time. Each construct is independently analyzed with only the sequences with the same barcode, in the same pooled library. Then the results are combined into a user friendly summary. This is used to identify which libraries of pooled sequences contains a perfect match, or fixable match to the reference file. This pipeline uses many freely available open source libraries, the value added is that in our application the steps of the pipeline are defined in Workflow Description Language (WDL) and run through the Cromwell workflow engine in Docker containers, for easy distribution and set up, as well as the user friendly html summary that is generated.

Simirenko, Lisa↗

Impact of the Exciter and Governor Parameters on Forced Oscillations

In recent years, the frequency of forced oscillation events due to control system malfunctions or improper parameter settings has increased. Tuning the parameters of exciters and governor models is crucial for maintaining power system stability. Traditional simulation studies typically involve small transient disturbances or step changes to find optimal parameter sets, but existing optimization algorithms often fall short in fine-tuning for forced oscillations. Identifying the sensitive parameters within these control models is essential for ensuring stability during large, sustained disturbances. This study focuses on identifying these critical exciter and governor model parameters by analyzing their influence on sustained forced oscillations. Using Kundur’s two-area system, we analyze common exciter models such as SCRX, ESST1A, and AC7B, along with governor models like GAST, HYGOV, and GGOV1, utilizing PSS®E software version 34. Sustained forced oscillations are injected at generator-1 of area-1, with individual parameter changes dynamically simulated. By considering a local oscillation frequency of 1.4 Hz and an inter-area oscillation mode of 0.25 Hz, we analyze the impact of each parameter change on the magnitude and frequency of forced oscillations as well as on active and reactive power outputs. This novel approach highlights the most influential parameters of each tested model—such as exciter, governor, and turbine gains, as well as time constant parameters—on the impact of forced oscillations. Based on our findings, the sensitive parameters of each tested model are ranked. These would provide valuable insights for industry operators to fine-tune control settings during oscillation events, ultimately enhancing system stability.

42 ENGINEERING↗

PARETO UI 1.1.0 Release

PARETO is an open-source Python-based software package for oilfield produced water management and beneficiary reuse optimization. PARETO supports produced water industry by providing cost-effective water management solutions. This version introduced an updated User Interface (UI) which makes it easier to navigate and understand the solution for industry users. New Features: - Map files are added for visualization - Added output export function button - Water residual view added - Workflow was streamlined - File extension was expanded - Minor bugfix

AS↗

Investigating Temperature Uniformity and Accuracy in PV Module Lamination: A Verification Study

This study investigates the temperature uniformity and accuracy of a photovoltaic (PV) module lamination process by addressing inconsistencies identified in 2017 data where irregular temperature changes were observed across setpoints. The 2017 data showed a notable drop in temperature upon bladder initiation, except for the 145 degrees Celsius profile. This inconsistency indicated potential inaccuracies in manual data recording methods. To address this concern, a verification experiment was conducted to evaluate temperature uniformity across the 2014 Bent River SPL2828 laminator platen and within test samples. Thermocouples, paired with Omega data acquisition software, were deployed to measure temperatures at multiple platen locations and within test samples. The experiment compared lamination temperatures of polyethylene-co-vinyl acetate (EVA) encapsulant when paired with solite glass or TPE backsheets. The methodology included verifying temperature uniformity directly on the platen and by using a large glass/EVA/glass sample using multiple thermocouples. Smaller samples were built with glass/EVA/glass and glass/EVA/backsheet configurations with one centered thermocouple to verify and compare sample temperatures. This verification aims to refine lamination temperature profiles, enhance data accuracy and provide insights into optimal process control for uniform module lamination. Ensuring consistent and uniform lamination may improve the accuracy and reliability of research outcomes.

14 SOLAR ENERGY↗

The Fluid Dynamics Uncertainty Quantification Challenge Problem: XFOIL vs. MFOIL

Uncertainty quantification (UQ) has become more critical in aerospace engineering due to the growing dependence on computational tools for design optimization and performance analyses of aerospace vehicles. Even though the significance of UQ in assessing the credibility of computational analyses is well recognized, its costs and complexity impede its integration into standard practices, particularly in computational fluid dynamics (CFD) and other fluid analyses. This paper presents a UQ study for low-fidelity computational aerodynamics analyses with XFOIL and mfoil (i.e., the MATLAB version of XFOIL with several implementation modifications); these tools are utilized widely in both research and education. The main contributions of this paper are as follows: 1) improved precision in quantifying the uncertainty of the baseline Monte Carlo results used to benchmark surrogate modeling techniques for UQ, 2) quantification of the effect of the implementation differences between XFOIL and mfoil on solution quantities of interest (QoIs), such as lift and pitching moment coefficients, and 3) development of an open-source UQ library for use with XFOIL and mfoil, which has educational values and helps promote UQ for fluid analyses with aerospace applications. Results and discussions revolve around cases 1-4 of the challenge problem posed by the AIAA Fluid Dynamics Technical Committee’s Uncertainty Quantification Discussion Group (UQDG). In case 3, this work employs CFDverify, an open-source solution verification software, to quantify the discretization error and evaluate the extrapolated QoIs based on the grid convergence index (GCI). This UQ study differentiates itself from previous studies in the rigor of handling baseline Monte Carlo uncertainty and in including mfoil, which is a more accessible alternative to XFOIL. Finally, despite the growing computing power, low-fidelity computational tools remain valuable, such as for aerodynamic shape optimization at Mach numbers below 0.65 and low-to-mid Reynolds numbers.

Lay, Aidan S [University of Tennessee, Knoxville (↗

Proton NMR spectra of lignin isolated from field grown transgenic poplar

Here we present a curated dataset of a series of 1H nuclear magnetic resonance (NMR) spectra of lignin isolated from transgenic monolignol 4-O-methyltransferase (MOMT4) engineered poplar. The transgenic poplar was collected from a 2-year-old rotation trees within a three-year field trial experiment. Two replicates were collected for each transgenic poplar for the 1H NMR analysis. The poplar samples were Soxhlet-extracted with toluene/ethanol to remove the extractives and the extractives-free poplar was then ball-milled in a Retsch PM100 planetary ball mill using a porcelain jar with ceramic balls at 600 rpm for 2 h. The ball-milled materials were subjected to enzymatic hydrolysis for 48 h followed by centrifugation and washing with deionized water. The solid residue was extracted twice with 96:4 (v/v) 1,4-dioxane/water mixture at room temperature overnight. The extracts were combined, rotary evaporated, and freeze-dried to recover lignin. The dry lignin samples were dissolved in deuterated dimethyl sulfoxide and transferred into a 5 mm NMR tube. 1H NMR experiments were performed in a Bruker Avance III HD 500 MHz NMR spectrometer operating at a frequency of 125.12 MHz for the 13C nucleus using a standard Bruker pulse sequence (zg) on a Prodigy platform cryoprobe. The NMR spectra were acquired with 16 ppm spectra width, 32k data points, 3s pulse delay, and 16 scans. All the data was processed using the Bruker’s TopSpin 3.6 software. Additional meta data is embedded in the raw spectra files.

1H NMR, lignin, poplar, field trial, MOMT4, CBI↗

Integrating PCTRAN with AI-Driven Host-Intrusion Detection and Secured Container Systems for Advanced Malware Analysis (Summer Internship Report)

This study presents a solution for enhancing the security of the Personal Computer Transient Analyzer (PCTRAN) PC-based Nuclear Power Plant Simulator by integrating the software with an artificial intelligence (AI)-driven host-intrusion detection system (HIDS), in addition to a secured container system, for malware analysis. PCTRAN is a Windows XP-based software package that has the ability to simulate a variety of accident and transient conditions for nuclear power plants (NPPs). It offers a high-resolution replica of the Nuclear Steam Supply System (NSSS) and displays the status of important parameters allowing for operator interaction. By including AI-driven HIDS for the NSSS, the framework can identify security threats in real-time, ensuring the integrity of the nuclear simulation environment. Additionally, the secured container system offers the ability to isolate and analyze malware, preventing potential threats from affecting core systems. The integration process involves extensive testing and validation in order to ensure accuracy, reliability, and compliance with security policies. This framework sets a new precedent for secure simulation and training in NPP operations, and offers insight for future advancements in cybersecurity.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Asynchronous GPU-based DEM solver embedded in commercial CFD software with polyhedral mesh support

A novel graphical processing unit-based discrete element method solver is introduced to improve stability, performance, and provide seamless integration into commercial or open-source computational fluid dynamics software. A key innovation is eliminating a need for network communication between solvers, which was previously required for cross-platform coupling. This is accomplished by a direct coupling method that employs dynamic-linked libraries. Furthermore, the solver optimizes memory usage by streamlining the particle-cell search algorithm by eliminating the cells' searching grid. This ensures the solver is compatible with a wide range of mesh types, providing high geometric flexibility. The approach simplifies the simulation process by directly incorporating computational fluid dynamics mesh information into the discrete element method solver. The performance analysis indicates about sixteen times boost in computational speed compared to benchmark central processing unit-based solvers. Finally, the solver's compatibility with polyhedral meshes, a vital advantage for complex geometries, is tested against a referenced study regarding the simulation of an immersed-tube fluidized bed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Tutorial: Machine-Learning-Based CREASE-2D Analysis of 2D SAXS Profiles to Characterize Anisotropic Nanostructures in Soft Materials

We present a tutorial to guide users on how to extend the Computational Reverse Engineering Analysis of Scattering Experiments-2D (CREASE-2D) framework to interpret their experimental two-dimensional small-angle scattering (SAS) data from soft materials (e.g., polymers, peptide amphiphiles, biomolecular fibrils). Unlike most traditional SAS analysis approaches, which typically rely on azimuthally averaged onedimensional (1D) profiles, CREASE-2D utilizes the complete 2D scattering profile to reveal information about anisotropy in the structure. In past applications, CREASE has provided insights into complex structural features, including the cross-sectional shapes of assembled nanostructures and dispersity in these features, which are difficult to discern with existing analytical models. While (1D- ) CREASE has been applied to SANS and SAXS data, this tutorial shares the steps for implementing CREASE-2D using an example of a dipeptide solution system, for which we have SAXS data. We present details for these steps involved in using CREASE-2D to interpret SAXS profiles: how to preprocess SAXS data, define relevant structural features, generate three-dimensional real-space structures for specific values of these features, train a machine learning (ML) surrogate model to predict scattering profiles for given structural features, and optimize these features using genetic algorithms (GA). Then, we use these steps to interpret complex 2DSAXS data collected from dipeptide solutions that, in microscopy images, exhibit nanoscale structures that could be elliptical tubes/ flat tapes/cylinders or a combination of these cross sections. Open-source codes, computational hardware, and software requirements, as well as the strengths and limitations of this protocol, are also presented. We expect researchers working with (soft) biomaterials, peptide amphiphiles, amphiphilic polymer solutions, polymer nanocomposites, and blends of particles/polymers will find this CREASE-2D method and this tutorial of use.

CREASE↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗