Search NASA⌕ Search

SEARCH · Search NASA

Results for “error correction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Parallelized telecom quantum networking with an ytterbium-171 atom array

The integration of quantum computers and sensors into a quantum network enables new capabilities in quantum information science. Most networks with atom-like qubits operate at visible or near-ultraviolet wavelengths and require conversion to the telecom band for long-distance communication, which reduces efficiency and potentially introduces noise. In this article we report high-fidelity entanglement between ytterbium-171 atoms and optical photons generated directly in the telecommunication band, where fibre loss is low. The nuclear spin of the atom is entangled with a single photon in the time-bin basis, yielding a high atom-measurement-corrected atom–photon Bell state fidelity. This can be further improved by addressing photon measurement errors. By imaging the atom array onto an optical fibre array, we also implement a parallelized networking protocol that can increase the remote entanglement rate proportionately with the number of channels. We also preserve coherence on a memory qubit during operations on communication qubits. These results support the integration of atomic systems into scalable quantum networks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Type Ia Supernova Growth-rate Measurement with LSST Simulations: Intrinsic Scatter Systematics

Measurement of the growth rate of structures (fσ 8 ) with Type Ia supernovae (SNe Ia) will improve our understanding of the nature of dark energy and enable tests of general relativity. In this paper, we generate simulations of the 10 yr SN Ia data set of the Rubin-LSST survey, including a correlated velocity field from an N-body simulation and realistic models of SNe Ia properties and their correlations with host-galaxy properties. We find, similar to SN Ia analyses that constrain the dark energy equation-of-state parameters w 0 w a , that constraints on fσ 8 can be biased depending on the intrinsic scatter of SNe Ia. While for the majority of intrinsic scatter models we recover fσ 8 with a precision of ∼13%–14%, for the most realistic dust-based model, we find that the presence of non-Gaussianities in Hubble diagram residuals leads to a bias on fσ 8 of ∼ −20%. When trying to correct for the dust-based intrinsic scatter, we find that the propagation of the uncertainty on the model parameters does not significantly increase the error on fσ 8 . We also find that while the main component of the error budget of fσ 8 is the statistical uncertainty (>75% of the total error budget), the systematic error budget is dominated by the uncertainty on the damping parameter, σ u , that gives an empirical description of the effect of redshift space distortions on the velocity power spectrum. Our results motivate a search for new methods to correct for the non-Gaussian distribution of the Hubble diagram residuals, as well as an improved modeling of the damping parameter.

Carreres, Bastien [Duke Univ., Durham, NC (United ↗

A Particle-in-Cell Method for Plasmas with a Generalized Momentum Formulation, Part II: Enforcing the Lorenz Gauge Condition

In a previous paper Christlieb et al. (A particle-in-cell method for plasmas with a generalized momentum formulation, part I: Model formulation, 2024), we developed a new particle-in-cell (PIC) method for the relativistic Vlasov–Maxwell system in which the electromagnetic fields and the equations of motion for the particles were cast in terms of scalar and vector potentials through a Hamiltonian formulation. This new method evolved the potentials under the Lorenz gauge using integral equation methods. New methods to construct spatial derivatives of the potentials that converge at the same rates as the fields were also presented. The new particle method was compared against standard explicit discretizations, including the well-known FDTD-PIC method, for a range of applications involving sheaths and particle beams. Here, this paper extends this new class of methods by focusing on the enforcement the Lorenz gauge condition in both exact and approximate forms using co-located meshes. A time-consistency property of the proposed field solver for the vector potential form of Maxwell’s equations is established, which is shown to preserve the equivalence between the semi-discrete Lorenz gauge condition and the analogous semi-discrete continuity equation. Using this property, we present three methods to enforce a semi-discrete gauge condition. The first method introduces an update for the continuity equation that is consistent with the discretization of the Lorenz gauge condition. Both the finite difference and spectral implementations satisfy this discrete gauge condition to machine precision. The second approach we propose enforces a semi-discrete continuity equation using the boundary integral solution to the field equations. The potential benefit of this approach is that it eliminates spatial derivatives that appear on the particle data, namely the current density, which is often calculated by linear combinations of low-order spline basis functions. This method is ideally suited to boundary integral equation methods that invert multi-dimensional operators without dimensional splitting techniques and will be the subject of future work. The third approach introduces a gauge correcting method that makes direct use of the gauge condition to modify the scalar potential and uses local maps for both the charge and current densities. This results in a gauge error, as the maps do not enforce the continuity equation. The vector potential coming from the current density is taken to be exact, and using the Lorenz gauge, we compute a correction to the scalar potential that makes the two potentials satisfy the gauge condition. This method also enforces the gauge condition to machine precision. We demonstrate two of the proposed methods in the context of periodic domains. Problems defined on bounded domains, including those with complex geometric features remain an ongoing effort. However, this work shows that it is possible to design computationally efficient methods that can effectively enforce the Lorenz gauge condition in a non-staggered PIC formulation.

97 MATHEMATICS AND COMPUTING↗

Kernel Manifolds: Nonlinear‐Augmentation Dimensionality Reduction Using Reproducing Kernel Hilbert Spaces

This paper generalizes recent advances on quadratic manifold (QM) dimensionality reduction by developing kernel methods-based nonlinear-augmentation dimensionality reduction. QMs, and more generally feature map-based nonlinear corrections, augment linear dimensionality reduction with a nonlinear correction term in the reconstruction map to overcome approximation accuracy limitations of purely linear approaches. While feature map-based approaches typically learn a least squares optimal polynomial correction term, we generalize this approach by learning an optimal nonlinear correction from a user-defined reproducing kernel Hilbert space. Our approach allows one to impose arbitrary nonlinear structure on the correction term, including polynomial structure, and includes feature map and radial basis function-based corrections as special cases. Furthermore, our method has relatively low training cost and has monotonically decreasing error as the latent space dimension increases. In conclusion, we compare our approach to proper orthogonal decomposition and several recent QM approaches on data from several example problems.

kernel methods↗

Accurate energies for ππ* excited states via exchange scaling: the XS-CASSCF method

The state-averaged complete-active space self-consistent field method (SA-CASSCF) is a widely employed electronic structure method used for studying photochemistry and dynamics owing to its ability to provide a reliable description even of complicated cases while still retaining computational efficiency. However, SA-CASSCF suffers from one Achilles heel, related to the description of ionic ππ* excited states, whose energy is often overestimated by 1–2 eV. In light of this challenge, we present the XS-CASSCF method, a new approach based on the idea of exchange scaling (XS) that screens the involved energy terms to improve the excitation energies of singlet ionic ππ* states. First, we illustrate the power of the XS-CASSCF method using hexatriene and para-quinodimethane as examples, showing that it corrects the targeted ionic states while leaving the other states largely unaffected, giving root-mean-square errors (RMSE) below 0.2 eV for the four lowest states in both cases. Subsequently, XS-CASSCF vertical excitation energies are tested against theoretical best estimates for a set of 11 molecules and 56 excited states. XS-CASSCF performs exceptionally well for the ππ* states of hydrocarbons, reducing the RMSE over 21 excitation energies from 0.96 to 0.27 eV. In the challenging subset of molecules with heteroatoms and a larger number of ππ* and nπ* states, we find that improvements can also be obtained, albeit not as pronounced. We conclude with an outlook into more realistic molecular materials focusing on their singlet–triplet (S 1 /T 1 ) gaps, finding that significant improvements can be obtained along the whole range of S 1 /T 1 gaps studied, going from 0.1 eV to more than 1.5 eV. Owing to notable improvements across significant classes of molecules combined with its conceptual simplicity, we believe that XS-CASSCF is a promising addition to the electronic structure toolbox, serving both as a standalone electronic structure method and as a starting point for further correlated treatment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bias-Variance Trade-Off in Physics-Informed Neural Networks with Randomized Smoothing for High-Dimensional PDEs

Physics-Informed Neural Networks (PINNs) have triggered a paradigm shift in scientific computing, leveraging mesh-free properties and robust approximation capabilities. While proving effective for low-dimensional partial differential equations (PDEs), the computational cost of PINNs remains a hurdle in high-dimensional scenarios. This is particularly pronounced when computing high-order and high-dimensional derivatives in the physics-informed loss. Randomized Smoothing PINN (RS-PINN) introduces Gaussian noise for stochastic smoothing of the original neural net model, enabling the use of Monte Carlo methods for derivative approximation, which eliminates the need for costly automatic differentiation. Despite its computational efficiency, especially in the approximation of high-dimensional derivatives, RS-PINN introduces biases in both loss and gradients, negatively impacting convergence, especially when coupled with stochastic gradient descent (SGD) algorithms. We present a comprehensive analysis of biases in RS-PINN, attributing them to the nonlinearity of the Mean Squared Error (MSE) loss as well as the intrinsic nonlinearity of the PDE itself. We propose tailored bias correction techniques, delineating their application based on the order of PDE nonlinearity. The derivation of an unbiased RS-PINN allows for a detailed examination of its advantages and disadvantages compared to the biased version. Specifically, the biased version has a lower variance and runs faster than the unbiased version, but it is less accurate due to the bias. To optimize the bias-variance trade-off, we combine the two approaches in a hybrid method that balances the rapid convergence of the biased version with the high accuracy of the unbiased version. In addition to methodological contributions, we present an enhanced implementation of RS-PINN. Extensive experiments on diverse high-dimensional PDEs, including Fokker-Planck, Hamilton-Jacobi-Bellman (HJB), viscous Burgers’, Allen-Cahn, and Sine-Gordon equations, illustrate the bias-variance trade-off and highlight the effectiveness of the hybrid RS-PINN. Empirical guidelines are provided for selecting biased, unbiased, or hybrid versions, depending on the dimensionality and nonlinearity of the specific PDE problem.

97 MATHEMATICS AND COMPUTING↗

Using a Large Language Model as a Building Block to Generate Usable Validation and Verification Suite for OpenMP

In the HPC area, both hardware and software move quickly. Often new hardware is developed and deployed, the corresponding software stack, including compilers and other tools, are under active development while leading edge software developers are working to port and tune their applications, all at the same time. While the software ecosystem is in flux, one of the key challenges for users is obtaining insight into the state of implementation of key features in the programming languages and models their applications are using – whether they have been implemented, and whether the implementation conforms to the specification, especially for newly implemented features (less tested by widespread use). OpenMP is one of the most prominent shared memory programming models used for on-node programming in HPC. With the shift towards accelerators (such as GPUs and FPGAs) and heterogeneous programming OpenMP features are getting more complex. It is natural to ask whether generative AI approaches, and large language models (LLMs) in particular, can help in producing validation and verification test suites to allow users better and faster insights into the availability and correctness of OpenMP features of interest. In this work, we explore the use of ChatGPT-4 to generate a suite of tests for OpenMP features. We have chosen a set of directives and clauses, a total of 78 combinations, which first appeared in OpenMP 3.0 (released in May 2008) but are also relevant for accelerators. We prompted ChatGPT to generate tests in the C and Fortran languages, for both host (CPU) and device (accelerator). On the Summit super-computer using the GNU implementation, we found that, of the 78 generated tests 67 C tests and 43 Fortran tests compiled successfully and fewer than those executed to completion. On further analysis we show that not all generated tests are valid. We document the process, results, and provide detailed analysis regarding the quality of tests generated. With the aim of providing input to a production quality validation and verification suite, we manually implement the corrections required to make the tests valid according to the current OpenMP specification. We quantify this effort as small, medium, or large, and record the lines of code changed to correct the invalid tests. With the corrected tests we validate recent implementations from HPE, AMD, and GNU on the Frontier supercomputer. Our experiment and subsequent analysis show that although LLMs are capable of producing HPC specific codes, they are limited by their understanding of the deeper semantics and restrictions of programming models such as OpenMP. Unsurprisingly more commonly used features have better support, while some OpenMP 3.0 directives such as sections and tasking are not universally supported on accelerators. We demonstrate that successful compilation and execution to completion are inadequate metrics for evaluating generated code and that, at this time, commodity LLMs require expert intervention for code verification. This points to gaps in the training data that is currently available for HPC. We demonstrate that with "small" effort 37% of generated invalid C tests and 63% of generated invalid Fortran tests could be corrected. This improves productivity of test generation as we circumvent writing from scratch and the common programming errors associated with it.

Pophale, Swaroop [ORNL] (ORCID:0000000185446367)↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Inverse bremsstrahlung absorption rate for super-Gaussian electron distribution functions including plasma screening

Here we provide analytic expressions for the effective Coulomb logarithm for inverse bremsstrahlung absorption which predict significant corrections to the Langdon effect and overall absorption rate compared to previous estimates. The calculation of the collisional absorption rate of laser energy in a plasma by the inverse bremsstrahlung mechanism usually makes the approximation of a constant Coulomb logarithm. We dispense with this approximation and instead take into account the velocity dependence of the Coulomb logarithm, leading to a more accurate expression for the absorption rate valid in both classical and quantum conditions. In contrast to previous work, the laser intensity enters into the Coulomb logarithm. In most laser-plasma interactions the electron distribution function is super-Gaussian [Langdon, Phys. Rev. Lett. 44, 575 (1980)], and we find the absorption rate under these conditions is increased by as much as ≈ 30% compared to previous estimates at low density. In many cases of interest the correction to Langdon's predicted reduction in absorption is large; for example at Z = 6 and Te = 400 eV the Langdon prediction for the absorption is in error by a factor of ≈ 2. However, we also account for the additional effect of plasma screening, which predicts a reduction in absorption by a similar amount (up to ≈ 30%). These two effects compete to determine the overall absorption, which may be increased or decreased, depending on the conditions. The corrections can be incorporated into radiation-hydrodynamics simulation codes by replacing the familiar Coulomb logarithm with an analytic expression which depends on the super-Gaussian order “M” and the screening length.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Two-stage formation-energy correction (NbZr, TaZr, VZr)

This bundle contains the scripts, the raw and corrected per-structure data, and the manuscript plots for the NbZr / TaZr / VZr BCC binary formation energies and the associated RMSDs. Why a two-stage correction is necessary: The "raw" formation energy of every relaxed VASP configuration is computed in the usual way, FE_raw(c) = E_alloy(c) - sum_i x_i * E_pure_i , where E_pure_i are the per-atom total energies of the pure-element reference structures (Nb, Ta, V, Zr in the same BCC supercell, with identical INCAR / KPOINTS / PAW choices). With perfectly consistent reference runs the raw FE should vanish at the two pure-element endpoints (x = 0 and x = 1) by construction. In practice this does not hold for two reasons that are present in our dataset: 1. Reference-energy inconsistency (composition-dependent bias). Even with identical input parameters, the pure-element runs (stored in `corrected_DFT_pure_element_runs/`) differ slightly from the values that would be implied by the alloy runs at near-pure compositions (a few meV/atom). This bias is approximately linear in concentration, because the residual error in E_pure_Nb (or E_pure_Ta / E_pure_V) propagates into FE_raw(c) as (1 - x) * dE_pure_1, and the corresponding error in E_pure_Zr propagates as x * dE_pure_2. Left uncorrected, this produces a non-physical "tilt" of FE_raw(x) and shifts the entire FE-vs-x cloud away from zero at the endpoints. 2. Endpoint anchoring against the audited true endpoints. The strict endpoint values (FE_x0_meVatom, FE_x1_meVatom in `corrected_fe_strict_endpoints_20260518/strict_endpoint_check_20260518.csv`) were re-derived from an independent cross-check of the pure-element runs. After stage 1 removes the linear bias, the near-pure compositions in the alloy dataset still extrapolate to values that differ slightly from these audited endpoints — because stage 1 is fit from a few near-end alloy bins, not from the audited pure-element references themselves. The README.txt file discusses how these issues are addressed by the two-stage correction, and describes folder layout, pipeline summary, and how to re-run.

36 MATERIALS SCIENCE↗

A moment-conserving discontinuous Galerkin representation of the relativistic Maxwellian distribution

Kinetic simulations of relativistic gases and plasmas are critical for understanding diverse astrophysical and terrestrial systems, but the accurate construction of the relativistic Maxwellian, the Maxwell–Jüttner distribution, on a discrete simulation grid is challenging. Difficulties arise from the finite velocity bounds of the domain, which may not capture the entire distribution function, as well as errors introduced by projecting the function onto a discrete grid. Here, we present a novel scheme for iteratively correcting the moments of the projected distribution applicable to all grid-based discretizations of the relativistic kinetic equation. In addition, we describe how to compute the needed nonlinear quantities, such as Lorentz boost factors, in a discontinuous Galerkin scheme through a combination of numerical quadrature and weak operations. The resulting method accurately captures the distribution function and ensures that the moments match the desired values to machine precision.

astrophysical plasmas↗

Optics tuning of the FCC-ee

The Future Circular Collider, FCC-ee, is a proposed next generation electron-positron collider aiming to provide large luminosities at beam energies from 45.6 up to 182.5 GeV. This collider faces a major challenge to deliver the design performance in the presence of realistic lattice errors. A commissioning strategy has been developed including dedicated optics designs, efficient beam-based alignment and optics corrections based on refined optics measurements. First specifications on main magnets, corrector circuits, and instrumentation have also been investigated. A summary of all these aspects is presented in this paper.

Tomas, Rogelio [CERN]↗

Evaluating Probabilistic Deep Learning Methods for Uncertainty Quantification of Precipitation Bias Correction

Climate models often exhibit biases in their precipitation predictions, particularly underestimating high-intensity events and overestimating low precipitation. Deep learning approaches offer promising solutions, but their epistemic uncertainty associated with a deep learning–based bias correction method has not previously been quantified for reliable downstream climate impact studies. While methods for capturing the epistemic uncertainty in deep learning frameworks exist, there is currently no consensus on the best method. In this work, we compare three uncertainty quantification (UQ) methods—Deep Ensembles (DEns), Monte Carlo Dropout (MCD), and Flipout—by assessing the reliability of their uncertainty estimates using standard measures such as sharpness and calibration. These UQ methods are applied to an existing deep learning precipitation bias correction model known as UFNet: a coupled U-Net and fully connected neural network. The methods utilized to assess the models’ uncertainties are 1) calibration, which ensures that the expected probabilities of the model align with reality and 2) sharpness, which is a measure of the precision of the model’s probabilistic predictions. Of the three UQ methods evaluated, the DEns and MCD methods demonstrated the best-calibrated performance (expected calibration error of 0.36 and 0.35, respectively), compared to Flipout (0.58). In contrast, Flipout had the sharpest predictions and the highest metric performance in bias correcting precipitation—especially for higher-order moments such as kurtosis with a spatial correlation of 72% compared to 32% and 55% spatial correlation for DEns and MCD, respectively. Of the three UQ methods, MCD was found to be the most suitable method for UQ purposes based on its calibration, sharpness, and computational requirements.

Bayesian methods↗

(Doublon) Benchmarking of Different Inverse Point Kinetics Implementations for an Autocorrected Reactimeter Algorithm

In November 2017, the Transient Reactor Test Facility returned to operation. Since that time, many transient test series have been completed, such as the Transient Heatsink Overpower Response capsule (THOR), the Transient Water Irradiation System for TREAT (TWIST), and Sirius. Each has provided valuable data for materials performance and reactor safety that can be applied in future designs. During each experimental series, detector count rates provided important information on the core behavior during transients. However, a limitation of these data is that variations in the neutron distribution during experiments can cause errors when attempting to infer reactivity evolution from detector signals. Neutron physics codes can be used to compute the flux shape variations. However, this is a poor solution when the experimental data is used for code verification, validation and uncertainty quantification. Indeed, if the output of the code is used both as a reference and to correct what the reference is compared to, the circular dependency limits the quality of the verification, validation and uncertainty quantification approach. To overcome this problem, the autocorrected reactimeter algorithm (ACRA) has been developed. This approach infers a time-dependent reactivity evolution by testing different spatial corrections and selecting the one that minimizes reactivity variations when the core is in a frozen configuration (i.e., when there is no variation in parameters affecting reactivity). However, the scope of this method was limited to transients where there were negligible thermal feedback. Indeed, the core is never in a frozen configuration when the fuel temperature varies during the whole transient. This is our motivation for developing an improved version of the ACRA that does not require frozen configurations. To develop this new algorithm, we need a precise and unbiased implementation of the inverse point kinetic equations (IPKEs) as any error in the reactivity evaluation will be propagated into the choice of the optimal spatial correction. Indeed, the previous reactimeter algorithm would use approximations, such as a negligible flux amplitude derivative, to focus on rapidity. For the numerical validation of ACRA, we aim at absolute error under for reactivity derived from signals similar to the one of this study. In this summary, we test eight different IPKE implementations. Each will process a mockup signal built for this study, similar to those that the future ACRA will process. Each reactivity output will be compared to the reference reactivity that has been used to generate the mockup signal. The implementation minimizing the difference with the reference reactivity will be used in the development of a new ACRA formulation.

73 - NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Feedforward-feedback ammonia control at a water resource recovery facility based on a digital twin with hybrid model

Ammonia-based aeration control (ABAC) at full-scale Water Resource Recovery Facilities (WRRFs) can be challenged by diurnal loading and transport delays. This work addressed these challenges using a hybrid feedforward–feedback controller built on Activated Sludge Model 1 (ASM1), marking the first full-scale deployment to pair a mechanistic feedforward core with data-driven corrections. The objectives were to improve ammonia setpoint tracking, assess performance of the mechanistic model when enhanced with data-driven corrections, and document full-scale operation. The hybrid model incorporates two data-driven components: (1) a Mechanistic Error Forecasting Engine (MEFE), consisting of a multivariate linear regressor and a long short-term memory (LSTM) ensemble. Defying expectations, low-parameter models outperformed more complex alternatives, reducing the mechanistic error by 71%. (2) A Residual Oscillation Forecasting Engine (ROFE), based on Fast Fourier Transform, reduced the remaining error by another 35%. Two proportional–integral (PI) feedback loops further (i) trim the feedforward output and (ii) eliminate residual controller error in the final aerobic zone. In full-scale operation, the controller reduced mean-squared error (MSE) by 94% over the baseline and produced more stable dissolved oxygen (DO) setpoints. Overall, it was proven that layering multi-timescale data-driven models on a mechanistic core can yield reliable ABAC performance at WRRFs.

54 ENVIRONMENTAL SCIENCES↗

Improving precision and accuracy of genetic mapping with genotyping‐by‐sequencing data in outcrossing species

Abstract Genotyping‐by‐sequencing (GBS) is a widely used strategy for obtaining large numbers of genetic markers in model and non‐model organisms. In crop plants, GBS‐derived marker datasets are frequently used to perform quantitative trait locus (QTL) mapping. In some plant species, however, high heterozygosity and complex genome structure mean that researchers must use care in handling GBS data to conduct QTL mapping most effectively. Such outbred crops include most of the perennial grass and tree species used for bioenergy. To identify strategies for increasing accuracy and precision of QTL mapping using GBS data in outbred crops, we conducted an empirical study of SNP‐calling and genetic map‐building pipeline parameters in a Miscanthus sinensis population, and a complementary simulation study to estimate the relationship between genome‐wide error rate, read depth, and marker number. The bioenergy grass Miscanthus is an obligate outcrossing species with a recent (diploidized) whole‐genome duplication. For the study of empirical M. sinensis data, we compared two SNP‐calling methods (one non‐reference‐based and one reference‐based), a series of depth filters (12×, 20×, 30×, and 40×) and two map‐construction methods (i.e., marker ordering: linkage‐only and order‐corrected based on a reference genome). We found that correcting the order of markers on a linkage map by using a high‐quality reference genome improved QTL precision (shorter confidence intervals). For typical GBS datasets of between 1000 and 5000 markers to build a genetic map for biparental populations, a depth filter set at 30× to 40× applied to outbred populations provided a genome‐wide genotype‐calling error rate of less than 1%, improved accuracy of QTL point estimates and minimized type I errors for identifying QTL. Based on these results, we recommend using a reference genome to correct the marker order of genetic maps and a robust genotype depth filter to improve QTL mapping for outbred crops.

59 BASIC BIOLOGICAL SCIENCES↗

Toward the “platinum standard” of quantum chemistry on quantum computers: Perturbative quadruple corrections in unitary coupled cluster theory

We propose a non-iterative, post-hoc correction to the unitary coupled cluster theory with the single, double, and triple excitations (UCCSDT) Ansatz, which considers the leading-order effects of neglected quadruple excitations. We present two ways to derive this correction, henceforth referred to as [Q-6], which leads to an improvement in the correlation energy shown to be truncated to sixth-order in many-body perturbation theory. Furthermore, a comparison between the UCC-based [Q-6] correction proposed in this work and analogous, “platinum standard” quadruple corrections proposed in conventional coupled cluster theory recognizes that [Q-6] is distinct from prior corrections since it is constructed entirely from internally connected components. Although trotterized (t) and full operator variants of UCCSDT exhibit errors in scans of small molecule potential energy surfaces that routinely exceed 1.6 mH, we find that t/UCCSDT[Q-6] is, nevertheless, able to achieve chemical accuracy as measured by the mean unsigned error.

Correlation energy↗

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES↗