Search NASASearch

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

PPI DataHub Project Data Package: S. elongatus PCC 7942 Circadian Control Bioproduction Transcriptomics (PB-DP3)

The purpose of this experiment was to evaluate how circadian clock regulation impacts carbon partitioning between storage, growth, and product synthesis in Synechococcus elongatus PCC 7942 in providing insights to strategies for enhanced bioproduction. Sample data was acquired using a Illumina HiSeq sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis. Transcriptomic differential expression analysis revealed coordinated circadian clock-driven adjustment of the cell cycle and rewiring of energy and carbon metabolism. Processed RNA-Seq datasets are openly accessible from the PNNL DataHub project dataset download page and contain secondary processed RNA-seq results files and supporting metadata materials linked to relevant source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES

A Curated Dataset of Regional Meteor Events with Simultaneous Optical and Infrasound Observations (2006–2011)

We present a curated, openly accessible dataset of 71 regional meteor events simultaneously recorded by optical and infrasound instrumentation between 2006 and 2011. These events were captured during an observational campaign using the all-sky cameras of the Southern Ontario Meteor Network and the co-located Elginfield Infrasound Array. Each entry provides optical trajectory measurements, infrasound waveforms, and atmospheric specification profiles. The integration of optical and acoustic data enables robust linkage between observed acoustic signals and specific points along meteor trajectories, offering new opportunities to examine shock wave generation, propagation, and energy deposition processes. This release fills a critical observational gap by providing the first validated, openly accessible archive of simultaneous optical–infrasound meteor observations that supports trajectory reconstruction, acoustic propagation modeling, and energy deposition analyses. By making these data openly available in a structured format, this work establishes a durable reference resource that advances reproducibility, fosters cross-disciplinary research, and underpins future developments in meteor physics, atmospheric acoustics, and planetary defense.

astrometry

Datasets of Faults in Variable Air Volume Terminal Units in a Multi-Zone Commercial Building

Faults in HVAC systems can decrease system efficiency and equipment lifespan, leading to 5%–30% of energy consumption being wasted in commercial buildings. We identified two common faults in HVAC variable air volume systems: a stuck damper fault in the variable air volume terminal unit and a discharge airflow sensor fault. We conducted three sets of damper stuck tests and two sets of airflow sensor tests, each including a fault-free scenario and scenarios with varying levels of faults, over one day. The faults were implemented in Oak Ridge National Laboratory’s two-story Flexible Research Platform building to generate a high-quality, well-controlled dataset covering fault-induced and fault-free scenarios. The test building, fault test scenarios, and data validation are described here. The open-source dataset includes 1 min intervals of weather and building data on the presence and absence of building faults. This dataset can be used to analyze the effects of HVAC system faults on system operation and indoor building conditions, and to develop or evaluate a fault detection and diagnosis algorithm.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Beyond Price-Taker: Multiscale Optimization of a Wind-Battery Integrated Energy System within the Wholesale Electricity Market

This work presents the optimization of a wind-battery IES using the multiscale optimization framework proposed in our previous work to quantify errors from the price-taker assumption. The framework, built over Prescient (an open-source package for solving production cost models), is applied to the RTS-GMLC dataset, an open-source dataset that is representative of the southwest U.S. wholesale electricity market. The framework provides detailed bidding, market clearing, and control processes of an IES, and it can quantify how the IES interacts with the market. In this work, we use the retrofit of a wind farm with a battery storage system as an example to show the difference in the market outcomes and revenues obtained from both price-taker and multiscale optimization approaches. Our work goes beyond price-taker and deep dives into quantifying IES-market interaction in optimizing IES. This framework enables users to explore how different design and operation decisions of energy systems interact with the market and provides a more accurate evaluation than the price-taker assumption.

Chen, Xinhe

Dataset describing two reference models for full-spectral lighting and daylight simulations together with implementations for two software systems

A dataset of two spectral lighting simulation reference models - one office and one factory hall - is presented. It aims to demonstrate and support full-spectral daylight and electric lighting simulations and facilitate evaluation of non-visual effects of light. The dataset includes Rhino CAD geometry, comprehensive spectral material and light source data and window system BSDF data. Example implementations in the two software tools, Radiance and OWL, enable reproducible workflows and support adoption in other software. The dataset is openly available on Zenodo. The office model reproduces Room 518 at the University of Innsbruck, including a west-facing façade and interior furnishings. The factory hall model follows the proposed geometry in the European standard 15193 for building energy performance. Interior reflectances in the office were measured in-situ using a handheld spectrometer. Exterior spectra and factory hall materials matching specified reflectances were obtained from an online spectral materials database. Glazing transmittance was derived from IGDB data using LBNL Optics/WINDOW. BSDFs for venetian blinds at various tilt angles, and for a diffusing pane adapted from the Complex Glazing Database, were generated in WINDOW. Luminaires in both models are specified with photometric files (Eulumdat/IES) and lamp spectra (Fluorescent 840, 4000 K LED). The provided example implementations (Radiance, OWL) include prepared input data and scripts to run first spectral simulations; example results are also included. The dataset is prepared to support reuse by researchers, designers and software developers for method validation, software engineering and comparison, and development of spectral metrics and controls.

Geisler-Moroder, David

Can protein expression be ‘solved’?

Recombinant protein expression is central to biotechnology’s application in academic exploration as well as human health, climate applications and the bioeconomy in general. However, not all proteins can be expressed in all organisms, and the field lacks a predictive model of soluble protein overexpression that could replace laborious experimental trial-and-error. Here, we discuss the state of the field and identify the lack of large, high-fidelity datasets as the primary bottleneck to progress. We review possible assays that could be used for data collection to identify a path toward an extensible experimental platform for collecting soluble recombinant protein overexpression data across organisms. We suggest that the resulting dataset should be used to train increasingly generalizable predictive models of protein expression to answer the question: “How can predictive protein expression be solved?”.

59 BASIC BIOLOGICAL SCIENCES

A Comprehensive Calibration Framework for the Northwest River Forecast Center

We present a comprehensive framework developed by the Northwest River Forecast Center for calibrating hydrologically diverse basins. The framework includes models for snow, soil moisture, routing, channel loss, and consumptive use. Data inputs include a wide range of open-access datasets for meteorology, land use, topography, and land cover. The framework uses conceptual hydrologic models to handle basins with various hydrologic regimes including rain-driven and snowmelt-dominated basins. We also develop a flexible automatic calibration system that can handle numerous unobservable model parameters in a computationally efficient manner. A single-basin automatic calibration run can typically be completed on a modern laptop in under 10 min. We found that model performance metrics for this new approach match the quality of the NWRFC's previous labor-intensive manual calibrations. The model performance also rivals that of a state-of-the-art deep learning model at a fraction of the computational cost. This framework presents a new standard for the quality of calibrations possible with lumped conceptual hydrologic models, combining careful data curation, an objective calibration framework, and expert local knowledge. In addition, we have made software packages available for the entire suite of National Weather Service River Forecast System models, including SAC-SMA, SNOW-17, and Lag-K. These modern interfaces are intended to increase accessibility and facilitate future research.

Forecasting

A Deep Learning Approach for Detection and Localization of Leaf Anomalies

The detection and localization of possible diseases in crops are usually automated by resorting to supervised deep learning approaches. In this work, we tackle these goals with unsupervised models, by applying three different types of autoencoders to a specific open-source dataset of healthy and unhealthy pepper and cherry leaf images. CAE, CVAE and VQ-VAE autoencoders are deployed to screen unlabeled images of such a dataset, and compared in terms of image reconstruction, anomaly removal, detection and localization. The vector-quantized variational architecture turns out to be the best performing one with respect to all these targets.

Calabro', Davide

Short-term electricity load forecasting: Application-driven evaluation of machine learning models across spatial and temporal scales

As we transition towards a decarbonized economy, the integration of variable renewable energy resources and new demands (e.g., electric vehicles, heat pumps) into the electricity grid places unprecedented pressure on grid operators to effectively anticipate and manage peak load. In this context, machine learning algorithms are proving to be indispensable for accurate short-term load forecasting, a crucial task to address these challenges. This study benchmarks 6 machine learning algorithms, including three neural networks and three tree-based algorithms, across various levels of spatial aggregation and time horizons (1, 4, 8, 24, and 48 h). The central contribution of this work is the comparison and analysis of load forecasting models not only based on statistical metrics, but also based on a novel error metric, which evaluates the cost implications of forecast errors for power system stakeholders. Results show that tree-based models outperform neural networks, based on statistical metrics, and yield less skewed error distributions for most spatial scales. However, through the lens of the novel error metric, neural networks are the more competitive choice, especially for forecast horizons that exceed 8 h. The study concludes with actionable recommendations to grid operators and highlights the need for the development of error metrics that link forecasting accuracy to operational costs. To promote transparency and open science, the datasets and Python code are open-sourced via a supplementary repository.

Houben, Nikolaus

Learning together: Towards foundation models for machine learning interatomic potentials with meta-learning

Abstract The development of machine learning models has led to an abundance of datasets containing quantum mechanical (QM) calculations for molecular and material systems. However, traditional training methods for machine learning models are unable to leverage the plethora of data available as they require that each dataset be generated using the same QM method. Taking machine learning interatomic potentials (MLIPs) as an example, we show that meta-learning techniques, a recent advancement from the machine learning community, can be used to fit multiple levels of QM theory in the same training process. Meta-learning changes the training procedure to learn a representation that can be easily re-trained to new tasks with small amounts of data. We then demonstrate that meta-learning enables simultaneously training to multiple large organic molecule datasets. As a proof of concept, we examine the performance of a MLIP refit to a small drug-like molecule and show that pre-training potentials to multiple levels of theory with meta-learning improves performance. This difference in performance can be seen both in the reduced error and in the improved smoothness of the potential energy surface produced. We therefore show that meta-learning can utilize existing datasets with inconsistent QM levels of theory to produce models that are better at specializing to new datasets. This opens new routes for creating pre-trained, foundation models for interatomic potentials.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Heterogeneous estimations of non-pharmaceutical mitigation behavior during the COVID-19 pandemic

The COVID-19 pandemic highlighted the importance of human behavior in mitigating the spread of disease. Nonetheless, human behavior is often overlooked in models of disease spread, particularly by underutilizing real-world data. We address this by estimating probabilities that individuals engage in behaviors that influence SARS-CoV-2 transmission risk during the COVID-19 pandemic, between September 2020 and June 2022. These behaviors include wearing a mask, using public transportation, spending time with others, avoiding contact with others, and going to work. Our estimates account for the age and sex of individuals and are generated for every county in the United States. We utilized multiple open-source datasets and United States Census data to produce these estimates. Multiple datasets were used for validation, showing our estimates demonstrated comparable accuracy and robustness. Our estimates aid in understanding human behavior dynamics during the COVID-19 pandemic and could be used to inform monthly or longer-term behavior in simulations of COVID-19. Moreover, the methods presented can be applied to other behaviors and features for future simulations of infectious disease.

97 MATHEMATICS AND COMPUTING

White paper on light sterile neutrino searches and related phenomenology

This white paper provides a comprehensive review of our present understanding of experimental neutrino anomalies that remain unresolved, charting the progress achieved over the last decade at the experimental and phenomenological level, and sets the stage for future programmatic prospects in addressing those anomalies. It is purposed to serve as a guiding and motivational "encyclopedic" reference, with emphasis on needs and options for future exploration that may lead to the ultimate resolution of the anomalies. We see the main experimental, analysis, and theory-driven thrusts that will be essential to achieving this goal being: 1) Cover all anomaly sectors -- given the unresolved nature of all four canonical anomalies, it is imperative to support all pillars of a diverse experimental portfolio, source, reactor, decay-at-rest, decay-in-flight, and other methods/sources, to provide complementary probes of and increased precision for new physics explanations; 2) Pursue diverse signatures -- it is imperative that experiments make design and analysis choices that maximize sensitivity to as broad an array of these potential new physics signatures as possible; 3) Deepen theoretical engagement -- priority in the theory community should be placed on development of standard and beyond standard models relevant to all four short-baseline anomalies and the development of tools for efficient tests of these models with existing and future experimental datasets; 4) Openly share data -- Fluid communication between the experimental and theory communities will be required, which implies that both experimental data releases and theoretical calculations should be publicly available; and 5) Apply robust analysis techniques -- Appropriate statistical treatment is crucial to assess the compatibility of data sets within the context of any given model.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo

When does global attention help: a unified empirical study on atomistic graph learning

Graph neural networks (GNNs) are widely used as surrogates for costly experiments and first-principles simulations to study the behavior of compounds at atomistic scale, and their architectural complexity is constantly increasing to enable the modeling of complex physics. While most recent GNNs combine more traditional message passing neural networks (MPNNs) layers to model short-range interactions with more advanced graph transformers (GTs) with global attention mechanisms to model long-range interactions, it is still unclear when global attention mechanisms provide real benefits over well-tuned MPNN layers due to inconsistent implementations, features, or hyperparameter tuning. We introduce the first unified, reproducible benchmarking framework–built on HydraGNN–that enables seamless switching among four controlled model classes: MPNN, MPNN with chemistry/topology encoders, GPS-style hybrids of MPNN with global attention, and fully fused localglobal models with encoders. Using seven diverse open-source datasets for benchmarking across regression and classification tasks, we systematically isolate the contributions of message passing, global attention, and encoder-based feature augmentation. Our study shows that encoder-augmented MPNNs form a robust baseline, while fused localglobal models yield the clearest benefits for properties governed by long-range interaction effects. We further quantify the accuracycompute trade-offs of attention, reporting its overhead in memory. Together, these results establish the first controlled evaluation of global attention in atomistic graph learning and provide a reproducible testbed for future model development.

Equivariant graph neural networks

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE

IEA Wind TCP Task 49: Reference Site Conditions for Floating Wind Arrays

This report, prepared within Work package 1 of IEA Wind Task 49, presents reference site conditions for floating wind arrays to serve as a design basis for the techno-economic design of reference floating wind arrays. Data of the reference sites presented here are publicly available in an open database and thus support fast development and comparable design of floating wind arrays for various relevant conditions. The development of these reference sites drew on existing open access datasets and ongoing research projects of task participants. Six classes were identified that describe relevant key conditions for the design and development of floating wind arrays: met-ocean conditions, seabed conditions, coastal infrastructure, environmental impact, socio-economic impact, as well as regulations and permissions. The reference sites for the techno-economic design of floating wind arrays are based on a concept with building blocks to synthesize purpose-built site representations. In each of the identified classes with influencing design factors, building blocks are used to describe the characteristic properties and their spread. However, the latter three classes (i.e., environmental impact, socio-economic impact, regulations and permissions) are not included in the reference site conditions due to limited knowledge and lack of reliable criteria to quantify their impact on the techno-economic design in numeric parameters. Building blocks with key parameters for the techno-economic design of floating wind arrays are provided for met-ocean conditions, seabed conditions, and coastal infrastructure. For met-ocean conditions, multiple sites were selected for detailed analysis that represent a range of conditions across the pipeline of floating wind projects. Wind conditions and sea states are separated, and each location considers both the severity of wind and waves e.g. one site may have a moderate wave condition but severe wind condition. From this pipeline, eleven representative sites were selected where both site-specific analysis was available within the consortium, and where they represent different parts of the global pipeline. The eleven sites are: Hannibal (Italy), Humboldt (US), Ulsan (South Korea), MoneyPoint One (Ireland), Havbredey (UK), Fukushima (Japan), Utsira Nord (Norway), Gulf of Maine (US), Sud de la Bretagne II (France), Sorlige Nordsjo II (Norway). Each of these sites is summarized in the main report while more details about the studies and analyses behind the datasets are provided in the appendix. For seabed conditions, general information about the geotechnical parameters is provided and a baseline is established for the geotechnical parameters and stratigraphy that may be encountered on the sites. A set of six 'synthetic cases' is defined as building blocks providing the different parameters required for design under each case/soil condition. For the coastal infrastructure, general information about the main requirements is provided that a port should comply with to provide a satisfactory service during the construction of floating offshore wind arrays. Minimum port infrastructural requirements are provided for three types of ports.

17 WIND ENERGY

Predicting Li-Ion Battery Capacity Fade Using Early-Life Data and a Hybrid Data-Driven Gaussian Process-Bayesian Regression Approach

Accurately predicting Li-ion battery capacity trajectories using early-life data can dramatically improve battery-life understandings and be used to rapidly evaluate design/cost/performance trade-offs when developing new battery materials. Accurate early-life predictions enable researchers to quickly iterate over cell designs and material precursor properties without consistently cycling cells to failure. To this end, we present a toolbox that uses a combined Gaussian Process and Bayesian regression approach that capitalizes on signals other than just capacity (e.g., dQ/dV, voltage drops) to rapidly predict capacity-fade trajectories. The prediction tool uses Bayesian regression to fit functional forms, e.g., power law, sigmoids, etc., to predict capacity-fade dynamics. By fitting functional forms, the capacity fade can be interrogated at any point in the future, allowing for early cell-failure prediction. Additionally, Bayesian regression allows for accurate uncertainty estimates that account for cell-to-cell variability (aleatoric uncertainty) and the lack of observation data (epistemic uncertainty). By only using early cycle data to predict the capacity fade trajectory, uncertainty bounds at end-of-life can be extremely large. The large uncertainty bounds are further exacerbated because there is no systematic way to define the prior distribution of the functional forms' parameters. We improve our the predicted trajectory confidence interval of our predicted trajectory using two methods. First, we shows that a small amount of held-out cycling data is sufficientuse some train cells, that have been cycled to failure to derive information regarding the appropriate prior distributions for the functional forms' parameters of the functional form, effectively leading to data-driven priors.. We propose constructing the data-driven priors by first running a Bayesian regression starting with uninformed priors to generate intermediate cell-specific posterior parameter distributions. These posterior distributions are combined using a Ggaussian mixture model for each parameter to create the data-driven priors. These mixture models serve as the data-driven prior distributions for the parameters for. Second, we derive multiple features, e.g., C_dchg 0.5 DoD 0.5, log (|mean(dQ/dV_(w_3-w_0 ) (V)|), etc., from the train cellsheld-out cycling data, identify which the features are that best predicting capacity at early/mid-life cycles, and then create Ggaussian process regression models that are used for predicting capacity at early/mid-life cycles for the test cells (see blue dots with error bars in Fig 1b). Finally, these predicted data-points are used in addition to the actual early cycle data capacity fade to construct the Bayesian regression trajectory for the test cell s. Notably. We note that these two methods are complementary and can be combined with each other. We evaluate the performance of our proposed method on an testing open-source dataset from Iowa State University and Iowa Lakes Community College (ISU-ILCC). This dataset comprises of 251 nickel-manganese-cobalt/graphite Lithium-ion cells that are cycled under 63 different conditions. We compute the mean average percentage error (MAPE) and negative log predictive density (NLPD) to quantify the efficacy of our method. Our initial findings suggest that, when only few observations are available, for test cells, when using only Bayesian regression with uninformed priors, a power law functional provides the most accurate predictions. with very few data points. However, asHowever, a the number of data points increases, a twin sigmoidal function becomes more accurate as the number of observations further increases. We also find that using as little as 10% of the data set towards generating data-driven priors can lead to significant improvement in prediction accuracy when using early cycle data. Lastly, we found that augmenting early-cycle data with Gaussian process-predicted capacity data for Bayesian regression greatly improves the prediction accuracy. We will present a comprehensive comparison of our methods to other methods available in the literature and apply this method to additional battery datasets.

42 ENGINEERING

Rhodotorula toruloides Nitrogen Limitation PTM Profiling Multi-Omics (TZ-DP1)

The purpose of this experiment was to evaluate the regulatory stress response of Oleaginous yeast species Rhodotorula toruloides NBRC 0880 (JGI strain IFO0880 v4.0) under nitrogen-rich and nitrogen-limited conditions over time. Time course experimental samples (0, 24, 48, and 72 hours after inoculation) were prepared using a semi-automated multi-PTM proteomic approach, using tandem mass tag 18-plex (TMT18), and lipidome remodeling for downstream multi-omics analysis. Processed datasets are openly accessible from PNNL DataHub and contain secondary processed proteomic (redox, phospho, and global TMT) and lipidomic (positive and negative ion mode) results files and experimental design metadata.

59 BASIC BIOLOGICAL SCIENCES