Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning and control”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Lattice Anisotropy-Driven Reduction of Phonon Velocities in Black Phosphorus

Phonon dynamics and transport determine how heat is utilized and dissipated in materials. In 2D systems for optoelectronics and thermoelectrics, the impact of nanoscale material structure on phonon propagation is central to controlling thermal conduction. Here, we directly observe in-plane coherent acoustic phonon propagation in black phosphorus (BP) using ultrafast electron microscopy. We identify a significant reduction of the group velocities in directions intermediate to the armchair and zigzag lattice directions. Using a machine learning-based model with an >8000 atom supercell, we find that this slowing results from the mixing of in-plane transverse and longitudinal acoustic phonons and is independent of broken symmetries of edge reconstructions. In conclusion, this work demonstrates how coherent phonon transport is sensitive to propagation direction in the lattice plane.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluating Cloud Properties at Scott Base: Comparing Ceilometer Observations With ERA5, JRA55, and MERRA2 Reanalyses Using an Instrument Simulator

This study compares CL51 ceilometer observations made at Scott Base, Antarctica, with statistics from the ERA5, JRA55, and MERRA2 reanalyses. To enhance the comparison we use a lidar instrument simulator to derive cloud statistics from the reanalyses which account for instrumental factors. The cloud occurrence in the three reanalyses is slightly overestimated above 3 km, but displays a larger underestimation below 3 km relative to observations. Unlike previous studies, we see no relationship between relative humidity and cloud occurrence biases, suggesting that the cloud biases do not result from the representation of moisture. We also show that the seasonal variation of cloud occurrence and cloud fraction, defined as the vertically integrated cloud occurrence, are small in both the observations and the reanalyses. We also examine the quality of the cloud representation for a set of weather states derived from ERA5 surface winds. The variability associated with grouping cloud occurrence based on weather state is much larger than the seasonal variation, highlighting weather state is a strong control of cloud occurrence. All the reanalyses continue to display underestimates below 3 km and overestimates above 3 km for each weather state. But the variability in ERA5 statistics matches the changes in the observations better than the other reanalyses. We also use a machine learning scheme to estimate the quantity of supercooled liquid water cloud from the ceilometer observations. Ceilometer low-level supercooled liquid water cloud occurrences are considerably larger than values derived from the reanalyses, further highlighting the poor representation of low-level clouds in the reanalyses.

54 ENVIRONMENTAL SCIENCES↗

Uncertainty-Guided Prediction Horizon of Phase-Resolved Ocean Wave Forecasting Under Data Sparsity: Experimental and Numerical Evaluation

Accurate short-term wave forecasting is critical for the safe and efficient operation of marine structures that rely on real-time, phase-resolved ocean wave information for control and monitoring purposes (e.g., digital twins). These systems often depend on environmental sensors (e.g., waverider buoys, wave-sensing LIDAR). Challenges arise when upstream sensor data are missing, sparse, or phase-shifted due to drift. This study investigates the performance of two machine learning models, time-series dense encoder (TiDE) and long short-term memory (LSTM), for forecasting phase-resolved ocean surface elevations under varying degrees of data degradation. We introduce the τ-trimming algorithm, which adapts the prediction horizon based on uncertainty thresholds derived from historical forecasts. Numerical wave tank (NWT) and wave basin experiments are used to benchmark model performance under short- and long-term data masking, spatially coarse sensor grids, and upstream phase shifts. Results show under a 50% probability of upstream data loss, the τ-trimmed TiDE model achieves a 46% reduction in error at the most upstream target, compared to 22% for LSTM. Furthermore, phase misalignment in upstream data introduces a near-linear increase in forecast error. Under moderate model settings, a ±3 s misalignment increases the mean absolute error by approximately 0.5 m, while the same error is accumulated at ±4 s using the more conservative approach. These findings inform the design of resilient, uncertainty-aware wave forecasting systems suited for realistic offshore sensing environments.

42 ENGINEERING↗

Machine learning based noise reduction for satellite products: application to solar-induced fluorescence retrievals using simulated and real data

In the past two decades, global satellite measurements of terrestrial chlorophyll solar-induced fluorescence (SIF) have been used widely for a number of different applications related to physiology, phenology, and productivity of plants. However, SIF retrievals are inherently noisy due to the relatively small SIF spectral signature in comparison with observational noise. In this work, we examine how a spectral-based approach that employs principal component analysis along with a relatively shallow artificial neural network can be used to reduce noise and other artifacts in satellite level 2 (L2) products. We first apply the approach in a controlled environment in which radiance spectra are simulated with a full atmospheric and surface radiative transfer model for different scenarios including various SIF values that are known. Various levels of noise can be added to the simulated spectra. Resulting noisy and noise-reduced SIF retrievals are compared with the true values to assess performance. We then apply the noise reduction approach to real SIF derived from instruments flying on meteorological satellites. The results are evaluated by comparing SIF retrievals from different platforms with each other and with other independent data sets, showing enhanced capability to capture seasonal and interannual variability in SIF.

Chlorophyll fluorescence↗

Using intrahost single nucleotide variant data to predict SARS-CoV-2 detection cycle threshold values

Over the last four years, each successive wave of the COVID-19 pandemic has been caused by variants with mutations that improve the transmissibility of the virus. Despite this, we still lack tools for predicting clinically important features of the virus. In this study, we show that it is possible to predict the PCR cycle threshold (Ct) values from clinical detection assays using sequence data. Ct values often correspond with patient viral load and the epidemiological trajectory of the pandemic. Using a collection of 36,335 high quality genomes, we built models from SARS-CoV-2 intrahost single nucleotide variant (iSNV) data, computing XGBoost models from the frequencies of A, T, G, C, insertions, and deletions at each position relative to the Wuhan-Hu-1 reference genome. Our best model had an R 2 of 0.604 [0.593–0.616, 95% confidence interval] and a Root Mean Square Error (RMSE) of 5.247 [5.156–5.337], demonstrating modest predictive power. Overall, we show that the results are stable relative to an external holdout set of genomes selected from SRA and are robust to patient status and the detection instruments that were used. This study highlights the importance of developing modeling strategies that can be applied to publicly available genome sequence data for use in disease prevention and control.

COVID19↗

Cognitive Communications for NASA Space Systems

The growing complexity of spacecraft constellations, communication relay offerings, and mission architectures drives the need for the development of autonomous communication systems. NASA has traditionally launched single spacecraft missions that are served by the Space Communication and Navigation (SCaN) program. Operations on SCaN networks are typically scheduled weeks in advance, and often each asset serves a single user spacecraft at a time. Recent movement towards swarm missions could make the current approach unsustainable. Additionally, the integration of commercial communication service providers will substantially increase the data transfer options available to new missions. NASA science missions have found benefit in launching swarms of spacecraft, allowing coordinated simultaneous observations from different perspectives. Inter-spacecraft communication (mesh networking) is an enabler for this architecture, as are CubeSats that allow cost-effective provisioning of distributed mission assets. As more complex swarm missions launch, one challenge is coordinating communication within the swarm and choosing the appropriate mechanism for telemetry, tracking, control, and data services to and from Earth. Cognitive communications research conducted by SCaN aims to mitigate the increasing communication complexity for mission users by increasing the autonomy of links, networks, and service scheduling. By considering automation techniques including recent advances in artificial intelligence and machine learning, cognitive algorithms and related approaches enable increased mission science return, improved resource utilization for service provider networks, and resiliency in unpredictable or unplanned environments. The Cognitive Communications Project at the NASA Glenn Research Center develops applications of data-driven, non-deterministic methods to improve the autonomy of space communication. The project emphasizes development of decentralized space networks with artificial intelligence agents optimizing communication link throughput, data routing, and system-wide asset management. This paper discusses the objectives, approaches, and opportunities of the research to address growing needs of the space communications community.

Chelmins, David↗

Data and scripts associated with a manuscript analyzing ELM-FATES parameter sensitivity under pre-fire and postfire scenarios using machine learning

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Fire Severity-Dependent Shifts in Vegetation Parameter Sensitivity: A Pre- and Post-Fire Analysis Using ELM-FATES and Explainable AI” submitted to Journal of Advances in Modeling Earth Systems (Zahura et al. 2026). The study examines vegetation physiological parameters controlling pre-fire and post-fire vegetation dynamics. To support this analysis, 73 vegetation parameters in Functionally Assembled Terrestrial Ecosystem Simulator (FATES) (Fisher et al., 2018) , which is coupled with E3SM (Energy Exascale Earth System Model) land model (ELM, ELM-FATES), were perturbed using a Sobol sequence to generate 1,024 ensemble members for two plant functional types: needleleaf evergreen extratropical trees (NEET) and C3 grass. Simulations were conducted for the pre-fire period (2016) and post-fire period (2018–2023). Burn severity was represented by modifying the Nesterov index in FATES to 75,000, 150,000, and 300,000 for low, moderate, and high severity, respectively. A no-fire scenario was also included. Simulations were performed for 16 grid cells in the American River Watershed across different burn severities and plant functional types. XGBoost (eXtreme Gradient Boosting) models were trained using the parameter ensembles and ELM-FATES-simulated outputs, including leaf area index (LAI), gross primary productivity (GPP), aboveground biomass, vegetation evaporation, transpiration, and soil evaporation. Models were trained separately for each year and burn severity, followed by SHAP (SHapley Additive exPlanations) analysis to identify changes in dominant parameters after fire disturbance. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package contains the ELM-FATES simulation data. The scripts and data related to the analysis will be added later. The inputs and outputs from ELM-FATES are inside the “FATES” folder. “FATES_domain_surface” contains the domain and surface netcdfs that were used to run ELM-FATES in the study area. “FATES_parameters” contains the 1024 ensembles that were generated using Sobol sequence. “FATES_outputs” folder contains ELM-FATES simulated variables. All files are .csv and .nc (NetCDF).

Aboveground biomass↗

Parametric Mechanism Design Through Numerical Optimization and Physics Simulation

Design-Build-Test approaches for developing spaceflight hardware are prohibitively time and cost intensive and often lead to suboptimal mechanism designs. Approaches that couple machine learning and high-fidelity physics simulation could eliminate the need for hardware prototyping and dramatically accelerate the engineering design cycle, ultimately reducing cost. This work presents a modular NASA-developed toolchain to optimize hardware mechanisms in a virtual environment using numerical optimization and multi-body physics simulation. The toolchain enables multi-objective optimization, generates parametric CAD files that can be further post-processed by an end user, and can be expanded to optimize full systems and non-mechanical parameters such as feedback control variables. We demonstrate the toolchain through an independently verifiable design problem that optimizes wheel radius to achieve a desired linear velocity in a rigid-body physics environment when the wheel rotates at a constant angular speed, and then post-process the parametric CAD file of the optimal design generated by the tool before ultimately manufacturing it via 3D printing. We end with a discussion of how the toolchain can incorporate other analysis tools, including finite element analysis, computational fluid dynamics, and granular media simulations.

Optimization↗

An Optimization-Based Toolchain for Parametric Mechanism Design

Design-Build-Test approaches for developing spaceflight hardware are prohibitively time and cost intensive and often lead to suboptimal mechanism designs. Approaches that couple machine learning and high-fidelity physics simulation could eliminate the need for hardware prototyping and dramatically accelerate the engineering design cycle, ultimately reducing cost. This work presents a modular NASA-developed toolchain to optimize hardware mechanisms in a virtual environment using numerical optimization and multi-body physics simulation. The toolchain enables multi-objective optimization, generates parametric CAD files that can be further post-processed by an end user, and can be expanded to optimize full systems and non-mechanical parameters such as feedback control variables. We demonstrate the toolchain through an independently verifiable design problem that optimizes wheel radius to achieve a desired linear velocity in a rigid-body physics environment when the wheel rotates at a constant angular speed, and then post-process the parametric CAD file of the optimal design generated by the tool before ultimately manufacturing it via 3D printing. We end with a discussion of how the toolchain can incorporate other analysis tools, including finite element analysis, computational fluid dynamics, and granular media simulations.

Optimization↗

Remote sensing images, DEM, and point clouds associated with “Accuracy evaluation of cost-effective 3D reconstruction approaches for hydrobiogeochemical processes in non-perennial stream riverbeds”

This data package is associated with the publication “Accuracy evaluation of cost-effective 3D reconstruction approaches for hydrobiogeochemical processes in non-perennial stream riverbeds” published in Frontiers in Environmental Science, Environmental Informatics and Remote Sensing (Bao et al., 2026; doi: 10.3389/fenvs.2026.1725258). This data package includes the drone photos for a section of Umtanum Creek in Washington, Unted States. The photos were used to reconstruct the 3-dimensional (3D) digital elevation model (DEM) of the riverbed for the investigated stream section. The reconstruction results from four approaches are provided: (1) unoccupied aerial vehicle (UAV, colloquially known as drone) imagery-based Structure-from-Motion (SfM), (2) a machine learning-based 3D reconstruction model, Visual Geometry Grounded Deep Structure from Motion (VGGSfM), (3) Visual Geometry Grounded Transformer for long sequence of images (VGGT-Long), and (4) handheld smartphone LiDAR scanning. The ground truth measurements by tripod-mounted optical level kit and ground control points GPS locations for evaluating the accuracy of the four reconstruction approaches are also provided in this data package. A preliminary version of this data package was published in October 2025 at the time of manuscript submission. It was updated in March 2026, at the time of manuscript acceptance, to include additional metadata (this readme, data dictionary, and file level metadata). The data did not change. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) 8 folders; (2) the detailed flight configuration html files; (3) field metadata; (4) a readme; (5) a data dictionary; and (6) file-level metadata. The folders “2024_10_18_d01” and “2024_10_18_d02” contain the original drone photos for the two drone flights (d01 and d02) on October 18, 2024. The reconstruction results from each of the approaches are in the folders called “ODM_SfM”, “VGGSfM”, “VGGTLong”, and “LiDAR”. The ground truth measurements are in the folder called “optical_level_kit”. Lastly, results comparing the different approaches are in the folder called “comparisons”. All files are .csv, .html, .jpg, .obj, .txt, and .npy. For information on using the .obj and .npy files, see the readme files within the same folder as the files.

54 ENVIRONMENTAL SCIENCES↗

Enabling the Next Generation of Smart Sensors in Coal Fired Power Plants using Cellular 5G Technology

An important need for coal fired power plants is the ability to monitor multiple systems with ease and accuracy. Common implementations of these monitoring systems come with drawbacks due to the nature of coal fired power plants. Harsh environments, High Temperatures, and lots of RF (Radio Frequency) noise can create issues for accurately recording and transmitting data across wireless signals. In addition, as renewable energy sources come online, existing fossil fueled plants will need to operate more flexibly with their maintenance schedules outside of standard conditions. Therefore, additional sensing and control mechanisms need placed in existing plants to provide operators with more information such that maintenance decisions can be made well in advance of failures. A solution to this problem is the Next Generation of Smart Sensors, which leverages the power of 5G cellular signals and machine learning to overcome the myriad of problems with current implementations

20 FOSSIL-FUELED POWER PLANTS↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Enhancing Air Traffic Control Planning with Automatic Speech Recognition

The decisions made during the Federal Aviation Administration Air Traffic Control System Command Center's planning teleconferences hold significant sway over the National Airspace System. Held every two hours, these teleconferences convene air traffic managers and stakeholders from across the nation to discuss airspace conditions, weather, and constraints, leading to the formulation and adjustment of traffic management initiatives. Given the critical nature of these decisions, the need for accurate and efficient record-keeping is paramount. In recent years, the application of automatic speech recognition has gained popularity across diverse industries, including aviation. While traditional applications focus on transcribing air traffic control communication, this paper explores a unique application of automatic speech recognition by converting the audio from planning teleconferences into text transcriptions. This innovative approach addresses key challenges in the field, presenting potential benefits for quality assurance, real-time participation, and downstream natural language processing tasks. A notable breakthrough in the machine learning community, namely the transformer neural network architecture, forms the backbone of the proposed solution in this paper. The transformer architecture's role in this research represents a paradigm shift in the efficiency of automatic speech recognition models. By reducing the amount of in-domain training data required, this architecture allows for the fine-tuning of such models like Whisper, originally pretrained on vast English speech datasets. The adaptability of the transformer architecture proves invaluable in capturing the nuances of aviation terminology and specific language used in planning teleconferences. Leveraging the Whisper model as a baseline, our research details the fine-tuning and validation using a dataset comprising 20 hours of meticulously transcribed planning teleconferences. Notably, the baseline pretrained Whisper model exhibited a word error rate of 18.77%. Through the fine-tuning process, the model achieved a substantial improvement, demonstrating an impressive performance with a reduced word error rate of 6.82%. This substantial decrease in WER not only highlights the effectiveness of the transformer architecture but also emphasizes the practical advancements achieved through the application of automatic speech recognition in this specific domain. The utilization of automatic speech recognition in planning teleconferences in this work introduces several novelties. Firstly, the creation of text transcriptions offers a valuable tool for quality assurance and facilitates the efficient review of teleconferences. This is an important aspect of the proposed solution, given the time-sensitive and high-stakes nature of decisions made during these meetings. Furthermore, text-searchable transcriptions provide a streamlined approach for locating and validating critical information, potentially saving hours of manual effort in searching through audio recordings. Moreover, our research identifies a key use case for external facilities and stakeholders. In situations where attendance at the planning teleconference is not feasible, having access to text transcriptions in real-time or shortly after the teleconference ends, proves to be a time-saving and informative resource. This feature enhances collaboration and ensures that stakeholders can stay abreast of important discussions and decisions even in their absence. Despite the efficiency gains facilitated by the transformer architecture in automatic speech recognition technology, it is essential to acknowledge the human factors in data creation. Subject matter experts play a crucial role in accurately transcribing planning teleconferences due to the specificity and complexity of the information discussed. The research dataset, consisting of 20 hours of transcribed planning teleconferences, forms the foundation for fine-tuning and validating the Whisper model. The achieved word error rate of 6.82% demonstrates promising advancements, particularly in recognizing essential aviation terminology within the teleconferences. In conclusion, this paper presents a comprehensive exploration of the application of automatic speech recognition in Air Traffic Control System Command Center planning teleconferences, leveraging the transformer architecture for enhanced efficiency. The novel contributions lie in the improved accessibility of decision-making records, real-time participation opportunities for external stakeholders, and the potential for downstream natural language processing advancements. As the aviation industry continues to evolve, the integration of automatic speech recognition technologies holds the promise of revolutionizing decision-making processes and contributing to the overall safety and efficiency of air traffic management.

ATM↗

Source Analysis of Ozone Pollution in Liaoyuan City’s Atmosphere Based on Machine Learning Models and HYSPLIT Clustering Method

Firstly, this study investigates the spatiotemporal distribution characteristics of the ozone (O 3 ) pollution in Liaoyuan City using monitoring data from 2015 to 2024. Then, three machine learning models (ML)—random forest (RF), support vector machine (SVM), and artificial neural network (ANN)—are employed to quantify the influence of meteorological and non-meteorological factors on O 3 concentrations. Finally, the HYSPLIT clustering method and CMAQ model are utilized to analyze inter-regional transport characteristics, identifying the causes of O 3 pollution. The results indicate that O 3 pollution in Liaoyuan exhibits a distinct seasonal pattern, with the highest concentrations found in spring and summer, peaking in the afternoon. Among the three ML models, the random forest model demonstrates the best predictive performance (R 2 = 0.9043). Feature importance identifies NO 2 as the primary driving factor, followed by meteorological conditions in the second quarter and land surface characteristics. Furthermore, regional transport significantly contributes to O 3 pollution, with approximately 80% of air mass trajectories in heavily polluted episodes originating from adjacent industrial areas and the sea. The combined effects of transboundary precursors and O 3 transport with local emissions and meteorological conditions further increase the O 3 pollution level. This study highlights the need to strengthen coordinated NO X and VOCs emission reductions and enhance regional joint prevention and control strategies in China.

HYSPLIT clustering↗

EFIT-Prime: Probabilistic and physics-constrained reduced-order neural network model for equilibrium reconstruction in DIII-D

We introduce EFIT-Prime, a novel machine learning surrogate model for EFIT (Equilibrium FIT) that integrates probabilistic and physics-informed methodologies to overcome typical limitations associated with deterministic and ad hoc neural network architectures. EFIT-Prime utilizes a neural architecture search-based deep ensemble for robust uncertainty quantification, providing scalable and efficient neural architectures that comprehensively quantify both data and model uncertainties. Physically informed by the Grad–Shafranov equation, EFIT-Prime applies a constraint on the current density J tor and a smoothness constraint on the first derivative of the poloidal flux, ensuring physically plausible solutions. Furthermore, the spatial location of the diagnostics is explicitly incorporated in the inputs to account for their spatial correlation. Extensive evaluations demonstrate EFIT-Prime's accuracy and robustness across diverse scenarios, most notably showing good generalization on negative-triangularity discharges that were excluded from training. Timing studies indicate an ensemble inference time of 15 ms for predicting a new equilibrium, offering the possibility of plasma control in real-time, if the model is optimized for speed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Predictive links between microbial communities and biological oxygen utilization in the Arctic Ocean

Microbial metabolism influences rates of net community production (NCP), exerting a direct biological control on marine oxygen and carbon fluxes. In the Arctic, it is increasingly important to understand and quantify this process, as ecological and oceanographic conditions shift due to changing climate. Here, we describe potential ecological links between pelagic microbial diversity and an NCP precursor, biological oxygen utilization, using machine learning and paired observations of community structure and metabolic activity from a seasonally and spatially variable transect of the Arctic Ocean (2019–2020 MOSAiC Expedition). Community structure was determined using 16S (prokaryotic) and 18S (eukaryotic) rRNA gene amplicon sequencing, and metabolic activity was derived from ΔO 2 /Ar. Using self-organizing maps, we identified clear successional patterns in observed microbial community structure that were seasonally driven in the upper ocean and vertically stratified with depth. Metabolic activity was also stratified, with a primarily net heterotrophic water column (median −1.5% biological oxygen saturation), excepting periodic oxygen supersaturation (maximum: 13.6%) within the mixed layer. Using DNA sequences as predictor variables, we then constructed a random forest regression model that reliably reconstructed biological oxygen concentrations (root mean squared error = 4.14 μmol kg −1 ). Top predictors from this model were from heterotrophic (bacteria) or potentially mixotrophic (dinoflagellate) taxa. These analyses highlight biologically driven diagnostic tools that can be used to expand biogeochemical datasets and improve the microbial perspectives and metabolisms represented in ecological models of net productivity and carbon flux in a changing Arctic Ocean.

Chamberlain, Emelia J. [Univ. of San Diego, San Di↗

Correlating Protein Dynamics and Catalytic Activity of a Model Hydrogenase Using Paramagnetic and Biological Nuclear Magnetic Resonance Spectroscopy

Rational catalyst design remains a significant challenge, with electronic structure, steric, and electrostatic effects known to contribute to activity. Recently, dynamics has been recognized as another factor that impacts catalysis, though identifying and predicting these effects has remained out of reach. Nickel-substituted rubredoxin (NiRd), a protein-based mimic of a hydrogenase enzyme, serves as a model catalytic system in which dynamics can be systematically investigated with respect to activity. While over 30 secondary-sphere mutants of NiRd have been shown to be catalytically active, no significant correlation was observed between the rates and catalytic overpotential or electronic structure, prompting questions about the protein-derived factors that modulate activity. Here, in this work, NMR spectroscopy was used to investigate the roles of substrate accessibility, protein dynamics, and protein stability in controlling catalysis. Significant paramagnetic effects from the nickel center (S = 1) isolate the methylene proton resonances of the metal-coordinating cysteine residues. The sensitivity of resonance positions and linewidths to local environment offers an opportunity to study dynamical molecular changes around the metal center with high resolution. Machine learning algorithms were employed to identify correlations between the catalytic activity and the paramagnetic NMR spectra. These analyses revealed spectroscopic features of specific cysteine protons that report on catalytic overpotential and increased turnover rates, which are further supported by the results obtained using high-field NMR techniques. Collectively, these studies indicate the potential for multifrequency NMR techniques to resolve key contributors to catalytic activity and highlight the importance of local and outer-sphere dynamics.

Protein Engineering↗

High-Fidelity Building Emulator

This dataset provides high-fidelity time series data for an emulated commercial office building sited in the Chicago, IL area during a Typical Meteorological Year (TMY). This dataset consists of air-side HVAC measurements and control inputs, and it includes normal operations as well as various implemented faults (with associated ground truth measurements) implemented on selected days. This data could be used to quantify and compare the impacts of different faults, and it could also be used as training or validation data for machine learning algorithms (e.g., reduced-order modelling, fault detection and diagnosis).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗