Search NASASearch

SEARCH · Search NASA

Results for “ensemble data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Influence of Local Water Vapor Analysis Uncertainty on Ensemble Forecasts of Tropical Cyclogenesis Using Hurricane Irma (2017) as a Testbed

Abstract Tropical cyclone formation is known to require abundant water vapor in the lower to middle troposphere within the incipient disturbance. In this study, we assess the impacts of local water vapor analysis uncertainty on the predictability of the formation of Hurricane Irma (2017). To this end, we reduce the magnitude of the incipient disturbance’s water vapor perturbations obtained from an ensemble-based data assimilation system that constrained moisture by assimilating all-sky infrared and microwave radiances. Five-day ensemble forecasts are initialized two days before genesis using each set of modified analysis perturbations. Growth of convective differences and intensity uncertainty are evaluated for each ensemble forecast. We observe that when initializing an ensemble forecast with only moisture uncertainty within the incipient disturbance, the resulting intensity uncertainty at every lead time exceeds half that of an ensemble containing initial perturbations to all variables throughout the domain. Although ensembles with different initial moisture uncertainty amplitudes reveal a similar pathway to genesis, uncertainty in genesis timing varies substantially across ensembles since moister members exhibit earlier spinup of the low-level vortex. These differences in genesis timing are traced back to the first 6–12 h of integration, when differences in the position and intensity of mesoscale convective systems across ensemble members develop more quickly with greater initial moisture uncertainty. In addition, the rapid growth of intensity uncertainty may be greatly modulated by the diurnal cycle. Ultimately, this study underscores the importance of targeting the incipient disturbance with high spatiotemporal water vapor observations for ingestion into data assimilation systems. Significance Statement Hurricanes form from clusters of thunderstorms that organize into a coherent system. One of the key ingredients for the formation process is an abundance of moisture. In this study, we test the sensitivity of hurricane formation to the initial moisture content in the vicinity of the cluster of thunderstorms that would become Hurricane Irma (2017). To do so, we initialize sets of forecasts each having a different variability of initial moisture content within the embryonic disturbance. Our results show that the predictability of hurricane formation is highly dependent on the uncertainty of the moisture content within the initial disturbance. Consequently, more high-quality observations of the moisture within the precursor disturbances to hurricanes are expected to improve forecasts of their formation.

Hartman, Christopher M.

Reconstruction of the 1997/1998 El Nino from TOPEX/POSEIDON and TOGA/TAO Data Using a Massively Parallel Pacific-Ocean Model and Ensemble Kalman Filter

Two massively parallel data assimilation systems in which the model forecast-error covariances are estimated from the distribution of an ensemble of model integrations are applied to the assimilation of 97-98 TOPEX/POSEIDON altimetry and TOGA/TAO temperature data into a Pacific basin version the NASA Seasonal to Interannual Prediction Project (NSIPP)ls quasi-isopycnal ocean general circulation model. in the first system, ensemble of model runs forced by an ensemble of atmospheric model simulations is used to calculate asymptotic error statistics. The data assimilation then occurs in the reduced phase space spanned by the corresponding leading empirical orthogonal functions. The second system is an ensemble Kalman filter in which new error statistics are computed during each assimilation cycle from the time-dependent ensemble distribution. The data assimilation experiments are conducted on NSIPP's 512-processor CRAY T3E. The two data assimilation systems are validated by withholding part of the data and quantifying the extent to which the withheld information can be inferred from the assimilation of the remaining data. The pros and cons of each system are discussed.

Keppenne, C. L.

Global Assimilation of Multi-Sensor Snow Observations for Improved Characterization of Snow Processes

Snow conditions on the land surface are recognized to be key components of the global hydrological cycle as they play a critical role in the determination of local and regional climate. In many mid-latitude and high-latitude regions, the seasonal water storage and associated spring snowmelt dominate the local hydrology. The contribution to the runoff and moisture conditions from snow is vital in supporting agriculture and in determining water resources management practices. Consequently, accurate characterization of snow properties becomes important for both end-use applications and weather and climate research. Recently a joint effort between the u.S. Air Force and NASA has enabled a blended, multi-sensor snow product known as the AFWA NASA Snow Algorithm (ANSA). This global snow dataset has been generated by utilizing the Earth Observation System (EOS) Moderate Resolution Imaging Spectroradiometer (MODIS) and Advanced Microwave Scanning Radiometer for EOS (AMSR-E) datasets. ANSA product includes estimates of snow cover extent, snow water equivalent (SWE) and SWE-derived snow depth fields. The MODIS-based products enable snow cover mappings under cloud-free conditions whereas the passive microwave data from AMSR-E provides measurements under cloudy conditions. These remotely-sensed snow observations are further augmented with the information from ground-based snow measurements through data fusion techniques. The resulting ANSA products are employed in the NASA Land Information System (LIS) data assimilation framework, which provides a comprehensive environment for integrating community land surface models, ground and satellite-based observations, and ensemble-based data assimilation tools. LIS incorporates the multisensor ANSA snow retrievals with the land surface model estimates to generate spatially and temporally continuous estimates of snow states, through data assimilation. A suite of experiments to assimilate ANSA snow cover, SWE and snow depth estimates with different land surface models in LIS are conducted and the resulting estimates of snow conditions are evaluated against a number of in-situ observational datasets, over several regions of the world. These evaluations are used to compare and contrast the advantages and disadvantages of these multi-sensor snow observations.

Kumar, Sujay

Uncertainty guided online ensemble for non-stationary data streams in fusion science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior with distribution drifts, resulted by both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with such non-stationary data streams. Online learning techniques have been leveraged in other domains, however it has been largely unexplored for fusion applications. In this paper, we investigate online learning for continuous adaptation to drifting data streams in the prediction of Toroidal Field (TF) coils deflection at the DIII-D fusion facility. We further address the short-term performance degradation inherent to standard online learning, which arises because ground truth is unavailable at prediction time. To mitigate this issue, we propose an uncertainty-guided online ensemble framework. The method leverages the Deep Gaussian Process Approximation (DGPA) for calibrated uncertainty estimation and uses these uncertainty measures to guide a meta-algorithm that aggregates predictions from learners trained over different historical horizons. Our results show that online learning reduces prediction error by 80% compared to a static model. The online ensemble and the proposed uncertainty-guided ensemble further reduce error by approximately 6%, and 10% respectively, relative to standard single-model online learning, while also providing calibrated uncertainty estimates to support operational decision-making.

AI

Assessment of Mars Atmospheric Temperature Retrievals from the Thermal Emission Spectrometer Radiances

Motivated by the needs of Mars data assimilation. particularly quantification of measurement errors and generation of averaging kernels. we have evaluated atmospheric temperature retrievals from Mars Global Surveyor (MGS) Thermal Emission Spectrometer (TES) radiances. Multiple sets of retrievals have been considered in this study; (1) retrievals available from the Planetary Data System (PDS), (2) retrievals based on variants of the retrieval algorithm used to generate the PDS retrievals, and (3) retrievals produced using the Mars 1-Dimensional Retrieval (M1R) algorithm based on the Optimal Spectral Sampling (OSS ) forward model. The retrieved temperature profiles are compared to the MGS Radio Science (RS) temperature profiles. For the samples tested, the M1R temperature profiles can be made to agree within 2 K with the RS temperature profiles, but only after tuning the prior and error statistics. Use of a global prior that does not take into account the seasonal dependence leads errors of up 6 K. In polar samples. errors relative to the RS temperature profiles are even larger. In these samples, the PDS temperature profiles also exhibit a poor fit with RS temperatures. This fit is worse than reported in previous studies, indicating that the lack of fit is due to a bias correction to TES radiances implemented after 2004. To explain the differences between the PDS and Ml R temperatures, the algorithms are compared directly, with the OSS forward model inserted into the PDS algorithm. Factors such as the filtering parameter, the use of linear versus nonlinear constrained inversion, and the choice of the forward model, are found to contribute heavily to the differences in the temperature profiles retrieved in the polar regions, resulting in uncertainties of up to 6 K. Even outside the poles, changes in the a priori statistics result in different profile shapes which all fit the radiances within the specified error. The importance of the a priori statistics prevents reliable global retrievals based a single a priori and strongly implies that a robust science analysis must instead rely on retrievals employing localized a priori information, for example from an ensemble based data assimilation system such as the Local Ensemble Transform Kalman Filter (LETKF).

Hoffman, Matthew J.

Mars gravity field derived from Viking-1 and Viking-2 - The navigation result

Viking-1 and Viking-2 Doppler tracking data taken during orbit phases characterized by 1500 km subperiapse altitudes have provided a basis for a determination of the Martian gravity field. Navigation results show that the linear combination of short-arc gravity estimates is an acceptable technique for obtaining gravity models over multiple data arcs. An ensemble field composed of Viking data and Mariner-9 a priori retains the inherent local accuracy of its constituent fields. At the same time, the model can be made to be valid globally by careful weighting of a priori Mariner-9 data. The sixth degree and order model presented reduces the error concerning the change in period by more than an order of magnitude during the high altitude (1500 km) phases of the Viking mission. The resulting areoid deviates by no more than 150 m from the areoid produced by the a priori Mariner-9 field.

Christensen, E. J.

Aerodynamic parameters of the X-31 drop model estimated from flight-data at high angles of attack

Lateral aerodynamic parameters of the X-31 drop model were estimated from flight data at angles of attack between 25 deg and 45 deg. Partitioned data from an ensemble of 12 maneuvers and data from 13 single maneuvers were analyzed by a stepwise regression technique to obtain an aerodynamic model structure and least squares parameter estimates. Because of data collinearity in several maneuvers, these maneuvers were reanalyzed by two biased estimation techniques, mixed estimation and fractional rank regression. The final parameter estimates in the form of stability and control derivatives were plotted against the angle of attack and compared with wind tunnel results and a limited number of estimates from full-scale aircraft data. There was no significant disagreement between parameters from the two sets of drop model data and the full-scale aircraft data. Some differences, however, existed between the dihedral, damping-in-roll, and aileron-effectiveness parameters from flight and wind tunnel data.

Klein, Vladislav

Uncertainty based Online Ensemble on Non-Stationary Data for Fusion Science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior due to drifts in the data. The drifts can arise from both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with non-stationary data streams.Online learning can be used to continuously adapt the models with new data as it is acquired. However, traditional online learning can suffer from short-term performance degradation, as ground truth are not available before making the prediction. To address this challenge, we propose uncertainty aware ensemble approach for online learning. We use Deep Gaussian Process Approximation (DGPA) technique for calibrated uncertainty estimation and use the uncertainty values to guide a meta-algorithm that produces predictions based on ensemble of learners. Moreover, DGPA also provides uncertainty estimation along with the predictions for decision makers. This paper demonstrates that the proposed method outperforms traditional online learning approach, and a naive ensemble without uncertainty guidance by about 7% and 6%, respectively, on B-coil deflection prediction at DIII-D Fusion Facility.

Rajput, Kishansingh [Thomas Jefferson National Acc

Uncertainty based Online Ensemble on Non-Stationary Data for Fusion Science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior due to drifts in the data. The drifts can arise from both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with non-stationary data streams.Online learning can be used to continuously adapt the models with new data as it is acquired. However, traditional online learning can suffer from short-term performance degradation, as ground truth are not available before making the prediction. To address this challenge, we propose uncertainty aware ensemble approach for online learning. We use Deep Gaussian Process Approximation (DGPA) technique for calibrated uncertainty estimation and use the uncertainty values to guide a meta-algorithm that produces predictions based on ensemble of learners. Moreover, DGPA also provides uncertainty estimation along with the predictions for decision makers. This paper demonstrates that the proposed method outperforms traditional online learning approach, and a naive ensemble without uncertainty guidance by about 7% and 6%, respectively, on B-coil deflection prediction at DIII-D Fusion Facility.

Rajput, Kishansingh [Thomas Jefferson National Acc

Assimilation of MODIS Snow Cover Through the Data Assimilation Research Testbed and the Community Land Model Version 4

To improve snowpack estimates in Community Land Model version 4 (CLM4), the Moderate Resolution Imaging Spectroradiometer (MODIS) snow cover fraction (SCF) was assimilated into the Community Land Model version 4 (CLM4) via the Data Assimilation Research Testbed (DART). The interface between CLM4 and DART is a flexible, extensible approach to land surface data assimilation. This data assimilation system has a large ensemble (80-member) atmospheric forcing that facilitates ensemble-based land data assimilation. We use 40 randomly chosen forcing members to drive 40 CLM members as a compromise between computational cost and the data assimilation performance. The localization distance, a parameter in DART, was tuned to optimize the data assimilation performance at the global scale. Snow water equivalent (SWE) and snow depth are adjusted via the ensemble adjustment Kalman filter, particularly in regions with large SCF variability. The root-mean-square error of the forecast SCF against MODIS SCF is largely reduced. In DJF (December-January-February), the discrepancy between MODIS and CLM4 is broadly ameliorated in the lower-middle latitudes (2345N). Only minimal modifications are made in the higher-middle (4566N) and high latitudes, part of which is due to the agreement between model and observation when snow cover is nearly 100. In some regions it also reveals that CLM4-modeled snow cover lacks heterogeneous features compared to MODIS. In MAM (March-April-May), adjustments to snowmove poleward mainly due to the northward movement of the snowline (i.e., where largest SCF uncertainty is and SCF assimilation has the greatest impact). The effectiveness of data assimilation also varies with vegetation types, with mixed performance over forest regions and consistently good performance over grass, which can partly be explained by the linearity of the relationship between SCF and SWE in the model ensembles. The updated snow depth was compared to the Canadian Meteorological Center (CMC) data. Differences between CMC and CLM4 are generally reduced in densely monitored regions.

data assimilation

Technical Report Series on Global Modeling and Data Assimilation: Soil Moisture Active Passive (SMAP) Project Assessment Report for the Beta-Release L4_SM Data Product - Volume 40

During the post-launch SMAP calibration and validation (Cal/Val) phase there are two objectives for each science data product team: 1) calibrate, verify, and improve the performance of the science algorithm, and 2) validate the accuracy of the science data product as specified in the science requirements and according to the Cal/Val schedule. This report provides an assessment of the SMAP Level 4 Surface and Root Zone Soil Moisture Passive (L4_SM) product specifically for the product's public beta release scheduled for 30 October 2015. The primary objective of the beta release is to allow users to familiarize themselves with the data product before the validated product becomes available. The beta release also allows users to conduct their own assessment of the data and to provide feedback to the L4_SM science data product team. The assessment of the L4_SM data product includes comparisons of SMAP L4_SM soil moisture estimates with in situ soil moisture observations from core validation sites and sparse networks. The assessment further includes a global evaluation of the internal diagnostics from the ensemble-based data assimilation system that is used to generate the L4_SM product. This evaluation focuses on the statistics of the observation-minus-forecast (O-F) residuals and the analysis increments. Together, the core validation site comparisons and the statistics of the assimilation diagnostics are considered primary validation methodologies for the L4_SM product. Comparisons against in situ measurements from regional-scale sparse networks are considered a secondary validation methodology because such in situ measurements are subject to upscaling errors from the point-scale to the grid cell scale of the data product. Based on the limited set of core validation sites, the assessment presented here meets the criteria established by the Committee on Earth Observing Satellites for Stage 1 validation and supports the beta release of the data. The validation against sparse network measurements and the evaluation of the assimilation diagnostics address Stage 2 validation criteria by expanding the assessment to regional and global scales.

ubRMSE

Soil Moisture Active Passive Mission L4_SM Data Product Assessment (Version 2 Validated Release)

During the post-launch SMAP calibration and validation (Cal/Val) phase there are two objectives for each science data product team: 1) calibrate, verify, and improve the performance of the science algorithm, and 2) validate the accuracy of the science data product as specified in the science requirements and according to the Cal/Val schedule. This report provides an assessment of the SMAP Level 4 Surface and Root Zone Soil Moisture Passive (L4_SM) product specifically for the product's public Version 2 validated release scheduled for 29 April 2016. The assessment of the Version 2 L4_SM data product includes comparisons of SMAP L4_SM soil moisture estimates with in situ soil moisture observations from core validation sites and sparse networks. The assessment further includes a global evaluation of the internal diagnostics from the ensemble-based data assimilation system that is used to generate the L4_SM product. This evaluation focuses on the statistics of the observation-minus-forecast (O-F) residuals and the analysis increments. Together, the core validation site comparisons and the statistics of the assimilation diagnostics are considered primary validation methodologies for the L4_SM product. Comparisons against in situ measurements from regional-scale sparse networks are considered a secondary validation methodology because such in situ measurements are subject to up-scaling errors from the point-scale to the grid cell scale of the data product. Based on the limited set of core validation sites, the wide geographic range of the sparse network sites, and the global assessment of the assimilation diagnostics, the assessment presented here meets the criteria established by the Committee on Earth Observing Satellites for Stage 2 validation and supports the validated release of the data. An analysis of the time average surface and root zone soil moisture shows that the global pattern of arid and humid regions are captured by the L4_SM estimates. Results from the core validation site comparisons indicate that "Version 2" of the L4_SM data product meets the self-imposed L4_SM accuracy requirement, which is formulated in terms of the ubRMSE: the RMSE (Root Mean Square Error) after removal of the long-term mean difference. The overall ubRMSE of the 3-hourly L4_SM surface soil moisture at the 9 km scale is 0.035 cubic meters per cubic meter requirement. The corresponding ubRMSE for L4_SM root zone soil moisture is 0.024 cubic meters per cubic meter requirement. Both of these metrics are comfortably below the 0.04 cubic meters per cubic meter requirement. The L4_SM estimates are an improvement over estimates from a model-only SMAP Nature Run version 4 (NRv4), which demonstrates the beneficial impact of the SMAP brightness temperature data. L4_SM surface soil moisture estimates are consistently more skillful than NRv4 estimates, although not by a statistically significant margin. The lack of statistical significance is not surprising given the limited data record available to date. Root zone soil moisture estimates from L4_SM and NRv4 have similar skill. Results from comparisons of the L4_SM product to in situ measurements from nearly 400 sparse network sites corroborate the core validation site results. The instantaneous soil moisture and soil temperature analysis increments are within a reasonable range and result in spatially smooth soil moisture analyses. The O-F residuals exhibit only small biases on the order of 1-3 degrees Kelvin between the (re-scaled) SMAP brightness temperature observations and the L4_SM model forecast, which indicates that the assimilation system is largely unbiased. The spatially averaged time series standard deviation of the O-F residuals is 5.9 degrees Kelvin, which reduces to 4.0 degrees Kelvin for the observation-minus-analysis (O-A) residuals, reflecting the impact of the SMAP observations on the L4_SM system. Averaged globally, the time series standard deviation of the normalized O-F residuals is close to unity, which would suggest that the magnitude of the modeled errors approximately reflects that of the actual errors. The assessment report also notes several limitations of the "Version 2" L4_SM data product and science algorithm calibration that will be addressed in future releases. Regionally, the time series standard deviation of the normalized O-F residuals deviates considerably from unity, which indicates that the L4_SM assimilation algorithm either over- or under-estimates the actual errors that are present in the system. Planned improvements include revised land model parameters, revised error parameters for the land model and the assimilated SMAP observations, and revised surface meteorological forcing data for the operational period and underlying climatological data. Moreover, a refined analysis of the impact of SMAP observations will be facilitated by the construction of additional variants of the model-only reference data. Nevertheless, the “Version 2” validated release of the L4_SM product is sufficiently mature and of adequate quality for distribution to and use by the larger science and application communities.

SMAP L4_SM

Application of laser velocimetry to unsteady flows in large scale, high speed tunnels

The optical, seeding, and data reduction procedures which have been used in the successful application of laser velocimetry to large-scale, unsteady, transonic wind-tunnel testing are described. Flowfield measurements obtained at several facilities are presented. These results include vortex flow measurements and ensemble-averaged data obtained under conditions of vortex-shedding, dynamic stall airfoil oscillation, oscillating trailing-edge control flaps, and helicopter rotor flow fields.

Owen, F. K.

Soil Moisture Active Passive (SMAP) Project Assessment Report for Version 4 of the L4_SM Data Product

This report provides an assessment of Version 4 of the SMAP Level 4 Surface and Root Zone Soil Moisture (L4_SM) product, released on 14 June 2018. The assessment includes comparisons of L4_SM soil moisture and temperature estimates with in situ measurements from core validation sites and sparse networks. The assessment further includes a global evaluation of the internal diagnostics from the ensemble-based data assimilation system that is used to generate the L4_SM product, including observation-minus-forecast (O-F) brightness temperature residuals and soil moisture analysis increments.Together, the core validation site comparisons and the statistics of the assimilation diagnostics areconsidered primary validation methodologies for the L4_SM product. Comparisons against in situ measurements from regional-scale sparse networks are considered a secondary validation methodology because such in situ measurements are subject to upscaling errors from the point-scale to the grid-cell scale of the data product.The Version 4 L4_SM product benefits from an improved land surface modeling system and from retrospective surface meteorological forcing data that are as consistent as possible with the present-day datain terms of their climatology. Specifically, the model changes include revised parameters and parameterizations for (i) the surface energy balance, (ii) recharge from below of the model's surface excess reservoir, and (iii) the snow depletion curve. Updated ancillary inputs include improved datasets for landcover, topography, and vegetation height. The Version 4 algorithm further includes a revised approach to precipitation corrections that improves the precipitation climatology in Africa and the high-latitudes. Moreover, for system calibration the model is forced retrospectively with MERRA-2 reanalysis data, which are more consistent with the near-real time GEOS forward processing (FP) data used during the SMAP period than the retrospective GEOS data that were available for previous L4_SM versions. An analysis of the time-average surface and root zone soil moisture shows that the global pattern ofarid and humid regions is captured by the Version 4 L4_SM estimates. Owing to the changes in the landsurface modeling system, surface soil moisture is typically drier by several volumetric percent in Version 4 compared to Version 3, whereas root zone soil moisture is wetter in Version 4 in some regions and drierin others. Because of these climatological differences, the Version 3 and Version 4 products should not be combined into a single dataset for use in applications.Results from the core validation site comparisons indicate that Version 4 of the L4_SM data product meets the self-imposed L4_SM accuracy requirement, which is formulated in terms of the RMSE after removal of the long-term mean difference (ubRMSE). The overall ubRMSE of the 3-hourly L4_SM dataat the 9 km scale is 0.039 m3 m-3 for surface soil moisture and 0.029 m3 m-3 for root zone soil moisture,below the 0.04 m3 m-3 requirement. The L4_SM estimates are an improvement over estimates from a model-only Nature Run version 7.2 (NRv7.2), which demonstrates the beneficial impact of the SMAP brightness temperature data. Overall, L4_SM surface and root zone soil moisture estimates are more skillful than NRv7.2 estimates, with statistically significant improvements at the 5% level for surface soil moisture R and anomaly R values. Results from comparisons of the L4_SM product to i

Reichle, Rolf H.

On learning what to learn: Heterogeneous observations of dynamics and establishing possibly causal relations among them

Abstract Before we attempt to (approximately) learn a function between two sets of observables of a physical process, we must first decide what the inputs and outputs of the desired function are going to be. Here we demonstrate two distinct, data-driven ways of first deciding “the right quantities” to relate through such a function, and then proceeding to learn it. This is accomplished by first processing simultaneous heterogeneous data streams (ensembles of time series) from observations of a physical system: records of multiple observation processes of the system. We determine (i) what subsets of observables are common between the observation processes (and therefore observable from each other, relatable through a function); and (ii) what information is unrelated to these common observables, therefore particular to each observation process, and not contributing to the desired function. Any data-driven technique can subsequently be used to learn the input–output relation—from k-nearest neighbors and Geometric Harmonics to Gaussian Processes and Neural Networks. Two particular “twists” of the approach are discussed. The first has to do with the identifiability of particular quantities of interest from the measurements. We now construct mappings from a single set of observations from one process to entire level sets of measurements of the second process, consistent with this single set. The second attempts to relate our framework to a form of causality: if one of the observation processes measures “now,” while the second observation process measures “in the future,” the function to be learned among what is common across observation processes constitutes a dynamical model for the system evolution.

Sroczynski, David W.

Using Multiple Isotope-Labeled Infrared Spectra for the Structural Characterization of an Intrinsically Disordered Peptide

Intrinsically disordered proteins (IDPs) rapidly interconvert between conformers, requiring an ensemble description. This complicates their experimental characterization, and force field limitations pose challenges for their simulation. Here, in this work, we use isotope-labeled and unlabeled infrared (IR) spectra to reweight simulated ensembles of the elastin-like peptide GVGVPGVG, a paradigmatic disordered peptide. By comparing the results obtained with different spectra, we explicitly show that the weights are underdetermined by the ensemble averaged data. We identify which labels and frequency regions maximize structural information while minimizing sensitivity to simulation error and show that these regions report on whether the peptide makes specific interactions. Our work shows the importance of incorporating simulations and simulated spectra at the planning stages of isotope-labeled IR experiments and more generally provides a framework for interpreting IR data for IDPs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH