Search NASA⌕ Search

SEARCH · Search NASA

Results for “ARIMA”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Utilizing physics-based input features within a machine learning model to predict wind speed forecasting error

Machine learning is quickly becoming a commonly used technique for wind speed and power forecasting. Many machine learning methods utilize exogenous variables as input features, but there remains the question of which atmospheric variables are most beneficial for forecasting, especially in handling non-linearities that lead to forecasting error. This question is addressed via creation of a hybrid model that utilizes an autoregressive integrated moving-average (ARIMA) model to make an initial wind speed forecast followed by a random forest model that attempts to predict the ARIMA forecasting error using knowledge of exogenous atmospheric variables. Variables conveying information about atmospheric stability and turbulence as well as inertial forcing are found to be useful in dealing with non-linear error prediction. Streamwise wind speed, time of day, turbulence intensity, turbulent heat flux, vertical velocity, and wind direction are found to be particularly useful when used in unison for hourly and 3 h timescales. The prediction accuracy of the developed ARIMA–random forest hybrid model is compared to that of the persistence and bias-corrected ARIMA models. The ARIMA–random forest model is shown to improve upon the latter commonly employed modeling methods, reducing hourly forecasting error by up to 5 % below that of the bias-corrected ARIMA model and achieving an R 2 value of 0.84 with true wind speed.

17 WIND ENERGY↗

Temporal Forecasting of Distributed Temperature Sensing in a Thermal Hydraulic System With Machine Learning and Statistical Models

We benchmark performance of long-short term memory (LSTM) network machine learning model and autoregressive integrated moving average (ARIMA) statistical model in temporal forecasting of distributed temperature sensing (DTS). Data in this study consists of fluid temperature transient measured with two co-located Rayleigh scattering fiber optic sensors (FOS) in a forced convection mixing zone of a thermal tee. We treat each gauge of a FOS as an independent temperature sensor. We first study prediction of DTS time series using Vanilla LSTM and ARIMA models trained on prior history of the same FOS that is used for testing. The results yield maximum absolute percentage error (MaxAPE) and root mean squared percentage error (RMSPE) of 1.58% and 0.06% for ARIMA, and 3.14% and 0.44% for LSTM, respectively. Next, we investigate zero-shot forecasting (ZSF) with LSTM and ARIMA trained on history of the co-located FOS only, which is advantageous when limited training data is available. The ZSF MaxAPE and RMSPE values for ARIMA are comparable to those of the Vanilla use case, while the error values for LSTM increase. We show that in ZSF, performance of LSTM network can be improved by training on most correlated gauges between the two FOS, which are identified by calculating the Pearson correlation coefficient. The improved ZSF MaxAPE and RMSPE for LSTM are 4.4% and 0.33%, respectively. Performance of ZSF LSTM can be further enhanced through transfer learning (TL), where LSTM is re-trained on a subset of the FOS that is the target of forecasting. We show that LSTM pre-trained on correlated dataset and re-trained on 30% of testing target dataset achieves MaxAPE and RMSPE values of 2.32% and 0.28%, respectively.

ARIMA↗

Comparison of Machine Learning-Based Predictive Models of the Nutrient Loads Delivered from the Mississippi/Atchafalaya River Basin to the Gulf of Mexico

Predicting nutrient loads is essential to understanding and managing one of the environmental issues faced by the northern Gulf of Mexico hypoxic zone, which poses a severe threat to the Gulf’s healthy ecosystem and economy. The development of hypoxia in the Gulf of Mexico is strongly associated with the eutrophication process initiated by excessive nutrient loads. Due to the complexities in the excessive nutrient loads to the Gulf of Mexico, it is challenging to understand and predict the underlying temporal variation of nutrient loads. The study was aimed at identifying an optimal predictive machine learning model to capture and predict nonlinear behavior of the nutrient loads delivered from the Mississippi/Atchafalaya River Basin (MARB) to the Gulf of Mexico. For this purpose, monthly nutrient loads (N and P) in tons were collected from US Geological Survey (USGS) monitoring station 07373420 from 1980 to 2020. Machine learning models—including autoregressive integrated moving average (ARIMA), gaussian process regression (GPR), single-layer multilayer perceptron (MLP), and a long short-term memory (LSTM) with the single hidden layer—were developed to predict the monthly nutrient loads, and model performances were evaluated by standard assessment metrics—Root Mean Square Error (RMSE) and Correlation Coefficient (R). The residuals of predictive models were examined by the Durbin–Watson statistic. The results showed that MLP and LSTM persistently achieved better accuracy in predicting monthly TN and TP loads compared to GPR and ARIMA. In addition, GPR models achieved slightly better test RMSE score than ARIMA models while their correlation coefficients are much lower than ARIMA models. Moreover, MLP performed slightly better than LSTM in predicting monthly TP loads while LSTM slightly outperformed for TN loads. Furthermore, it was found that the optimizer and number of inputs didn’t show effects on the LSTM performance while they exhibited impacts on MLP outcomes. This study explores the capability of machine learning models to accurately predict nonlinearly fluctuating nutrient loads delivered to the Gulf of Mexico. Further efforts focus on improving the accuracy of forecasting using hybrid models which combine several machine learning models with superior predictive performance for nutrient fluxes throughout the MARB.

54 ENVIRONMENTAL SCIENCES↗

Use of Physics to Improve Solar Forecast: Part II, Machine Learning and Model Interpretability

Machine learning (ML) models have been applied to forecast solar energy; however, they often lack clarity of interpretability and underlying physics. This work addresses such challenges by developing a hierarchy of ML models that gradually introduce predictors to improve the forecast accuracy based on a physics-based framework. Three ML models (ARIMA, LSTM, and XGBoost) are examined and compared with four physics-informed persistence models reported in Part I and the simple persistence model to assess the improvement of different models. The 7-year measurements at the U.S. Department of Energy's Atmospheric Radiation Measurement's Southern Great Plains Central Facility site are used for forecasts and evaluations. The results reveal that the step-by-step introduction of predictors leads to different improvements for models at different hierarchical levels. Comparison of the ML models with persistence models shows that LSTM and XGBoost outperform all the persistence models, with LSTM having the overall best performance; however, ARIMA underperforms the four physics-informed persistence models. This study demonstrates the importance and utility of incorporating physics into ML models in improving forecast accuracy by introducing a hierarchy of physics-based predictors, distinguishing predictor contributions, and enhancing the ML interpretability. The combined use of Global Horizontal Irradiance (GHI) and Direct Normal Irradiance (DNI) significantly improves the forecast accuracy compared to using individual irradiances alone because the pair contains more information on cloud-radiation interactions.

interpretability↗

Physics-Informed Gaussian Process Regression for States Estimation and Forecasting in Power Grids

Real-time state estimation and forecasting are critical for the efficient operation of power grids. In this paper, a physics-informed Gaussian process regression (PhI-GPR) method is presented and used for forecasting and estimating the phase angle, angular speed, and wind mechanical power of a three-generator power grid system using sparse measurements. In standard data-driven Gaussian process regression (GPR), parameterized models for the prior statistics are fit by maximizing the marginal likelihood of observed data. In the PhI-GPR method, we propose to compute the prior statistics offline by solving stochastic differential equations (SDEs) governing the power grid dynamics. The short-term forecast of a power grid system dominated by wind generation is complicated by the stochastic nature of the wind and the resulting uncertainty in wind mechanical power. Here, we assume that the power grid dynamics are governed by swing equations, with the wind mechanical power fluctuating randomly in time. We solve these equations for the mean and covariances of the power grid states using the Monte Carlo simulation method. We demonstrate that the proposed PhI-GPR method can accurately forecast and estimate observed and unobserved states. For the considered problem, PhI-GPR has computational advantages over the ensemble Kalman filter (EnKF) method: In PhI-GPR, ensembles are computed offline and independently of the data acquisition process, whereas for EnFK, ensembles are computed online with data acquisition, rendering real-time forecast more challenging. We also demonstrate that the PhI-GPR forecast is more accurate than the EnKF forecast when the random mechanical wind power is non-Markovian. In contrast, the two methods produce similar forecasts for the Markovian mechanical wind power. For observed states, we show that PhI-GPR provides a forecast comparable to the standard data-driven GPR; both forecasts are significantly more accurate than the autoregressive integrated moving average (ARIMA) forecast. We also show that the ARIMA forecast is more sensitive to observation frequency and measurement errors than the PhI-GPR forecast.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Cross-Market Price Difference Forecast Using Deep Learning for Electricity Markets

Price forecasting is in the center of decision making in electricity markets. Many researches have been done in forecasting energy prices while little research has been reported on forecasting price difference between day-ahead and realtime markets due to its high volatility, which however plays a critical role in virtual trading. To this end, this paper takes the first attempt to employ novel deep learning architecture with Bidirectional Long-Short Term Memory (LSTM) units to forecast the price difference between day-ahead and real-time markets for the same node. The raw data is collected from PJM market, processed and fed into the proposed network. The Root Mean Squared Error (RMSE) and customized performance metric are used to evaluate the performance of the proposed method. Case studies show that it outperforms the traditional statistical models like ARIMA, and machine learning models like XGBoost and SVR methods in both RMSE and the capability of forecasting the sign of price difference. Additionally, to cross-market price difference forecast, the proposed approach has the potential to be applied to solve other forecasting problems such as price spread forecast in DA market for Financial Transmission Right (FTR) trading purpose.

DA/RT price difference↗

A Time Series Sustainability Assessment of a Partial Energy Portfolio Transition

Energy portfolios are overwhelmingly dependent on fossil fuel resources that perpetuate the consequences associated with climate change. Therefore, it is imperative to transition to more renewable alternatives to limit further harm to the environment. This study presents a univariate time series prediction model that evaluates sustainability outcomes of partial energy transitions. Future electricity generation at the state-level is predicted using exponential smoothing and autoregressive integrated moving average (ARIMA). The best prediction results are then used as an input for a sustainability assessment of a proposed transition by calculating carbon, water, land, and cost footprints. Missouri, USA was selected as a model testbed due to its dependence on coal. Of the time series methods, ARIMA exhibited the best performance and was used to predict annual electricity generation over a 10-year period. The proposed transition consisted of a one-percent annual decrease of coal’s portfolio share to be replaced with an equal share of solar and wind supply. The sustainability outcomes of the transition demonstrate decreases in carbon and water footprints but increases in land and cost footprints. Decision makers can use the results presented here to better inform strategic provisioning of critical resources in the context of proposed energy transitions.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Local Weather Station Design and Development for Cost-Effective Environmental Monitoring and Real-Time Data Sharing

Current weather monitoring systems often remain out of reach for small-scale users and local communities due to their high costs and complexity. This paper addresses this significant issue by introducing a cost-effective, easy-to-use local weather station. Utilizing low-cost sensors, this weather station is a pivotal tool in making environmental monitoring more accessible and user-friendly, particularly for those with limited resources. It offers efficient in-site measurements of various environmental parameters, such as temperature, relative humidity, atmospheric pressure, carbon dioxide concentration, and particulate matter, including PM 1, PM 2.5, and PM 10. The findings demonstrate the station’s capability to monitor these variables remotely and provide forecasts with a high degree of accuracy, displaying an error margin of just 0.67%. Furthermore, the station’s use of the Autoregressive Integrated Moving Average (ARIMA) model enables short-term, reliable forecasts crucial for applications in agriculture, transportation, and air quality monitoring. Furthermore, the weather station’s open-source nature significantly enhances environmental monitoring accessibility for smaller users and encourages broader public data sharing. With this approach, crucial in addressing climate change challenges, the station empowers communities to make informed decisions based on real-time data. In designing and developing this low-cost, efficient monitoring system, this work provides a valuable blueprint for future advancements in environmental technologies, emphasizing sustainability. The proposed automatic weather station not only offers an economical solution for environmental monitoring but also features a user-friendly interface for seamless data communication between the sensor platform and end users. This system ensures the transmission of data through various web-based platforms, catering to users with diverse technical backgrounds. Furthermore, by leveraging historical data through the ARIMA model, the station enhances its utility in providing short-term forecasts and supporting critical decision-making processes across different sectors.

54 ENVIRONMENTAL SCIENCES↗

Technical note: Using long short-term memory models to fill data gaps in hydrological monitoring networks

Abstract. Quantifying the spatiotemporal dynamics in subsurface hydrological flows over a long time window usually employs a network of monitoring wells. However, such observations are often spatially sparse with potential temporal gaps due to poor quality or instrument failure. In this study, we explore the ability of recurrent neural networks to fill gaps in a spatially distributed time-series dataset. We use a well network that monitors the dynamic and heterogeneous hydrologic exchanges between the Columbia River and its adjacent groundwater aquifer at the U.S. Department of Energy's Hanford site. This 10-year-long dataset contains hourly temperature, specific conductance, and groundwater table elevation measurements from 42 wells with gaps of various lengths. We employ a long short-term memory (LSTM) model to capture the temporal variations in the observed system behaviors needed for gap filling. The performance of the LSTM-based gap-filling method was evaluated against a traditional autoregressive integrated moving average (ARIMA) method in terms of error statistics and accuracy in capturing the temporal patterns of river corridor wells with various dynamics signatures. Our study demonstrates that the ARIMA models yield better average error statistics, although they tend to have larger errors during time windows with abrupt changes or high-frequency (daily and subdaily) variations. The LSTM-based models excel in capturing both high-frequency and low-frequency (monthly and seasonal) dynamics. However, the inclusion of high-frequency fluctuations may also lead to overly dynamic predictions in time windows that lack such fluctuations. The LSTM can take advantage of the spatial information from neighboring wells to improve the gap-filling accuracy, especially for long gaps in system states that vary at subdaily scales. While LSTM models require substantial training data and have limited extrapolation power beyond the conditions represented in the training data, they afford great flexibility to account for the spatial correlations, temporal correlations, and nonlinearity in data without a priori assumptions. Thus, LSTMs provide effective alternatives to fill in data gaps in spatially distributed time-series observations characterized by multiple dominant frequencies of variability, which are essential for advancing our understanding of dynamic complex systems.

54 ENVIRONMENTAL SCIENCES↗

Reliable statistics-based detection and investigation of anomalies in a SMART valve system

Reliable anomaly detection and diagnosis are critical for the safe operation of complex engineered systems. This study presents a unified framework that integrates statistical, model-based, and data-driven techniques for anomaly detection and investigation, demonstrated on SMART valve systems in hybrid energy applications. Four detection methods—mean deviation, seasonal extreme studentized deviate, ARIMA forecasting, and matrix profiling—were implemented and compared. Matrix profiling was particularly effective in revealing subtle deviations and hidden relationships among variables. Anomaly investigation was performed by analyzing variable-level and grouped signal profiles, with system topology incorporated to distinguish primary faults from propagated effects. Grouping signals by type enhanced interpretability, enabling accurate localization of anomalies across multi-dimensional datasets. Experimental results confirmed the framework's capability to consistently detect and isolate anomalies while providing actionable insights into system interdependencies. The proposed methodology offers a robust, interpretable, and scalable solution for condition monitoring, with potential applications in safety-critical domains such as nuclear energy, aerospace, and process industries.

ARIMA models↗

Stochastic Price Generation for Evaluating Wholesale Electricity Market Bidding Strategies

This work presents a novel method for generating electricity price scenarios from statistical properties of past electricity prices using a hybrid statistical and reduced-form stochastic model. Previous work in applying stochastic differential equations (SDE) to model electricity prices has focused on daily average prices. To extend stochastic price generation methods to hourly or sub-hourly pricing, we address several weaknesses in the state-of-the-art: (1) we replace the mean-reversion component of the SDE with an ARIMA process that is better able to characterize the daily and weekly trends; (2) we extend the price-spike, or jump process to account for conditional probabilities of price spikes occurring in consecutive time steps by replacing the traditional Poisson process for modeling jumps with a generalized point process model inspired by brain neuron models; and (3) we replace the traditional method of estimating spike intensity with empirical variance with a Markov process based on observed price spike intensity transitions. The method is demonstrated with electricity prices from the US ERCOT market and a use-case example is provided for bidding an energy storage unit into the day-ahead and real-time energy markets of ERCOT using stochastic optimization methods. Results show that the the synthetic price model out performs a (naive) persistence forecast model by resulting in 24% to 47% more in profits over 168 simulated days.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

Online evolutionary neural architecture search for multivariate non-stationary time series forecasting

Time series forecasting (TSF) is one of the most important tasks in data science. TSF models are usually pre-trained with historical data and then applied on future unseen datapoints. However, real-world time series data is usually non-stationary and models trained offline usually face problems from data drift. Models trained and designed in an offline fashion can not quickly adapt to changes quickly or be deployed in real-time. To address these issues, this work presents the Online NeuroEvolution-based Neural Architecture Search (ONE-NAS) algorithm, which is a novel neural architecture search method capable of automatically designing and dynamically training recurrent neural networks (RNNs) for online forecasting tasks. Without any pre-training, ONE-NAS utilizes populations of RNNs that are continuously updated with new network structures and weights in response to new multivariate input data. ONE-NAS is tested on real-world, large-scale multivariate wind turbine data as well as the univariate Dow Jones Industrial Average (DJIA) dataset. These results demonstrate that ONE-NAS outperforms traditional statistical time series forecasting methods, including online linear regression, fixed long short-term memory (LSTM) and gated recurrent unit (GRU) models trained online, as well as state-of-the-art, online ARIMA strategies. Additionally, results show that utilizing multiple populations of RNNs which are periodically repopulated provide significant performance improvements, allowing this online neural network architecture design and training to be successful.

97 MATHEMATICS AND COMPUTING↗

A data-driven operational model for traffic at the Dallas Fort Worth International Airport

Airports are on the front line of significant innovations, allowing the movement of more people and goods faster, cheaper, and with greater convenience. As air travel continues to grow, airports will face challenges in responding to increasing passenger vehicle traffic, which leads to lower operational efficiency, poor air quality, and security concerns. This paper evaluates methods for traffic demand forecasting combined with traffic microsimulation, which will allow airport operations staff to accurately predict traffic and congestion. Using two years of detailed data describing individual vehicle arrivals and departures, aircraft movements, and weather at Dallas-Fort Worth (DFW) International Airport, we evaluate multiple prediction methods including the Auto Regressive Integrated Moving Average (ARIMA) family of models, traditional machine learning models, and DeepAR, a modern recurrent neural network (RNN). We find that these algorithms are able to capture the diurnal trends in the surface traffic, and all do very well when predicting the next 30 minutes of demand. Longer forecast horizons are moderately effective, demonstrating the challenge of this problem and highlighting promising techniques as well as potential areas for improvement. Traffic demand is not the only factor that contributes to terminal congestion, because temporary changes to the road network, such as a lane closure, can make benign traffic demand highly congested. Combining a demand forecast with a traffic microsimulation framework provides a complete picture of traffic and its consequences. The result is an operational intelligence platform for exploring policy changes, as well as infrastructure expansion and disruption scenarios. To demonstrate the value of this approach, we present results from a case study at DFW Airport assessing the impact of a policy change for vehicle routing in high demand scenarios. This framework can assist airports like DFW as they tackle daily operational challenges, as well as explore the integration of emerging technology and expansion of their services into long term plans.

97 MATHEMATICS AND COMPUTING↗

Impact of duration and missing data on the long-term photovoltaic degradation rate estimation

Accurate quantification of photovoltaic (PV) system degradation rate (R D ) is essential for lifetime yield predictions. Although R D is a critical parameter, its estimation lacks a standardized methodology that can be applied on outdoor field data. The purpose of this paper is to investigate the impact of time period duration and missing data on R D by analyzing the performance of different techniques applied to synthetic PV system data at different linear R D patterns and known noise conditions. The analysis includes the application of different techniques to a 10-year synthetic dataset of a crystalline Silicon PV system, with emulated degradation levels and imputed missing data. Here, the analysis demonstrated that the accuracy of ordinary least squares (OLS), year-on-year (YOY), autoregressive integrated moving average (ARIMA) and robust principal component analysis (RPCA) techniques is affected by the evaluation duration with all techniques converging to lower R D deviations over the 10-year evaluation, apart from RPCA at high degradation levels. Moreover, the estimated R D is strongly affected by the amount of missing data. Filtering out the corrupted data yielded more accurate R D results for all techniques. It is proven that the application of a change-point detection stage is necessary and guidelines for accurate R D estimation are provided.

14 SOLAR ENERGY↗

A comprehensive framework to assess elemental mercury in the Department of Energy: A time series analysis

Objective: This study investigated whether seasonal categories affect airborne mercury concentrations in the U.S. Department of Energy operations. Methods: We conducted an initial assessment of the general variability of airborne elemental mercury time-weighted average (TWA) samples. Then, we performed a two-component time series analysis to determine whether long-term, cyclical temperature change patterns affect mercury concentrations. Results: Both ARIMA time series models demonstrated stationary, non-random means (χ² = 83.8, p < 0.001) and standard deviation (χ² = 55.8, p < 0.001) of mercury concentrations. Here, our results indicate that the seasonal factors did not influence mercury concentration. Conclusions: Our results demonstrate that mercury concentrations primarily emanate from operational activities, work practices, and/or transient environmental conditions rather than seasonal fluctuations.

Cannady, Ryan T. [Oak Ridge National Laboratory (O↗

The first two chromosome‐scale genome assemblies of American hazelnut enable comparative genomic analysis of the genus Corylus

Summary The native, perennial shrub American hazelnut ( Corylus americana ) is cultivated in the Midwestern United States for its significant ecological benefits, as well as its high‐value nut crop. Implementation of modern breeding methods and quantitative genetic analyses of C. americana requires high‐quality reference genomes, a resource that is currently lacking. We therefore developed the first chromosome‐scale assemblies for this species using the accessions ‘Rush’ and ‘Winkler’. Genomes were assembled using HiFi PacBio reads and Arima Hi‐C data, and Oxford Nanopore reads and a high‐density genetic map were used to perform error correction. N50 scores are 31.9 Mb and 35.3 Mb, with 90.2% and 97.1% of the total genome assembled into the 11 pseudomolecules, for ‘Rush’ and ‘Winkler’, respectively. Gene prediction was performed using custom RNAseq libraries and protein homology data. ‘Rush’ has a BUSCO score of 99.0 for its assembly and 99.0 for its annotation, while ‘Winkler’ had corresponding scores of 96.9 and 96.5, indicating high‐quality assemblies. These two independent assemblies enable unbiased assessment of structural variation within C. americana , as well as patterns of syntenic relationships across the Corylus genus. Furthermore, we identified high‐density SNP marker sets from genotyping‐by‐sequencing data using 1343 C. americana , C. avellana and C. americana × C. avellana hybrids, in order to assess population structure in natural and breeding populations. Finally, the transcriptomes of these assemblies, as well as several other recently published Corylus genomes, were utilized to perform phylogenetic analysis of sporophytic self‐incompatibility (SSI) in hazelnut, providing evidence of unique molecular pathways governing self‐incompatibility in Corylus .

54 ENVIRONMENTAL SCIENCES↗

The predictive skill of convolutional neural networks models for disease forecasting

In this paper we investigate the utility of one-dimensional convolutional neural network (CNN) models in epidemiological forecasting. Deep learning models, in particular variants of recurrent neural networks (RNNs) have been studied for ILI (Influenza-Like Illness) forecasting, and have achieved a higher forecasting skill compared to conventional models such as ARIMA. In this study, we adapt two neural networks that employ one-dimensional temporal convolutional layers as a primary building block—temporal convolutional networks and simple neural attentive meta-learners—for epidemiological forecasting. We then test them with influenza data from the US collected over 2010-2019. We find that epidemiological forecasting with CNNs is feasible, and their forecasting skill is comparable to, and at times, superior to, plain RNNs. Thus CNNs and RNNs bring the power of nonlinear transformations to purely data-driven epidemiological models, a capability that heretofore has been limited to more elaborate mechanistic/compartmental disease models.

59 BASIC BIOLOGICAL SCIENCES↗