Search NASASearch

SEARCH · Search NASA

Results for “Resource prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Machine learning-driven predictive resource management in complex science workflows

Here, the collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.

97 MATHEMATICS AND COMPUTING

Impact of El Niño‐Southern Oscillation and Madden‐Julian Oscillation on the US Puget Sound Regional Hydroclimate

El Niño-Southern Oscillation (ENSO) and Madden-Julian Oscillation (MJO) are two major modes of climate variability with global hydroclimate impacts. However, their impacts often depend on the local climate and geography, resulting in large regional differences. In this study, we examined the connection of ENSO and MJO to the hydroclimate conditions and extremes in the Puget Sound (PS) basin located in the US Pacific Northwest coast. The results indicate that ENSO significantly modulates the cold season temperature and temperature-mediated hydrologic processes. El Niño cold seasons feature less snow accumulation and intensified surface runoff, even if the precipitation amount is similar to La Niña cold seasons. Therefore, El Niño causes more snow drought (in the form of compound dry and warm snow drought) and shifts the surface runoff seasonality by reducing runoff in the subsequent warm season. MJO phases 6–7 trigger more extreme precipitation, temperature, snowmelt, and runoff in the PS region at 0–9–day lags, and such connections are robust regardless of how the ENSO signals are removed. Meanwhile, MJO modulates large-scale extreme weather systems (e.g., atmospheric rivers) with significant enhancement during phases 6–7. ENSO impacts have intensified in the 2001–2020 period, whereas MJO impacts showed some phase shift in this period. This study reveals ENSO and MJO phases 6–7 as useful predictors of the PS hydroclimate anomalies/extremes at seasonal and daily scales, respectively. Utilizing these findings holds the potential to improve regional water resources prediction and management.

ENSO

Predicting runtime and resource utilization of jobs on integrated cloud and HPC systems

Recent advances in virtualization technologies used in cloud computing offer performance that closely approaches bare-metal levels. Combined with specialized instance types and high-speed networking services for cluster computing, cloud platforms have become a compelling option for high-performance computing (HPC). However, most current batch job schedulers in HPC systems are designed for homogeneous clusters and make decisions based on limited information about jobs and system status. Scientists typically submit computational jobs to these schedulers with a requested runtime that is often over- or under-estimated. More accurate runtime predictions can help schedulers make better decisions and reduce job turnaround times. Here, they can also support decisions about migrating jobs to the cloud to avoid long queue wait times in HPC systems.

97 MATHEMATICS AND COMPUTING

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han

Regime Characterization of Offshore Wind Resource Using Unsupervised Learning

Predictability of wind resource conditions is critical for offshore wind design and operations. While many studies of extreme wind conditions focus on specific events such as low-level jets or ramps, these rely on threshold definitions that limit generality. Here we present a data-driven framework that combines principal component analysis (PCA), self-organizing maps (SOM), and k-means clustering to classify wind resource conditions as typical and anomalous from climatological data. Anomalies are defined not by fixed thresholds but by flagging samples located far from SOM node centers inside the baseline SOM structure. This reframes extremes as rare ebents and hence, likely difficult to anticipate by numerical weather prediction models. We applied this approach to 23 years (2000–2022) of hourly profiles from the NOW-23 hindcast model at the Humboldt Wind Energy Area. Classification is conducted on a feature space consisting of 10 m wind speed and direction, bulk shear and veer across 30–270 m, and a low-level jet index. Dimensionality reduction is achieved through PC. A 2 × 3 OM lattice trained on the PCA vectors identified six baseline regimes spanning weak to strong flow states. High quantization-error profiles are identified and re-clustered into four anomalous regimes. The baseline regimes exhibited clear seasonal and diurnal cycles. Meanwhile, the anomalous regimes represented <10 % of all hours but showed distinct combinations of speed, shear, and veer, when compared to the baseline regimes. Anomalous regimes are typically short-lived (~few hours), yet their transitions can lead to hub-height wind changes of −18 to +9 m s -1 . For a representative 15 MW turbine, these shifts imply rapid swings in capacity factor from near-full output to negligible generation. Validation with lidar buoy data showed 51% agreement in SOM labels across ~6,000 overlapping hours, with most mismatches confined to adjacent speed classes. HRRR comparisons further revealed that anomalous regimes were disproportionately associated with forecast biases exceeding 5 m s -1 . Together, these results reframe extremes in offshore wind from absolute maxima or minima to weather states that are difficult to anticipate from models.

17 WIND ENERGY

Population structure limits the use of genomic data for predicting phenotypes and managing genetic resources in forest trees

There is overwhelming evidence that forest trees are locally adapted to climate. Thus, genecological models based on population phenotypes have been used to measure local adaptation, infer genetic maladaptation to climate, and guide assisted migration. However, instead of phenotypes, there is increasing interest in using genomic data for gene resource management. We used whole-genome resequencing and common-garden experiments to understand the genetic architecture of adaptive traits in black cottonwood. We studied the potential of using genome-wide association studies (GWAS) and genomic prediction to detect causal loci, identify climate-adapted phenotypes, and inform gene resource management. We analyzed population structure by partitioning phenotypic and genomic (single-nucleotide polymorphism) variation among 840 genotypes collected from 91 stands along 16 rivers. Most phenotypic variation (60 to 81%) occurred among populations and was strongly associated with climate. Population phenotypes were predicted well using genomic data (e.g., predictive abilityr> 0.9) but almost as well using climate or geography (r> 0.8). In contrast, genomic prediction within populations was poor (r< 0.2). We identified many GWAS associations among populations, but most appeared to be spurious based on pooled within-population analyses. Hierarchical partitioning of linkage disequilibrium and haplotype sharing suggested that within-population genomic prediction and GWAS were poor because allele frequencies of causal loci and linked markers differed among populations. Given the urgent need to conserve natural populations and ecosystems, our results suggest that climate variables alone can be used to predict population phenotypes, delineate seed zones and deployment zones, and guide assisted migration.

Science & Technology - Other Topics

Evaluating the effect of meso/submesoscale current–wave interactions on wave energy resource characterization at northeast U.S. coast

Wave energy is a promising renewable resource, but accurate assessment is difficult in regions with strong currents due to wave–current interactions (WCI). Here, this study develops a two-way coupled WCI model within the Coupled Ocean Atmosphere Wave Sediment Transport (COAWST) framework at 2 km resolution to improve wave energy characterization along the northeastern U.S. coast, including the Mid-Atlantic Bight and Gulf of Maine. The model integrates WaveWatchIII (WWIII) and the Regional Ocean Modeling System (ROMS) to enhance wave hindcasting by accounting for Doppler-shift, refraction, and nonlinear energy exchanges. Validation against buoy and satellite observations confirms model accuracy. Analysis shows that Doppler-shifting can alter wave power density by over 20%, while strong current gradients and shear distort wave crests via focusing/defocusing and stretching/squeezing, modifying wave direction and frequency. These processes together can induce wave power fluctuations of up to 40% on synoptic scales. Applying a 2.5 MW Ocean Energy Converter power matrix shows that WCI may change harvested energy by up to 100% in shallow-waters and 60% in deep-waters. These results underscore the importance of incorporating nonlinear WCI for reliable wave climate predictions and resource assessments in energetic coastal regions.

doppler-shift

Evaluating mesoscale model predictions of diurnal speedup events in the Altamont Pass Wind Resource Area of California

Mesoscale model predictions of wind, turbulence, and wind energy capacity factors are evaluated in the Altamont Pass Wind Resource Area of California (APWRA), where the diurnal regional sea breeze and associated terrain-driven speedup flows drive wind energy production during the summer months. Results from the Weather Research and Forecasting model version 4.4 using a novel three-dimensional planetary boundary layer (3D PBL) scheme, which treats both vertical and horizontal turbulent mixing, are compared to those using a well-established one-dimensional (1D) scheme that treats only vertical turbulent mixing. Each configuration is evaluated over a nearly 3-month-long period during the Hill Flow Study, and due to the recurring nature of the observed speedup flows, diurnal composite averaging is used to capture robust trends in model performance. Both model configurations showed similar overall skill. The general timing and direction of the speedup flows is captured, but their magnitude is overestimated within a typical wind turbine rotor layer. Both also fail to capture a persistent observed near-surface jet-like flow, likely due to the limited grid resolution that is typical of mesoscale models. However, the 3D PBL configuration shows several minor improvements over the 1D PBL configuration, including improved wind speed and turbulence kinetic energy profiles during the accelerating phase of the speedup events, as well as reduced positive wind speed bias at surface stations across the APWRA region. Using a mesoscale wind farm parameterization, modeled capacity factors are also compared to monthly data reported to the US Energy Information Administration (EIA) during the study period. Although the monthly trend in the data is captured, both model configurations overestimate capacity factors by roughly 7 %–11 %. Through model evaluation, this study provides confidence in the 3D PBL scheme for wind energy applications in complex terrain and provides guidance for future testing.

17 WIND ENERGY

A unifying equation for fermentation sustainability across the titer-rate-yield landscape

Industrial fermentation is central to the sustainable production of fuels and chemicals, yet commercial viability of emerging technologies hinges on improving fermentation titer, rate, and yield (TRY). How these metrics shape system cost remains difficult to generalize due to complex interactions among feedstocks, fermentation, separations, catalytic upgrading, waste management, and facility design. Here, we systematically map theoretical fermentation performance spaces (formed by all potential TRY combinations) for 32 representative biomanufacturing facilities—spanning distinct choices for feedstocks, fermentation regimes and products, separations, and catalytic upgrading—by simulating and evaluating them (via techno-economic analysis, TEA) under uncertainty (600,000 Monte Carlo simulations) and across TRY combinations (7500 TRY combinations for each of 32 configurations). Across this wide design and thermodynamic simulation space, we find the relationship between fermentation TRY and system cost is captured by a simple, generalizable mathematical equation (R 2 of 0.992 − 1.000 across our simulations; 0.954 − 1.000 when validated against prior studies that used different tools). We use this equation to elucidate key drivers that shape cost sensitivity to fermentation performance, generating widely applicable insights. By demonstrating a unifying relationship governs the impact of fermentation on biomanufacturing economics, this work establishes a foundation for agile, holistically predictive, resource-efficient strategies to prioritize fermentation research and development needs and accelerate commercialization of emerging biomanufacturing technologies.

applied mathematics

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin [Fermilab] (ORCID:0000000157000288

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multipleefforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of synthesized ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680 000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin G. [Fermilab]

Assessing the reliability of medical resource demand models in the context of COVID-19

Abstract Background Numerous medical resource demand models have been created as tools for governments or hospitals, aiming to predict the need for crucial resources like ventilators, hospital beds, personal protective equipment (PPE), and diagnostic kits during crises such as the COVID-19 pandemic. However, the reliability of these demand models remains uncertain. Methods Demand models typically consist of two main components: hospital use epidemiological models that predict hospitalizations or daily admissions, and a demand calculator that translates the outputs of the epidemiological model into predictions for resource usage. We conducted separate analyses to evaluate each of these components. In the first analysis, we validated various hospital use epidemiological models using a recent validation framework designed for epidemiological models. This allowed us to quantify the accuracy of the models in predicting critical aspects such as the date and magnitude of local COVID-19 peaks, among other factors. In the second analysis, we evaluated a range of demand calculators for ventilators, medical gowns, and COVID-19 test kits. To achieve this, we decoupled these demand calculators from the underlying epidemiological models and provided ground truth data for their inputs. This approach enabled a direct comparison of the demand calculators, comparing them against each other and actual usage data when available. The code is available athttps://doi.org/10.5281/zenodo.13712387. Results Performance varied greatly across the epidemiological models, with greater variability in COVID-19 hospital use predictions than for COVID-19 deaths as analyzed previously. Some models did not have any peaks. Among those that did, the models under-estimated date of peak approximately as often as they over-estimated, but were more likely to under-estimate magnitude of peak, with typical relative errors around 50%. Regarding demand calculator predictions, there was significant variability, including five-fold differences in predictions for gown models. Validation against actual or surrogate usage data illustrated the potential value of demand models while demonstrating their limitations. Conclusions The emerging field of demand modeling holds promise in averting medical resource shortages during future public health emergencies. However, achieving this potential necessitates focused efforts on standardization, transparency, and rigorous model validation before placing reliance on demand models in critical public health decision-making.

Medical Informatics

Spatially structured bacterial interactions alter algal carbon flow to bacteria

Phytoplankton account for nearly half of global photosynthetic carbon fixation, and the fate of that carbon is regulated in large part by microbial food web processing. We currently lack a mechanistic understanding of how interactions among heterotrophic bacteria impact the fate of photosynthetically fixed carbon. Here, we used a set of bacterial isolates capable of growing on exudates from the diatom Phaeodactylum tricornutum to investigate how bacteria-bacteria interactions affect the balance between exudate remineralization and incorporation into biomass. With exometabolomics and genome-scale metabolic modeling, we estimated the degree of resource competition between bacterial pairs. In a sequential spent media experiment, we found that pairwise interactions were more beneficial than predicted based on resource competition alone, and 30% exhibited facilitative interactions. To link this to carbon fate, we used single-cell isotope tracing in a custom cultivation system to compare the impact of different "primary" bacterial strains in close proximity to live P. tricornutum on a distal "secondary" strain. We found that a primary strain with a high degree of competition decreased secondary strain carbon drawdown by 51% at the single-cell level, providing a quantitative metric for the "cost" of competition on algal carbon fate. Additionally, a primary strain classified as facilitative based on sequential interactions increased total algal-derived carbon assimilation by 7.6 times, integrated over all members, compared to the competitive primary strain. Our findings suggest that the degree of interaction between bacteria along a spectrum from competitive to facilitative is directly linked to algal carbon drawdown.

genome-scale metabolic model

Leveraging public AI tools to explore systems biology resources in mathematical modeling

Predictive mathematical modeling is an essential part of systems biology and is interconnected with information management. Systems biology information is often stored in specialized formats to facilitate data storage and analysis. These formats are not designed for easy human readability and thus require specialized software to visualize and interpret results. Therefore, comprehending modeling and underlying networks and pathways is contingent on mastering systems biology tools, which is particularly challenging for users with no or little background in data science or system biology. To address this challenge, we investigated the usage of public Artificial Intelligence (AI) tools in exploring systems biology resources in mathematical modeling. We tested public AI’s understanding of mathematics in models, related systems biology data, and the complexity of model structures. Our approach can enhance the accessibility of systems biology for non-system biologists and help them understand systems biology without a deep learning curve.

59 BASIC BIOLOGICAL SCIENCES

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy

RAFT: Reconfigurable Array of High-Efficiency Ducted Turbines for Hydrokinetic Energy Harvesting

Diversifying the energy harvesting portfolio is crucial to achieving the ambitious goal of transitioning to clean energy by 2030. Marine hydrokinetic energy has garnered renewed interest due to its high harvesting potential in the U.S., and the resource's reliability and predictability—remaining relatively constant on a daily basis and available 24/7. However, there are currently few commercial devices capable of harnessing the energy from flowing water. This project aims to bridge that gap by designing and evaluating a novel hydrokinetic turbine concept that can efficiently harvest energy from both rivers and tidal streams. The RAFT (Reconfigurable Array of High-Efficiency Ducted Turbines) concept introduces a duct surrounding the turbine rotor and creates an array of small 5-kW units. The duct serves two primary purposes: (1) it enhances hydrodynamic efficiency by accelerating flow to the rotor, and (2) it functions as a structural component, facilitating the formation of modular arrays that lower costs. This project focuses on demonstrating this concept and validating these benefits through simulations and scaled prototype testing. The project team includes 8 faculty members and over 20 students from 3 universities, organized into three core areas: hydrodynamics, electrical systems, and structural analysis, with additional teams dedicated to system integration, environmental assessment and risk management, and tech-to-market strategy. The team successfully demonstrated the increased hydrodynamic efficiency of a ducted turbine compared to an unducted version using high-fidelity simulations and prototype tests. Moreover, design optimization efforts led to surpassing the SHARKS program's goal of 60% reduction of the levelized cost of energy with a significant margin.

13 HYDRO ENERGY

Evaluating the Accuracy of Machine Learning Forecasts

To improve the accuracy of forecasting in machine learning, we must investigate multiple machine learning models and see how accurately they can predict values after training. We used seven machine learning models to try and get more accurate predictions. The models that were used were ARIMA, SES, MLP, CART, LightGBM, and XGBoost. We used a processed dataset from a Terminal at LAX that had the number of people traveling through terminal X every hour in March from 2015-2019. We trained our models with the dates March 6 - March 19 to predict the value for March 20th and the hours 6:00 am to 6:00 pm since those are the most popular traveling hours. By using the different models, we had varying results of accuracy when estimating the amount of people traveling through terminal X on March 20th. We know that machine learning models are helpful for forecasting and by seeing how accurately these models can predict, we can see how forecasting can be helpful for other issues. Using these methods, airports can use forecasting to predict the amount of people coming in and out and can use these predictions to prepare their resource management, operational efficiency, and overall passenger experience.

97 MATHEMATICS AND COMPUTING