Search NASA⌕ Search

SEARCH · Search NASA

Results for “big data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗

Big Data For Operation and Maintenance Cost Reduction

The purpose of this research is to develop a first-of-a-kind framework for integrating Big Data capability into the daily activities of our current fleet of nuclear power plants. Big Data is traditionally defined as data sets with high volume, velocity, and heterogeneity, and the existing Big Data analytics capabilities are now widely popular in fields such as finance, weather, e-commerce, healthcare and sports. In the nuclear industry, while the volume and velocity of data may present computational challenges for existing analytics capabilities, data heterogeneity are seen to present the major challenge. This research project mainly focuses on incorporating the wide range of data heterogeneities in nuclear power plants into an integrated Big Data Analytics capability. The primary end-product of this project is a Big Data framework that is capable of dealing with the large volume and heterogeneity of the data found in nuclear power plants to extract timely and valuable information on equipment performance. The framework can generate system insights that are actionable relations between measurable impacts and the corresponding maintenance action plans and enable optimization of plant operation and maintenance based on the extracted information. The developed framework is capable of handling heterogeneous data including both image data and time-series sensor data. Specifically, this developed framework includes the following components. The first component is an overarching maintenance ontology which includes system insights required by maintenance optimization. The maintenance ontology interacts with other components in the developed framework. The second component handles Piping & Instrumentation Diagram (P&ID) data. It can be used to extract system components and their relations automatically from the P&IDs. This extracted information is stored in the first component, i.e., maintenance ontology, and is also used as input to the third component, i.e., a tool for generating the fault tree for the corresponding system. The generated fault tree in turn is stored in the ontology for assessing risk that is used as a criterion in maintenance policy optimization. The fourth component is a tool for inferring the parameters in the Markov degradation model for a nuclear system. It uses basic information from the ontology. The fifth component is a tool for assessing the degradation level using sensor measurement data, for example, pressure, flowrate. This tool can be used for determining corrective maintenance actions. The results obtained from components four and five are returned to the ontology. The sixth component of the framework is a tool for optimizing the maintenance policy for a nuclear system of interest. It takes certain basic information from the ontology, e.g., costs of maintenance actions and system failures, as input, and returns the optimal maintenance policy to the ontology. This tool can be used for determining predictive maintenance actions. A set of experiments have also been conducted to verify the algorithms developed in this project for nuclear system degradation monitoring. The experiments are based on four solenoid valves, similar to the ones used in nuclear power plants. The analyses based on the experimental data using two algorithms, i.e., the Randomized Window Decomposition (RWD) algorithm and the particle filtering algorithm, and the results are introduced in the report. The Big Data framework developed in this project can be used as a support tool in daily activities of plant operation and maintenance and will reduce current costs while maintaining or improving safety levels. Overall, the project will not only benefit existing reactors, however it will open new frontiers to realize the long overdue value of Big Data Analytics in the nuclear sphere.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Understanding collective human movement dynamics during large-scale events using big geosocial data analytics

Conventional approaches for modeling human mobility pattern often focus on human activity and movement dynamics in their regular daily lives and cannot capture changes in human movement dynamics in response to large-scale events. With the rapid advancement of information and communication technologies, many researchers have adopted alternative data sources (e.g., cell phone records, GPS trajectory data) from private data vendors to study human movement dynamics in response to large-scale natural or societal events. Big geosocial data such as georeferenced tweets are publicly available and dynamically evolving as real-world events are happening, making it more likely to capture the real-time sentiments and responses of populations. However, precisely-geolocated geosocial data is scarce and biased toward urban population centers. In this research, we developed a big geosocial data analytical framework for extracting human movement dynamics in response to large-scale events from publicly available georeferenced tweets. The framework includes a two-stage data collection module that collects data in a more targeted fashion in order to mitigate the data scarcity issue of georeferenced tweets; in addition, a variable bandwidth kernel density estimation(VB-KDE) approach was adopted to fuse georeference information at different spatial scales, further augmenting the signals of human movement dynamics contained in georeferenced tweets. To correct for the sampling bias of georeferenced tweets, we adjusted the number of tweets for different spatial units (e.g., county, state) by population. To demonstrate the performance of the proposed analytic framework, we chose an astronomical event that occurred nationwide across the United States, i.e., the 2017 Great American Eclipse, as an example event and studied the human movement dynamics in response to this event. Finally, this analytic framework can easily be applied to other types of large-scale events such as hurricanes or earthquakes.

54 ENVIRONMENTAL SCIENCES↗

Development of a Unified Taxonomy for HVAC System Faults

Detecting and diagnosing HVAC faults is critical for maintaining building operation performance, reducing energy waste, and ensuring indoor comfort. An increasing deployment of commercial fault detection and diagnostics (FDD) software tools in commercial buildings in the past decade has significantly increased buildings’ operational reliability and reduced energy consumption. A massive amount of data has been generated by the FDD software tools. However, efficiently utilizing FDD data for ‘big data’ analytics, algorithm improvement, and other data-driven applications is challenging because the format and naming conventions of those data are very customized, unstructured, and hard to interpret. This paper presents the development of a unified taxonomy for HVAC faults. A taxonomy is an orderly classification of HVAC faults according to their characteristics and causal relations. The taxonomy includes fault categorization, physical hierarchy, fault library, relation model, and naming/tagging scheme. The taxonomy employs both a physical hierarchy of HVAC equipment and a cause-effect relationship model to reveal the root causes of faults in HVAC systems. A structured and standardized vocabulary library is developed to increase data representability and interpretability. The developed fault taxonomy can be used for HVAC system ‘big data’ analytics such as HVAC system fault prevalence analysis or the development of an HVAC FDD software standard. A common type of HVAC equipment-packaged rooftop unit (RTU) is used as an example to demonstrate the application of the developed fault taxonomy. Two RTU FDD software tools are used to show that after mapping FDD data according to the taxonomy, the meta-analysis of the multiple FDD reports is possible and efficient.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine

03 NATURAL GAS↗

Findings on Subtask 3.1 - Bakken Rich Gas Enhanced Oil Recovery Project

Total in-place oil for the Bakken petroleum system (BPS) (which includes the Bakken and Three Forks Formations) has been estimated to be 600 billion barrels (bbl). However, BPS wells have decline rates as high as 85% over the first 3 years of their lives, and primary recovery factors typically range from 3% to 10% of original oil in place. Given the low initial recovery rates, even small incremental productivity improvements could dramatically increase technically recoverable oil in the BPS. One potential solution is enhanced oil recovery (EOR) using gas injection, such as carbon dioxide (CO2) or hydrocarbon (HC) gases. While commonly used in conventional reservoirs, CO2 EOR in unconventional tight oil reservoirs has been limited to pilot tests. EOR using rich gas (mixture of methane, ethane, and propane) has also been employed in numerous pilots in several unconventional plays and has recently been successfully applied in the Eagle Ford play. If successful, large-scale gas-based EOR in the BPS could dramatically increase oil productivity and recovery factors and extend the life of the play for decades. While CO2 may be a technically suitable working fluid for EOR in the BPS, supplies are limited and costs for using CO2 in EOR pilots are prohibitively high. Meanwhile, produced gas flaring has presented challenges for BPS operators in North Dakota. Analysis conducted by the North Dakota Pipeline Authority indicates that the current gas-gathering infrastructure in North Dakota is insufficient to accommodate all of the associated gas that is produced from the BPS. The geographically isolated location of North Dakota relative to large natural gas markets, combined with suppressed natural gas prices, has made it economically challenging for industry to invest capital in expanding gas-gathering infrastructure in the state. These circumstances led to a research program conducted by the Energy & Environmental Research Center (EERC) in partnership with Liberty Resources Management Company LLC (LR) to examine the potential to use rich gas injection for EOR and mitigate flaring. A rich gas EOR pilot test was designed and executed by LR at its Stomping Horse development area in Williams County, North Dakota. From July 2018 through May 2019, a total of 160 million standard cubic feet (MMscf) of rich produced gas was injected into the BPS using five different wells in a sequential injection strategy. LR’s Leon–Gohrick drill spacing unit (DSU) was used as the test site. Regulatory oversight was provided by the North Dakota Industrial Commission (NDIC). Technical support was provided by the EERC through a series of laboratory, modeling, and field-based activities, and additional post-pilot research activities incorporated learnings from the test, developed new laboratory data, improved fracture modeling methods, and developed machine learning and big data analytics. The results from the Stomping Horse rich gas EOR pilot activities indicate that developing an effective, economical EOR approach for the BPS will require more field tests. Another key lesson learned from the Stomping Horse tests is that detailed pre- and posttest data on reservoir conditions and fluids production are essential. Robust reservoir characterization provides information that is crucial to creating realistic geomodels and conducting valid dynamic simulations of potential EOR scenarios. A detailed understanding of the completions and production history of offset wells is also necessary for valid test result interpretations. This knowledge is essential to designing the operational parameters of injectivity tests and interpreting the results. A conformance control strategy is also essential to success. Laboratory-based examinations of rich gas interactions with reservoir fluids and rocks were conducted, with an emphasis on determining the ability to mobilize oil in the tight reservoir rocks and shales of the BPS. Injection fluid composition was shown to have a positive impact on reducing reservoir oil minimum miscibility pressure (MMP), reducing interfacial tension (IFT), and altering wettability. IFT and contact angle measurements demonstrated that wettability can be altered in the presence of rich gas, suggesting the potential to improve oil recovery. Iterative modeling of surface infrastructure and reservoir performance using data generated by the various project activities was conducted. A geologic model of the Stomping Horse area was built; history-matched oil, gas, and water production was used in simulations of various EOR scenarios. Early programmatic modeling results were used to support LR’s design and operation of the EOR pilot and to provide insight regarding optimization of future commercial-scale BPS EOR design and operations. Post-pilot modeling focused on alternative methods of understanding complex fracture networks and accelerating simulation time. These led to improved simulation run times and provide excellent history-matching results. Several of these iterative models were used as the bases for developing algorithms into machine learning and big data analytics. History matching in reservoir simulation is time-consuming and computer processing-intensive. Machine learning algorithms were created, and an automated history-matching tool was developed. A large set of synthetic reservoir simulations were created to generate well responses (oil, gas, and water production, well bottomhole pressure [BHP], and tracer or propane breakthrough) for a set of EOR operating parameters that included offset well status (open or closed), injectate (rich gas or propane), injection rate, and injection well BHP. A user interface was developed to provide real-time visualization. Machine learning-based models were developed to provide rapid forecasting of well performance given a set of user-defined EOR operating parameters. These predictive models allow the user to modify the offset well status, injection rate, and injection well BHP and rapidly forecast future production performance. The combination of real-time visualization tools with real-time forecasting tools provides a framework for real-time control—operational changes that the EOR site operator can enact (e.g., changing gas injection rates) to affect the observed performance and potentially improve the EOR outcome. There is great reason to be optimistic about the future of EOR in the Bakken. The results of the laboratory studies suggest significant potential for high rates of oil mobilization using produced field gas injection under the right conditions. The results of the lab studies, combined with rigorous statistical analysis of well production data and associated modeling efforts, confirm the notion that fluid mobility within the reservoir is controlled by fractures. As more knowledge is gained about the nature and distribution of fracture networks in the Bakken, the industry will be in a better position to predict and, ultimately, influence fluid mobility. New field tests are necessary to develop a more complete understanding of those conditions. Thoughtful and creatively engineered field tests within a well-characterized geologic setting will yield the fundamental knowledge needed to take Bakken oil production to the next level. This subtask was cofunded through the EERC–U.S. Department of Energy Joint Program on Research and Development for Fossil Energy-Related Resources Cooperative Agreement No. DE-FE0024233. Nonfederal funding was provided by the North Dakota Industrial Commission’s Oil and Gas Research Program and Computer Modelling Group.

04 OIL SHALES AND TAR SANDS↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data (Final Report)

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components: (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine.

20 FOSSIL-FUELED POWER PLANTS↗

Elucidating and predicting the dynamic evolution of water and land systems due to natural and energy-related forcings

Focal Area(s): 3. Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI; & 1. Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: Interactions between water, land, and energy systems are complex and occur on a variety of scales, ranging from local to basinal to regional. Accurately predicting the behavior of ground water and surface water systems for 5-10 years and beyond requires an understanding of the current system and the ability to model both the natural system at scale and human-induced forcings related to energy and other activities. Artificial intelligence and machine learning (AI/ML) combined with modern compilation and integration efforts for U.S. groundwater and surface water systems present potential solutions to bolstering detailed physics-based models of these systems. Big data tied with ML and physics-based modeling can drive breakthroughs in understanding the earth system, but research is often impeded by data access (e.g., privacy issues), quality, formats, gaps, multi-source, multi-scale, integration, and spatiotemporal challenges. Effective integration of real data and simulated (synthetic) data that fill gaps is critical. Overcoming these complex data and model integration challenges will enable a transformational approach to acquiring enhanced understanding of environmental systems.

54 ENVIRONMENTAL SCIENCES↗

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

Dimensionality reduction using elastic measures

With the recent surge in big data analytics for hyperdimensional data, there is a renewed interest in dimensionality reduction techniques. In order for these methods to improve performance gains and understanding of the underlying data, a proper metric needs to be identified. This step is often overlooked, and metrics are typically chosen without consideration of the underlying geometry of the data. Here, in this paper, we present a method for incorporating elastic metrics into the t-distributed stochastic neighbour embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP). We apply our method to functional data, which is uniquely characterized by rotations, parameterization and scale. If these properties are ignored, they can lead to incorrect analysis and poor classification performance. Through our method, we demonstrate improved performance on shape identification tasks for three benchmark data sets (MPEG-7, Car data set and Plane data set of Thankoor), where we achieve 0.77, 0.95 and 1.00 F1 score, respectively.

97 MATHEMATICS AND COMPUTING↗

Aggregate attack surface management for network discovery of operational technology

Interconnectivity has become a substratum of technology as the benefits of data-driven functionality are being realized in nearly all industries. Increased connectivity of Operational Technology (OT) exacerbates cyber risks because Industrial Control Systems (ICS) are becoming exposed to the Internet. These exposures are often done inadvertently through misconfigurations as additional network devices come online. Attack surface management (ASM) platforms can be used to identify vulnerabilities by performing external network discovery over the Internet using web spiders. These web spiders enable big data analytics of Internet of Things (IoT) devices as identifiable information of Internet-exposed equipment are archived in searchable databases that are made publicly available. There are a multitude of ASM service providers on the market. Here, this study was conducted to evaluate several commonly known tools to determine the aggregate attack surface of control systems. Queries were crafted by targeting commonly known manufacturers and communication protocols found in OT networks. Identified devices were that categorized based on technology types. Each query was replicated between several tools to target identical ICS equipment. Findings in this paper suggested a significant variance in the exposures discovered by each tool, but unique contributions were identified for each tool when a merged attack surface was derived. Therefore, all tools should be used in aggregate.

97 MATHEMATICS AND COMPUTING↗

Machine learning models inaccurately predict current and future high-latitude C balances

The high-latitude carbon (C) cycle is a key feedback to the global climate system, yet because of system complexity and data limitations, there is currently disagreement over whether the region is a source or sink of C. Recent advances in big data analytics and computing power have popularized the use of machine learning (ML) algorithms to upscale site measurements of ecosystem processes, and in some cases forecast the response of these processes to climate change. Due to data limitations, however, ML model predictions of these processes are almost never validated with independent datasets. To better understand and characterize the limitations of these methods, we develop an approach to independently evaluate ML upscaling and forecasting. We mimic data-driven upscaling and forecasting efforts by applying ML algorithms to different subsets of regional process-model simulation gridcells, and then test ML performance using the remaining gridcells. In this study, we simulate C fluxes and environmental data across Alaska using ecosys, a process-rich terrestrial ecosystem model, and then apply boosted regression tree ML algorithms to training data configurations that mirror and expand upon existing AmeriFLUX eddy-covariance data availability. We first show that a ML model trained using ecosys outputs from currently-available Alaska AmeriFLUX sites incorrectly predicts that Alaska is presently a modeled net C source. Increased spatial coverage of the training dataset improves ML predictions, halving the bias when 240 modeled sites are used instead of 15. However, even this more accurate ML model incorrectly predicts Alaska C fluxes under 21st century climate change because of changes in atmospheric CO 2 , litter inputs, and vegetation composition that have impacts on C fluxes which cannot be inferred from the training data. Our results provide key insights to future C flux upscaling efforts and expose the potential for inaccurate ML upscaling and forecasting of high-latitude C cycle dynamics.

54 ENVIRONMENTAL SCIENCES↗

Parallel I/O Evaluation Techniques and Emerging HPC Workloads: A Perspective

Emerging workloads such as artificial intelligence, big data analytics and complex multi-step workflows alongside future exascale applications are anticipated future HPC workloads, which will result in a more diverse I/O system workload and even less predictable I/O behavior and access patterns. Along with the ever increasing gap between the compute and storage performance capabilities, the in-depth understanding of extreme-scale I/O behavior and the I/O performance modeling and prediction are essential tools of the large-scale I/O evaluation process for addressing the needs of extreme-scale hybrid workloads. In this survey article, we focus on the state-of-the-art of the I/O behavior and performance analysis process for HPC systems in a 5-year time window and identify future research challenges.

Neuwirth, Sarah↗

Strategies for Integrating Deep Learning Surrogate Models with HPC Simulation Applications

The emerging trend of the convergence of high performance computing (HPC), machine learning/deep learning (ML/DL), and big data analytics presents a host of challenges for large-scale computing campaigns that seek best practices to interleave traditional scientific simulation-based workloads with ML/DL models. A portfolio of systematic approaches to incorporate deep learning into modeling and simulation serves a vital need when we support AI for science at a computing facility. In this paper, we evaluate several strategies for deploying deep learning surrogate models in a representative physics application on supercomputers at the Oak Ridge Leadership Computing Facility (OLCF). We discuss a set of recommended deployment architectures and implementation approaches. We analyze and evaluate these alternatives and show their performance and scalability up to 1000 GPUs on two mainstream platforms equipped with different deep learning hardware and software stacks.

Yin, Junqi↗

Enabling HPC Scientific Workflows for Serverless

The convergence of edge computing, big data analytics, and AI with traditional scientific calculations is increasingly being adopted in HPC workflows. Workflow management systems are crucial for managing and orchestrating these complex computational tasks. However, it is difficult to identify patterns within the growing population of HPC workflows. Serverless has emerged as a novel computing paradigm, offering dynamic resource allocation, quick response time, fine-grained resource management and auto-scaling. In this paper, we propose a framework to enable HPC scientific workflows on serverless. Our approach integrates a widely used traditional HPC workflow generator with an HPC serverless workflow management system to create benchmark suites of scientific workflows with diverse characteristics. These workflows can be executed on different serverless platforms. We comprehensively compare executing workflows on traditional local containers and serverless computing platforms. Our results show that serverless can reduce CPU and memory usage respectively by 78.11% and 73.92% without compromising performance.

Andrei da silva, Anderson↗

Dimensionally reduced machine learning model for predicting single component octanol–water partition coefficients

Abstract MF-LOGP, a new method for determining a single component octanol–water partition coefficients ( $$LogP$$ LogP ) is presented which uses molecular formula as the only input. Octanol–water partition coefficients are useful in many applications, ranging from environmental fate and drug delivery. Currently, partition coefficients are either experimentally measured or predicted as a function of structural fragments, topological descriptors, or thermodynamic properties known or calculated from precise molecular structures. The MF-LOGP method presented here differs from classical methods as it does not require any structural information and uses molecular formula as the sole model input. MF-LOGP is therefore useful for situations in which the structure is unknown or where the use of a low dimensional, easily automatable, and computationally inexpensive calculations is required. MF-LOGP is a random forest algorithm that is trained and tested on 15,377 data points, using 10 features derived from the molecular formula to make $$LogP$$ LogP predictions. Using an independent validation set of 2713 data points, MF-LOGP was found to have an average $$RMSE$$ RMSE = 0.77 ± 0.007, $$MAE$$ MAE = 0.52 ± 0.003, and $${R}^{2}$$ R 2 = 0.83 ± 0.003. This performance fell within the spectrum of performances reported in the published literature for conventional higher dimensional models ( $$RMSE$$ RMSE = 0.42–1.54, $$MAE$$ MAE = 0.09–1.07, and $${R}^{2}$$ R 2 = 0.32–0.95). Compared with existing models, MF-LOGP requires a maximum of ten features and no structural information, thereby providing a practical and yet predictive tool. The development of MF-LOGP provides the groundwork for development of more physical prediction models leveraging big data analytical methods or complex multicomponent mixtures. Graphical Abstract

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗