Search NASASearch

SEARCH · Search NASA

Results for “data-driven resource management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

35 records · Page 2

Interparticle Characterization of Mechanical Biomass Particle-Particle and Particle-Wall Interactions

The biomass materials industry faces significant challenges in managing material variability and its impact on storage and handling systems. Physical properties such as moisture content, particle size, and density fluctuate considerably, leading to operational issues like bridging and ratholing that disrupt material flow. These variations create a complex cascade effect throughout the process chain, affecting transportation, storage, and conversion processes. The economic consequences of this variability manifest in increased operational costs, maintenance requirements, and system downtime. Environmental factors further complicate the situation, as weather conditions and seasonal availability influence material properties and system performance. Engineers employ specialized equipment design, material characterization protocols, and pre-processing steps like size reduction and homogenization to address these challenges. A critical knowledge gap exists between continuous-level constitutive models and particle-scale behavior. This project developed a novel device to quantify interparticle mechanics between biomass particles, measuring friction and adhesion forces between particles and wall materials. The research focused on corn stover and southern pine forest residue, creating a comprehensive database of particle interactions. This breakthrough enables direct application in particle-based computational modeling, advancing the field's understanding of biomass handling characteristics and supporting the development of more reliable and efficient storage and handling systems. The project's outcomes contribute significantly to understanding biomass's mechanical and flow characteristics, particularly how variability at the particle level affects larger-scale handling operations. This knowledge is crucial for engineering feedstock supply systems that consistently meet quality and cost specifications for various conversion processes. The innovative experimental setup developed through this research represents a significant advancement in biomass characterization methodology. Providing precise measurements of particle-level interactions establishes a foundation for more accurate predictive modeling of bulk material behavior. This enhanced understanding of fundamental particle mechanics enables engineers to anticipate better and address handling challenges before they manifest in full-scale operations. This research opens new avenues for optimizing biomass handling systems through data-driven design approaches. The comprehensive database of particle interactions serves as a valuable resource for future research and development efforts, potentially leading to more efficient and cost-effective biomass processing solutions. This advancement in particle-level mechanics could revolutionize how biomass handling systems are designed and operated, contributing to more sustainable and reliable renewable energy production.

09 BIOMASS FUELS

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence

Artificial Intelligence-Driven Management of Sustainable Energy Resources: Visibility, Operation, and Control

The rapid global transition toward sustainable energy resources (SERs) is reshaping how modern power systems are observed, optimized, and controlled. While SERs have significantly advanced decarbonization, their weather dependence, variability, and inverter-dominated characteristics challenge traditional, centralized, and deterministic grid operation. At the same time, the proliferation of high-resolution data from inverters, smart meters, and sensors offers unprecedented visibility into system dynamics. Yet, it also exceeds the analytical capability of conventional model-based approaches. Artificial intelligence (AI) provides a new foundation for addressing these challenges by bridging physical laws with data-driven learning, enabling accurate state awareness, adaptive operation, and coordinated control across distributed assets. This article examines how AI transforms the management of SER-rich power systems along three critical dimensions: 1) enhancing visibility by inferring behind-the-meter (BTM) activities, assessing SER flexibility, and reconstructing system states from sparse or noisy measurements; 2) improving operation through AI-enhanced SER service provision, volt/var control (VVC), and dynamic operating envelopes (DOE) for efficiency and security; and 3) advancing control by embedding learning-based intelligence into inverter coordination, voltage and frequency regulation, and long-term dispatch. Together, these developments reveal how AI can convert the variability of SERs from an operational challenge into a source of flexibility, resilience, and intelligence, paving the way toward sustainable, adaptive, and self-optimizing power systems.

24 POWER TRANSMISSION AND DISTRIBUTION

Exploring the Frontiers of Energy Efficiency using Power Management at System Scale

In the face of surging power demands for exascale HPC systems, this work tackles the critical challenge of understanding the impact of software-driven power management techniques like Dynamic Voltage and Frequency Scaling (DVFS) and Power Capping. These techniques have been actively developed over the past few decades. By combining insights from GPU benchmarking to understand application power profiles, we present a telemetry data-driven approach for deriving energy savings projections. This approach has been demonstrably applied to the Frontier supercomputer at scale. Our findings based on three months of telemetry data indicate that, for certain resource-constrained jobs, significant energy savings (up to 8.5%) can be achieved without compromising performance. This translates to a substantial cost reduction, equivalent to 1438 MWh of energy saved. The key contribution of this work lies in the methodology for establishing an upper limit for these best-case scenarios and its successful application. This work enables HPC professionals to optimize the power-performance trade-off within constrained power budgets, not only for the exascale era but also beyond.

Karimi, Ahmad Maroof

Data-Driven Mean-Corrected Recursive Estimation-Based Optimal DER Dispatch for Distribution System Voltage Control

Recent advances in smart inverters offer opportunities to mitigate adverse grid impacts caused by high penetrations of distributed photovoltaics (PV) in distribution grids, such as voltage violations. Here, this paper proposes a novel measurement-driven optimal power flow (OPF)-based distributed energy resource management system (DERMS) voltage regulation via recursive sensitivity estimation informed coordinated control of distributed PV inverters. The proposed approach leverages available grid and controllable DER measurements, eliminating reliance on system model information while being adaptive and robust to volatile operating conditions. A mean-corrected recursive ridge regression (MCRRR) algorithm is proposed for sensitivity estimation, continuously refining the sensitivity model through a closed-form solution. It effectively manages varying grid operating conditions, such as changes in power injections and topology reconfiguration, to facilitate a time-varying update of the Load Sensitivity Factors (LSF). The proposed approach is formulated as a linear programming (LP) problem and is thus scalable to larger-scale distribution systems. Its effectiveness and efficiency are demonstrated on a realistic distribution feeder with high PV penetrations in Southern California, USA.

14 SOLAR ENERGY

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING

Distribution Substation Planning Toolkit (dsp-toolkit) v1.0

The Distribution Substation Planning Toolkit (DSP Toolkit) is a software suite designed to streamline the planning and optimization of distribution substations. This toolkit offers a comprehensive set of tools and APIs for data curation, short-term electric load forecasting, and weather-sensitive load adjustment, making it an essential resource for utility companies, engineers, and researchers. Features • Data Preprocessing and Curation: Efficiently manage and preprocess large datasets to ensure high-quality input for analysis. • Short-Term Load Forecasting: Utilize data-driven models to predict short-term electric loads accurately. • Weather-Sensitive Modeling: Automatically adjust load forecasts based on weather data to predict future peak demands more precisely. Uses The DSP Toolkit is ideal for planning and optimizing distribution substations, providing a user-friendly interface and comprehensive documentation. It is suitable for both novice and experienced users, facilitating efficient and accurate planning processes. Advantages • Efficiency: Automates complex planning tasks, reducing manual effort and minimizing errors. • Scalability: Handles large datasets and complex models, making it suitable for large-scale projects. • Community and Support: Open-source with active community contributions, ensuring continuous improvement and support. • Extensibility: Easily extendable with custom modules and plugins, allowing users to tailor the toolkit to their specific needs. The DSP Toolkit stands out by offering a robust, flexible, and user-friendly solution for distribution substation planning. Public Abstract

Li, Han [Lawrence Berkeley National Laboratory (LB

Advancing Urban Water Resilience: Coproducing Knowledge through Civic–Academic Global Partnerships on Water and Climate

As extreme weather events become more pronounced, the vulnerabilities associated with the urban water supply and wastewater systems in megacities are intensified in multiple interconnected dimensions. These multifaceted water challenges can benefit from enhanced cross-sectoral collaboration and sharing of critical knowledge, which are essential for sustainable and adaptive water governance frameworks. In this context, the Megacity Alliance for Water and Climate (MAWAC)–Europe and North America Region (ENAR) Working Group convened a workshop in March 2023, followed by a subsequent workshop in London, United Kingdom, from 11 to 13 September 2024. These workshops aimed to investigate and devise solutions for the cascading hazards with water systems. The solutions examined various aspects focused on climate adaptation and mitigation, stormwater management, and the governance of water and wastewater systems. Additionally, discussions highlighted the importance of community engagement, economic considerations, equity, and effective communication in addressing these pressing challenges. Over the course of 3 days, experts from academia, government agencies, and industry engaged in meaningful discussions on digital modeling for integrated water management, climate-informed urban planning, and public–private–academic partnerships (Fig. 1). Case studies from cities such as New York, Los Angeles, London, Paris, and Chicago highlighted innovative governance strategies for managing water and wastewater systems, promoting water reuse, planning infrastructure, and fostering stakeholder-driven and stakeholder-informed adaptation. The workshop participants emphasized the need for data-driven decision-making, scalable governance models, and knowledge-sharing networks to enhance urban water governance for sustainability and resilience. This workshop report presents the key takeaways from the 3-day convening, providing a roadmap for integrating scientific research, policy frameworks, and emerging technologies to address water challenges faced by megacities.

Hydrologic models

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING

Teamwork for Oversight of Processes and Systems (TOPS). Implementation guide for TOPS version 2.0, 10 August 1992

As the nation redefines priorities to deal with a rapidly changing world order, both government and industry require new approaches for oversight of management systems, particularly for high technology products. Declining defense budgets will lead to significant reductions in government contract management personnel. Concurrently, defense contractors are reducing administrative and overhead staffing to control costs. These combined pressures require bold approaches for the oversight of management systems. In the Spring of 1991, the DPRO and TRW created a Process Action Team (PAT) to jointly prepare a Performance Based Management (PBM) system titled Teamwork for Oversight of Processes and Systems (TOPS). The primary goal is implementation of a performance based management system based on objective data to review critical TRW processes with an emphasis on continuous improvement. The processes are: Finance and Business Systems, Engineering and Manufacturing Systems, Quality Assurance, and Software Systems. The team established a number of goals: delivery of quality products to contractual terms and conditions; ensure that TRW management systems meet government guidance and good business practices; use of objective data to measure critical processes; elimination of wasteful/duplicative reviews and audits; emphasis on teamwork--all efforts must be perceived to add value by both sides and decisions are made by consensus; and synergy and the creation of a strong working trust between TRW and the DPRO. TOPS permits the adjustment of oversight resources when conditions change or when TRW systems performance indicate either an increase or decrease in surveillance is appropriate. Monthly Contractor Performance Assessments (CPA) are derived from a summary of supporting system level and process-level ratings obtained from objective process-level data. Tiered, objective, data-driven metrics are highly successful in achieving a cooperative and effective method of measuring performance. The teamwork-based culture developed by TOPS proved an unequaled success in removing adversarial relationships and creating an atmosphere of continuous improvement in quality processes at TRW. The new working relationship does not decrease the responsibility or authority of the DPRO to ensure contract compliance and it permits both parties to work more effectively to improve total quality and reduce cost. By emphasizing teamwork in developing a stronger approach to efficient management of the defense industrial base TOPS is a singular success.

Strand, Albert A.

Cognitive Communications for NASA Space Systems

The growing complexity of spacecraft constellations, communication relay offerings, and mission architectures drives the need for the development of autonomous communication systems. NASA has traditionally launched single spacecraft missions that are served by the Space Communication and Navigation (SCaN) program. Operations on SCaN networks are typically scheduled weeks in advance, and often each asset serves a single user spacecraft at a time. Recent movement towards swarm missions could make the current approach unsustainable. Additionally, the integration of commercial communication service providers will substantially increase the data transfer options available to new missions. NASA science missions have found benefit in launching swarms of spacecraft, allowing coordinated simultaneous observations from different perspectives. Inter-spacecraft communication (mesh networking) is an enabler for this architecture, as are CubeSats that allow cost-effective provisioning of distributed mission assets. As more complex swarm missions launch, one challenge is coordinating communication within the swarm and choosing the appropriate mechanism for telemetry, tracking, control, and data services to and from Earth. Cognitive communications research conducted by SCaN aims to mitigate the increasing communication complexity for mission users by increasing the autonomy of links, networks, and service scheduling. By considering automation techniques including recent advances in artificial intelligence and machine learning, cognitive algorithms and related approaches enable increased mission science return, improved resource utilization for service provider networks, and resiliency in unpredictable or unplanned environments. The Cognitive Communications Project at the NASA Glenn Research Center develops applications of data-driven, non-deterministic methods to improve the autonomy of space communication. The project emphasizes development of decentralized space networks with artificial intelligence agents optimizing communication link throughput, data routing, and system-wide asset management. This paper discusses the objectives, approaches, and opportunities of the research to address growing needs of the space communications community.

Chelmins, David

Towards a program of record of inland water quality: Exploiting present and heritage multispectral sensors for maximum information extraction

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This research exploits recent advancements in bio-optical modeling, cloud computing, and machine learning to enhance our capacity to leverage present and heritage satellite data. Recent research suggests that sensors with low spectral resolution, such as Sentinel 2 and Landsat missions, contain enough hidden spectral variation which can be exploited using data-driven approaches. The availability of three decades of archival imagery will open doors to discover global trends of eutrophication and increased cyanobacteria dominance and provide valuable insight to the development of predictive methodologies. Preliminary efforts in synthetic emulation of global natural inland waters will be discussed and contextualized against satellite radiometric measurement uncertainty, satellite data product uncertainties and causal signal ambiguity over the visible wavelength range, supported by high quality field and image data for selected inland aquatic sites. Insights on water quality estimation via data-driven machine learning models versus matrix inversions will be discussed, and how we can exploit spectral-spatial relationships in high spatial resolution data. A cross-sensor synergistic approach with detailed uncertainty analysis based on optical water types, will allow for unprecedented global snapshots of fine scale ecological dynamics of inland waters.

Inland

A Robust Machine Learning Schema for Developing, Maintaining, and Disseminating Machine Learning Models

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of ML models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based modeling of material behavior at various length scales and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using ML techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus, effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train ML models and the defining model parameters and architectures within the Granta MI Platform. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in the prediction of material behavior, while following outlined best practices for effective data management. An effective schema for ML data and models can help prevent the recreation of virtual/real training data and surrogate models, help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Brandon L. Hearley

Predicting Fiber Failure of Plain Weave Fabric with Recursive Multiscale Micromechanics

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of machine learning models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based, modeling of material behavior at various length scales, and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using machine learning (ML) techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train machine learning models and the defining model parameters and architectures. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in for various types of machine learning models while following outlined best practices for effective data management. An effective schema for machine learning data and models can help prevent the recreation of virtual/real training data and surrogate models, can help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Failure

EVs@Scale Next-Gen Profiles - Fleet Utilization 2024

As part of the U.S. Department of Energy’s EVs@Scale initiative, the Next-Gen Profiles (NGP) project provides a comprehensive, data-driven analysis of electric vehicle (EV) and electric vehicle supply equipment (EVSE) operations across real-world fleet deployments. This paper presents findings from the NGP’s Fleet Utilization study, which investigates operational behavior and asset usage across seventeen EV fleets and two EVSE fleets, encompassing a wide range of vehicle types and use cases. Data collected from diverse sources—varying in format and temporal resolution—are first reformatted into a unified structure. From this harmonized dataset, a suite of rigorously defined performance metrics is calculated at an hourly cadence, enabling consistent cross-comparison of charging, routing, and other key operational behaviors. Amid rapidly increasing EV adoption and growing demands for energy-efficient fleet operations, the analysis reveals clear utilization trends—including diurnal and weekly activity cycles, differences in short versus long charging session dependencies, and route-specific energy usage patterns. These findings highlight the need for tailored infrastructure strategies and the deployment of advanced energy management systems, such as Distributed Energy Resource Management Systems (DERMS) and Site Energy Management Systems (SEMS), which can optimize charging schedules and mitigate peak loads. By leveraging anonymized, harmonized datasets and standardized metrics, this study offers critical insights into fleet behavior and performance, providing a foundation to improve operational efficiency, reduce costs, and enable the scalable deployment of electrified transportation.

Wells, Landon

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Ultra-filtration(UF) units are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square (RMSE) metric. Accurate prediction of initial TMP is critical for optimizing CCRO operations, as it enables the development of robust modelling frameworks that enhance process efficiency and reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)