Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Bundling measures for food systems transformation: a global, multimodel assessment

Background Current food systems leave one in ten individuals at risk of hunger while driving unsustainable environmental impacts. Inaction risks further exacerbating negative impacts on both human and planetary health. These challenges emerge from complex system interactions, requiring approaches that engage with this complexity and consider how transformation measures interact across food systems. We aimed to quantify the magnitude and uncertainty of the impacts of key food systems transformation measures both individually and in a bundle using an ensemble of global economic models. Methods In this global multimodel assessment, we applied an ensemble of ten state-of-the-art global economic models to evaluate the potential of four key measures in transforming food systems: increasing agricultural productivity, halving food loss and waste, shifting towards healthier diets, and economy-wide climate mitigation policies aligned with limiting warming to 1·5°C. The scenarios used a middle-of-the-road shared socioeconomic pathway for population and gross domestic product growth, climate impact data from Jägermeyr and colleagues, Thornton and colleagues, and Nelson and colleagues, and dietary targets based on the EAT–Lancet healthy reference diet, with model simulations conducted from 2020 to 2050. We then assessed the effect of these measures in isolation and in combination in a bundled scenario. To further understand the interactions between these measures, we conducted a decomposition analysis that distinguishes between the individual effects of a measure (effect when implemented alone), total effects (its contribution within the bundle), and interaction effects (the difference between total and individual effects). This approach aimed to show complementarities and trade-offs that emerge when multiple measures are implemented simultaneously. Findings Our analysis showed that individual measures in isolation are insufficient to achieve high-level environmental objectives and might generate unintended consequences. In contrast, bundling measures produces co-benefits: avoiding 50% of projected agricultural greenhouse gas emissions by 2050 and almost 20% of anticipated land conversion, while moderating food price increases associated with ambitious climate change mitigation policies. Our decomposition analysis further shows that measures can have varying effects across different dimensions. Although dietary shifts and climate mitigation policies are the largest drivers of environmental benefits (each contributing to a median decline of >10 percentage points in non-CO 2 emissions and 5 percentage points in agricultural land use globally), productivity improvements and reducing food loss and waste play essential roles in moderating price increases (each contributing to a median decline of >5 percentage points in average prices). Interpretation This study highlights the importance of implementing coordinated approaches to food system transformation and climate change mitigation rather than relying on isolated interventions. Comprehensive transformation requires understanding how supply-side and demand-side changes can interact with climate mitigation policies, enabling policy makers to design intervention packages that maximise benefits while minimising trade-offs across environmental, economic, and social dimensions.

Sundiang, Marina [Cornell Univ., Ithaca, NY (Unite↗

High-Mileage Courier Fleet Vehicle Laboratory Battery Pack Testing

For one of each of the AVTA's plug-in hybrid electric vehicles, battery electric vehicles, and some hybrid electric vehicles tested in high-mileage courier fleets, the high-voltage traction battery packs were removed from the vehicle and tested at the beginning and end of fleet testing. For some vehicles, batteries also were tested at periodic intervals during fleet testing. Standard reference performance tests were conducted to characterize battery degradation over time. This dataset contains results from two or more rounds of battery tests for 21 distinct year/make/model vehicles (see reference ["INL Advanced Vehicle Testing Activity: On-road Logger and Laboratory Battery Pack Testing Vehicle List"](https://avt.inl.gov/sites/default/files/pdf/reports/DatasetVehicleList.pdf) for full list of vehicles). Each round of battery testing included the "Static Capacity Test" and the "Hybrid Pulse Power Characterization (HPPC) Test", conducted according to test procedures published in the United States Advanced Battery Consortium ["Battery Test Manual For Power-Assist Hybrid Electric Vehicles"](https://www.uscar.org/commands/files_download.php?files_id=57), ["Battery Test Manual For Plug-In Hybrid Electric Vehicles"](https://www.uscar.org/commands/files_download.php?files_id=168), and ["Electric Vehicle Battery Test Procedures Manual"](https://www.uscar.org/commands/files_download.php?files_id=5) prior to the time of testing. These tests were performed by Intertek Testing Services, North America. This dataset is shared by API; a small sample of the vehicle battery and test data has been extracted and is also available for download.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Designing resilient IoT and Edge Computing with federated tinyML

The rapid growth of the Internet of Things (IoT) and Edge Computing (EC) has brought significant conveniences to modern society but has also greatly expanded the cyber attack surfaces, particularly as these technologies are being increasingly integrated into critical systems such as power grids, healthcare, and smart homes. Here, to improve IoT/EC’s cybersecurity posture, we leveraged Artificial Intelligence (AI) and Machine Learning (ML) by employing tinyML to monitor voluminous IoT data for cyber threats while addressing devices’ resource constraints, and utilizing Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. Building on our three-layer architecture combining tinyML and FL to enhance autonomous cyber attack detection, this paper demonstrated that the architecture improves detection accuracy, reduces resource consumption, and enables lightweight, secure IoT device monitoring. These results were validated using the public N-BaIoT dataset as well as real IoT network traffic data collected under multiple attack scenarios from our testbeds. Additionally, we introduced an enhanced FL methodology with a novel preprocessing stage, including federated feature selection and global preprocessor construction, to address IoT/EC data heterogeneity. We developed a physical IoT testbed for attack simulations and data collection, implemented a tinyML-powered detector for realistic model validation, and also built a virtual testbed for scalable evaluations of FL models across diverse network environments.

Cognitive cyber↗

FunM2C: A Filter for Uncertainty Visualization of Multivariate Data on Multi-Core Devices

Uncertainty visualization is an emerging research topic in data visualization because neglecting uncertainty in visualization can lead to inaccurate assessments. In this paper, we study the propagation of multivariate data uncertainty in visualization. Although there have been a few advancements in probabilistic uncertainty visualization of multivariate data, three critical challenges remain to be addressed. First, the state-of-the-art probabilistic uncertainty visualization framework is limited to bivariate data (two variables). Second, existing uncertainty visualization algorithms use computationally intensive techniques and lack support for cross-platform portability. Third, as a consequence of the computational expense, integration into production visualization tools is impractical. In this work, we address all three issues and make a threefold contribution. First, we take a step to generalize the state-of-the-art probabilistic framework for bivariate data to multivariate data with an arbitrary number of variables. Second, through utilization of VTK-m’s shared-memory parallelism and cross-platform compatibility features, we demonstrate acceleration of multivariate uncertainty visualization on different many-core architectures, including OpenMP and AMD GPUs. Third, we demonstrate the integration of our algorithms with the ParaView software. We demonstrate the utility of our algorithms through experiments on multivariate simulation data with three and four variables.

Hari, Gautam↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

Top Research Challenges and Opportunities for Near Real-Time Extreme-Scale Visualization of Scientific Data

The rapid advancement in scientific simulations and experimental facilities has resulted in the generation of vast amounts of data at unprecedented scales. The analysis and visualization of large amounts of data is a challenge in and of itself, but the requirements for timeliness significantly magnify these difficulties. Near real-time visualization is critical to monitor and analyze the data produced by these large facilities, but current production tools are not well-suited to these requirements. In this position paper, we share our perspective on some of the challenges, and thus, opportunities for research that stand in the way of near-real-time visualization of large scientific data.

Pugmire, Dave↗

PISCES two-detector covariance matrix fit for the NOvA Experiment

NOvA is a long-baseline neutrino oscillation experiment with two functionally identical detectors: a Near Detector (ND) at Fermilab, placed 1 km from the neutrino source, and a Far Detector (FD) located 810 km away from the ND in Minnesota. NOvA's primary physics goals are the precise measurements of neutrino oscillation parameters $\theta_{23}$ and $\Delta m^2_{32}$ , determine the neutrino mass ordering, and constrain the value of $\delta_{CP}$, via the study of muon neutrino to electron neutrino oscillation. In the standard NOvA three-flavor analysis, oscillation parameters are extracted using an extrapolation technique in which the ND data constrain the FD prediction through a ratio method. While this allows for systematic uncertainties sharing the same effects in both detectors to cancel, it remains an FD-only fit and does not fully leverage the constraining power of the high-statistics ND. This analysis proposes a simultaneous ND+FD fit using the PISCES method. PISCES (Parameter Inference with Systematic Covariance and Exact Statistics) is a framework designed to support complex configurations such as a joint ND+FD fit. This allows PISCES to take full advantage of the ND data to directly constrain systematic uncertainties across all samples. In PISCES, systematic uncertainties are encoded in a fractional covariance matrix, and statistical uncertainties are handled with a Poisson likelihood, making the approach well suited for low-statistics samples. For interpretability, we further use a Newton–Raphson + PCA method to recover per-systematic pulls from the covariance formulation. This poster presents the full PISCES joint ND+FD fit for the NOvA three-flavor analysis, describes its implementation and evaluates its performance through extensive robustness tests and fake data studies. It also provides a comparison between the PISCES joint ND+FD results and the standard NOvA extrapolation method.

Rajaoalisoa, Miriama [Cincinnati U.] (ORCID:000000↗

NEPATEC2.0: NEPA Text Corpus v2.0

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

environmental review↗

NEPATEC v2.0: Standardized Metadata and Text Corpus of National Environmental Policy Act Documents

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

54 ENVIRONMENTAL SCIENCES↗

Core Model Proposal 401: Ukraine as an independent region in GCAM

The goal of this core model proposal (CMP) is to break out Ukraine from the Europe_Eastern region. This work aims to establish Ukraine as an independent region in the GCAM core (region 14) while moving Belarus and Moldova to region 15 (Europe_Non_EU). We have: 1) Updated several mappings to recode region 14 (formerly Europe_Eastern) as Ukraine and moved Belarus and Moldova to region 15 (Europe_Non_EU); 2) Updated several assumptions in the raw data files which provide information by region to reframe Ukraine as the 14th region, including coefficients, base year values, share weight interpolation values and rules, pipeline networks for gas trade, elasticities, shares, etc. 3) Changed documentation and in-code comments at several places referring to fixed 32 regions in GCAM to indicate that GCAM can have any number of regions; 4) Updated code base in gcamdata to dynamically process data for Ukraine given special cases.

Global Change Analysis Model (GCAM)↗

Combinatorial transcription factor binding encodes cis -regulatory wiring of mouse forebrain GABAergic neurogenesis

Transcription factors (TFs) bind combinatorially to cis-regulatory elements, orchestrating transcriptional programs. Although studies of chromatin state and chromosomal interactions have demonstrated dynamic neurodevelopmental cis-regulatory landscapes, parallel understanding of TF interactions lags. To elucidate combinatorial TF binding driving mouse basal ganglia development, we integrated chromatin immunoprecipitation sequencing (ChIP-seq) for twelve TFs, H3K4me3-associated enhancer-promoter interactions, chromatin and gene expression data, and functional enhancer assays. We identified sets of putative regulatory elements with shared TF binding (TF-pRE modules) that orchestrate distinct processes of GABAergic neurogenesis and suppress other cell fates. The majority of pREs were bound by one or two TFs; however, a small proportion were extensively bound. These sequences had exceptional evolutionary conservation and motif density, complex chromosomal interactions, and activity as in vivo enhancers. Our results provide insights into the combinatorial TF-pRE interactions that activate and repress expression programs during telencephalon neurogenesis and demonstrate the value of TF binding toward modeling developmental transcriptional wiring.

59 BASIC BIOLOGICAL SCIENCES↗

Accelerating Structure–Property Relationship Discovery with Multimodal Machine Learning and Self-Driving Microscopy

Microscopy combined with local spectroscopy is widely used to correlate nanoscale structure with functional properties in materials, but conventional measurements rely heavily on human-selected sampling locations and predefined targets, limiting data set diversity and the potential for discovery. Here, we present a framework that integrates autonomous microscopy with dual-novelty deep kernel learning (DN-DKL) for adaptive data acquisition and a dual variational autoencoder (VAE) for representation learning. DN-DKL actively guides the microscopy toward structurally and spectroscopically novel regions, enabling efficient collection of large spectral data sets. Dual-VAE embeds local structures and spectroscopic responses into a shared latent manifold that serves as a structure–property relationship map. We applied this framework for the investigation of halide perovskite films by using conductive atomic force microscopy. The results reveal distinct hysteresis behaviors that are linked to specific nanoscale structural motifs, including grain boundary junction points that show hysteresis under different bias conditions and asymmetric grain boundaries that suppress the charge transport. This framework establishes a general strategy that leverages the complementary strengths of self-driving microscopy, machine learning, and human expertise to accelerate scientific discovery in functional materials.

atomic force microscopy↗

HPC Campaign Management: Remote data access with user-defined error bound using ADIOS and ZFP

Remote access to large-scale scientific datasets, like those generated by combustion simulations or other high-performance computing (HPC) applications, presents a significant challenge. Downloading entire datasets is often impractical due to their size and the bandwidth limitations of typical networks. To address this challenge, we propose a novel approach that enables efficient remote access to large datasets distributed across multiple facilities. Our method enables technologies to download only the data values of a select variable, in a select region of interest, to a user-defined accuracy. For this purpose, we extended the ADIOS IO library to provide read functions with user-defined accuracy, a remote data server that understands multidimensional selections of specific variables, steps and accuracy from an ADIOS dataset, and which uses lossy compression on the remote site to reduce the data to be transferred back to the client. In addition, our extension of the ADIOS library collects metadata from multiple datasets in small files called Campaign Archives, which can be shared among project participants on any HPC, cloud or laptop, and which can easily facilitate the discovery of content and pointers to the data location as well as remote access to the data by local tools as if data was local. This feature called Campaign Management, enables a group of scientists to manage related datasets stored in multiple files, across multiple facilities as if it was in a single file/database. We demonstrate the effectiveness of our approach using a 1.5 TB dataset from the S3D combustion simulation on Frontier at the Oak Ridge Leadership Facility. Even a single variable from this dataset, at 64 GB, is too large to be processed on a standard laptop. We show two different reading patterns for 2D plots and 3D visualization, with careful settings that a scientist studying combustion data would do and show that running the same Python scripts on Frontier directly takes comparable time than running them on the local laptop with remote access to the data on Frontier.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X↗

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection Description This dataset contains input and output data for the manuscript Mongird, K. et al. (under review) titled "Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection". Input data corresponds to gridded spatial siting attributes that are necessary to conduct a random forest machine learning analysis of siting feature importance. Output data includes SHAP feature analysis outputs, and classification report values. For data on power plant siting results referred to in the manuscript, please refer to the CERF: IM3 Projected Western US Power Plant Locations data download page. The downloadable data includes values for eight different future scenarios for the Western US. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 Technical Information The dataset includes two sets of data files: (1) CERF gridded siting parameters and (2) Feature analysis outputs and classification reports. All downloadable data is in csv file format. Files with x/y coordinate information use the Albers Equal Area Conic projection (ESRI:102003). 1. CERF Gridded Siting Parameters This directory provides a balanced sample of gridded CERF siting parameters data for eight different scenarios for the Western US through 2055, seven different technologies, and eight timesteps. This data serves as input to the feature analysis. It contains the following parameters. region_name - name of region (i.e., state) sited - binary value representing whether the grid cell received a siting of that technology type (1=True) rcp - binary value representing scenario resource concentration pathway (0 = RCP4.5, 1 = RCP8.5) ssp - binary value representing scenario shared socioeconomic pathway (0 = SSP3, 1 = SSP5) climate - binary value representing cooler (0) or hotter (1) GCM forcing tech_name - generation technology name sited_year - year that values correspond to transmission_cost - cost of transmission interconnection pipeline_cost - cost of natural gas pipeline interconnection interconnection_cost - total interconnection cost (sum of transmission cost and gas pipeline cost) lmp - associated locational marginal value ($/MWh) associated with the grid cell, timestep, scenario, and technology xcoord - x-coordinate of location ycoord - y-coordinate of location 2a. Feature Analysis Output The dataset includes the feature analysis shap output for locational marginal price and interconnection cost. It contains the following parameters. technology - generator technology name scenario - name of scenario feature - name of feature, either locational_marginal_price or interconnection_cost value - the mean of absolute value of SHAP values for given feature 2b. Feature Analysis Classification Report This download includes the classification report associated with each random forest model. The dataset contains the following parameters. technology - generation technology name scenario - name of scenario test - one of precision (the proportion of predicted positives that are actually correct), recall (the proportion of actual positives that were correctly identified), f1-score (the harmonic mean of precision and recall) 0.0 - value of test for classification of 0 (grid cell not chosen for siting) 1.0 - value of test for classification of 1 (grid cell chosen for siting) accuracy - accuracy of model (i.e., fraction of all predictions that were right) macro avg - Simple average of test values for all classes weighted avg - Weighted average of test values for all classes, weighted based on Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

Towards provision of regularly updated climate data from the Coupled Model Intercomparison Project

The Coupled Model Intercomparison Project (CMIP) is a flagship of the World Climate Research Programme (WCRP). CMIP has become a recognised ‘brand’ in climate circles evolving over the last thirty years from a targeted research activity by a small number of climate modelling centres intercomparing their Earth System Model (ESM) simulations to a broad international coordinated research effort (Durack et al, 2025). CMIP is organized as a research activity leveraging funded and in-kind contributions from experts within modelling centres and the broader scientific community supported more recently by a fully-funded International Project Office. Within CMIP, Model Intercomparison Projects (MIPs) are community-designed to understand past, present and future climate. CMIP data provides a valuable resource for climate research and is routinely used to assess model representation of climate processes and test scientific hypotheses in the context of model uncertainty and (forced and internal) variability as evident from its prolific use in scientific publications1 . The impact relies on enabling infrastructure (most prominently via the Earth System Grid Federation (ESGF)), which allows sharing of simulation output, provision of the boundary conditions used in each simulation, and definition of the data standards that are essential to facilitating wide use of the data. The impact is supplemented by the wide-ranging scrutiny to which model simulations are subjected. Beyond its use in research, CMIP data is a key resource for communities producing derived climate information from downscaling and impact studies, such as the Coordinated Regional Downscaling Experiment (CORDEX; Gutowski et al., 2016) and the Intersectoral Impacts MIP (ISIMIP; Frieler et al., 2024). Government, academic and commercial entities also increasingly rely on CMIP and its downstream data for climate risk assessments and climate services (for example, Copernicus Climate Change Service and World Bank portal). This means that, although CMIP is a research activity, it increasingly serves a secondary and very relevant role as a provider of climate data – a long-recognised dichotomy (Stevens, 2024). Research and applications have distinct needs, with the former requiring flexibility and generality and the latter consistency. Here we explain how the design of the research activity has been adapted to reduce the burdens imposed by applications and how the research infrastructure might evolve to further enable scientific inquiry. We propose one possible approach to consistently providing model information and projections for applications in the future.

Environmental sciences↗

Towards continual machine learning for particle accelerators

This talk covers our work on errant beam prognostics at the Spallation Neutron Source (SNS), focusing on the end-to-end process from data collection to the development and deployment of predictive models in specific. A short overview of AIML work done for accelerators and current trends will be presented. We will walk through key steps involved in creating robust Machine Learning (ML) models, including model training, validation, and deployment in an operational setting. In addition to presenting our technical approach, we will share valuable lessons learned, emphasizing the importance of infrastructure to support the continuous adaptation of models to evolving data and system behaviors. This talk will provide insights into the challenges and solutions involved in applying ML to real-world operational environments, with a particular focus on managing data drift and changes in accelerator setup while ensuring model resilience over time.

Accelerator Physics↗

Supervisory Control and Data Acquisition for Electrochemical Separation Experimentation

The Python-based program is a laboratory automation tool designed to control and monitor electrochemical systems. The tool was developed for capacitive deionization (CDI) experiments, but it can be used for any system that requires controlled voltage or current segments and multi-parameter monitoring. The program integrates hardware components to run user-defined experimental parameters, providing operational control of a programmable power supply, peristaltic pump, and data acquisition devices. Currently, the program is structured with a workflow that includes an initialization (or pre-run) phase, a main loop, and a post-experiment stabilization (or post-run) phase. The initialization phase prepares and stabilizes the cell, ensuring that the electrodes and solution reach a baseline state before the experiment begins. The main loop consists of multiple voltage segments that repeat, controlling the experiment while recording key parameters such as time, voltage, current, pH, and conductivity. Finally, the post-experiment stabilization phase allows the system to stabilize after the experiment, returning the cell and solution to equilibrium conditions before ending the sequence. The program is designed with four variations, each tailored to different experimental needs. All variations include both the initialization and post-experiment stabilization stages, which run for a set amount of time, voltage, current, and flow rate before and after the main experiment block. The main loop runs for a set number of cycles, as defined by the user input, and each cycle is composed of 2 or 4 segments. The 4 program variations are described as follows: Program 1: The main program includes 2 segments. Each segment is defined to have a set duration, flow rate, voltage, and current. This program measures conductivity, flow rate, voltage, and current. Program 2: The main program expands Program 1 to include 4 segments. Each segment has a specified duration, flow rate, voltage, and current. Like Program 1, it measures conductivity, flow rate, voltage, and current. Program 3: The main program consists of 2 segments, each defined by time, flow rate, voltage, and current. In addition to conductivity, flow rate, voltage, and current, Program 3 collects pH and temperature data through a 4-channel data acquisition device. Program 4: This program independently controls two channels of a multi-channel power supply simultaneously. While conductivity can only be measured for one cell at a time, the dual-channel control makes it possible to operate two cells simultaneously under different voltage/current conditions. The main program includes 2 segments.For each program, all measurements are automatically logged and integrated into a single Excel output file. Data are displayed in numerical format and plotted, both in real time, to track system performance. A key feature of the program is its ability to synchronize all outputs so that every measurement shares a single timestamp, ensuring accurate alignment of voltage, current, pH, conductivity, and pH data.By combining hardware control, real-time monitoring, and unified data collection, this program significantly reduces manual workload and minimizes errors, making it a reliable platform for researchers, engineers, and laboratory technicians conducting CDI experiments, among other electrochemical tests.

Valentino, Lauren [Argonne National Laboratory (AN↗

Characterization of SimulCam, a standoff Raman system for scientific support of SuperCam operations on Mars

During the development activities of SuperCam Calibration Target, target intended for one of the two first Raman instruments to be deployed on another planetary body, our group developed a laboratory instrument that could simulate to some extent the Raman capabilities of one of such instruments and could provide data with similar quality. The use of this kind of laboratory instruments has demonstrated its utility in the evaluation of potential calibration targets or anticipating the science outcome that an instrument could provide. The present work describes our laboratory setup to support SuperCam, evaluating similarities between both instruments, despite of differences in the hardware. Evaluation of data gathered by SuperCam on Mars and the availability of one replica of SuperCam’s Calibration Target allowed the comparison on the same set of targets, demonstrating how similar Signal-to-Noise Ratio (SNR) could be achieved from both instruments. The higher energy per pulse on SimulCam is compensated by a greater analytical footprint and the use of smaller collection optics. The results show how spectra obtained at representative distances of SuperCam are comparable. Operational principles are also comparable in terms of time resolution, and close in terms of spectral resolution. This similarity has allowed different science support works using SimulCam data, as well as the support to Mars detections using our setup. We provide examples of this support that will be shared with the community in different papers, as well as examples of possible operations activities that could benefit from experiments performed with SimulCam. We show how this setup can complement the two laboratory replicas in Los Alamos and Toulouse in providing support data to different experiments.

47 OTHER INSTRUMENTATION↗