Search NASA⌕ Search

SEARCH · Search NASA

Results for “Automated Testing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Automation of Vulnerability and Patch Management: Information Extraction, Association, and Optimization

Vulnerability and patch management is an integral part of a robust cybersecurity program, yet it grows increasingly complex due to the sheer amount of data that must be analyzed. Particularly in Operational Technology (OT) environments, analysis must be done manually because of the lack of automated solutions. Additionally, there are many steps in this process, from the initial discovery of the vulnerability to the implementation of its remediation, and each step in the process requires different data in order to be performed effectively. In this work, we provide approaches and strategies to assist operators in industrial or OT environments throughout the vulnerability management cycle. Security advisories provide key information about mitigation strategies, or actions that can be taken when a patch is unavailable or cannot be installed. Details of these strategies are not shared in public vulnerability databases and must be found manually. We approach this problem by designing a solution to automatically identify that information within vendor security advisories and retrieve it for operator use. We start with an approach that requires domain-specific knowledge of certain frequently-seen reference websites. Next, an approach that can work on an arbitrary website but relies on certain keywords. Finally, an approach that uses Natural Language Processing (NLP) methods and does not require specific knowledge or keywords. Each of these approaches is more general than its predecessor; we demonstrate high accuracy for all approaches Advisories also often contain details of affected products in non-standard or natural language formats. While this information can be easily understood when read by an operator, the non-standard format acts as a barrier to effective automation. We provide an approach for the first step in this process: identifying vendors in security advisories and mapping them to a standard framework for representing digital assets and software products. We evaluate five established string similarity algorithms, plus one of our own design that combines string similarity and information theory, on the task of mapping vendors to their corresponding entries in the Common Platform Enumeration (CPE) repository. Our results show that our proposed metric outperforms all others. Due to the constraints on time, finances, and personnel for organizations, Large Language Models (LLMs) may seem like attractive opportunities for security operators to speed up information gathering; however, it is still not clear whether LLMs can handle vulnerability management tasks well. To answer this question, we perform an empirical study of LLMs’ ability to provide consistent, accurate information about vulnerabilities in order to guide organizations in their adoption of LLMs. We observe poor performance for all models tested, suggesting that these models are not well-suited to the consistent retrieval of accurate vulnerability information. Finally, once vulnerabilities have been identified and any additional information has been obtained, operators must decide which remediation actions to implement based on their available resources. This already-complex problem becomes even more so when we consider that a vulnerability may have multiple avenues for remediation. We formulate this scenario as two knapsack problems and provide solutions, which we then compare against several existing strategies for vulnerability prioritization seen in real operational environments.

McClanahan, Kylie↗

An Empirical Validation of a Constrained Bin Packing Algorithm for a Home Energy Management System

The increasing number of intelligent electrical appliances and home energy management systems provide a big opportunity for demand response services from residential and small commercial buildings to the grid. Simultaneously, direct control of individual devices by utilities can cause communication bottlenecks, as well as coordination and privacy concerns. These challenges can be addressed by combining the constituent devices into a single house battery equivalent for the purposes of demand response, using Minkowski sum and a 2d bin packing problem. However, the well-studied traditional problems have not been tested in a real house, as implementation carries significant challenges of its own. We deploy the packing problem on residential devices in a controllable house. We report the barriers we found, such as charge forecast and scalability of the algorithm, and discuss our solutions. The study serves as an intermediate step between existing theoretical research and possible future steps, such as prototype deployments of systems that provide residential demand response.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Aerosol Envelope Sealing of Existing Residences

This report explores the best methods for aerosol envelope sealing of unoccupied, existing residences and documents typical leakage reductions. The project consisted of three distinct efforts: (1) scaled field demonstrations of the sealing process, tracking from setup to sealing and cleanup, (2) laboratory testing of new sealants that dry clear, making them more appropriate for retrofit applications, and (3) BEopt™ (Building Energy Optimization Tool) modeling of the energy implications of the measured reductions in leakage. Overall, aerosol sealing performance in existing homes was effective, with an average leakage reduction of 47% across all 34 sites. Conventional approaches typically only produce leakage reductions of 25%–30%. The aerosol sealing technology could provide a process for existing homes and multifamily units to gain the benefits of a well-sealed home at a reasonable cost with minimal disruption.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Flux REaction TArget Prioritization (Flux RETAP) v1

Metabolic engineering is evolving rapidly as a result of new advances in synthetic biology and automation, as well as the irruption of machine learning (ML). ML has been shown to provide the predictive power synthetic biology lacked and needed, and to be able to effectively guide the metabolic engineering process. However, current technical limitations prevent the independent application of ML approaches to metabolic engineering without the use of previous biological knowledge in the form of a prioritized list of desirable engineering targets. Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale metabolic models (GSMs) for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing metabolite production. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production in the literature accessible to us, 50% of targets that experimentally improved taxadiene production in E. coli and ~60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets which can also be utilized in ML pipelines.

Czajka, Jeffrey [Battelle Memorial Institute, Paci↗

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Czajka, Jeffrey J↗

Transformers and Long Short-Term Memory Transfer Learning for GenIV Reactor Temperature Time Series Forecasting

Automated monitoring of the coolant temperature can enable autonomous operation of generation IV reactors (GenIV), thus reducing their operating and maintenance costs. Automation can be accomplished with machine learning (ML) models trained on historical sensor data. However, the performance of ML usually depends on the availability of large amount of training data, which is difficult to obtain for GenIV, as this technology is still under development. We propose the use of transfer learning (TL), which involves utilizing knowledge across different domains, to compensate for this lack of training data. TL can be used to create pre-trained ML models with data from small-scale research facilities, which can then be fine-tuned to monitor GenIV reactors. In this work, we develop pre-trained Transformer and long short-term memory (LSTM) networks by training them on temperature measurements from thermal hydraulic flow loops operating with water and Galinstan fluids at room temperature at Argonne National Laboratory. The pre-trained models are then fine-tuned and re-trained with minimal additional data to perform predictions of the time series of high temperature measurements obtained from the Engineering Test Unit (ETU) at Kairos Power. The performance of the LSTM and Transformer networks is investigated by varying the size of the lookback window and forecast horizon. The results of this study show that LSTM networks have lower prediction errors than Transformers, but LSTM errors increase more rapidly with increasing lookback window size and forecast horizon compared to the Transformer errors.

LSTM↗

Drug providers’ perspectives on antibiotic misuse practices in eastern Ethiopia: a qualitative study

Objective Antibiotic misuse includes using them to treat colds and influenza, obtaining them without a prescription, not finishing the prescribed course and sharing them with others. Although drug providers are well positioned to advise clients on proper stewardship practices, antibiotic misuse continues to rise in Ethiopia. It necessitates an understanding of why drug providers failed to limit such risky behaviours. This study aimed to explore drug providers’ perspectives on antibiotic misuse practices in eastern Ethiopia. Setting The study was conducted in rural Haramaya district and Harar town, eastern Ethiopia. Design and participants An exploratory qualitative study was undertaken between March and June 2023, among the 15 drug providers. In-depth interviews were conducted using pilot-tested, semistructured questions. The interviews were transcribed verbatim, translated into English and analysed thematically. The analyses considered the entire dataset and field notes. Results The study identified self-medication pressures, non-prescribed dispensing motives, insufficient regulatory functions and a lack of specific antibiotic use policy as the key contributors to antibiotic misuse. We found previous usage experience, a desire to avoid extra costs and a lack of essential diagnostics and antibiotics in public institutions as the key drivers of non-prescribed antibiotic access from private drug suppliers. Non-prescribed antibiotic dispensing in pharmacies was driven by client satisfaction, financial gain, business survival and market competition from informal sellers. Antibiotic misuse in the setting has also been linked to traditional and ineffective dispensing audits, inadequate regulatory oversights and policy gaps. Conclusion This study highlights profits and oversimplified access to antibiotics as the main motivations for their misuse. It also identifies the traditional antibiotic dispensing audit as an inefficient regulatory operation. Hence, enforcing specific antibiotic usage policy guidance that entails an automated practice audit, a responsible office and insurance coverage for persons with financial limitations can help optimise antibiotic use while reducing resistance consequences.

General & Internal Medicine↗

Lab-Scale Cable-Driven Parallel Robot Prototype for Automated Prefabricated Component Manipulation

This paper presents the design and evaluation of a lab-scale cable-driven parallel robot (CDPR) developed as a flexible platform for automated installation of prefabricated components onto exterior building envelopes. Traditional manual installation methods for prefabricated components, which depend on scaffolding, cranes, cherry pickers, and verbal coordination, are not only labor-intensive and error-prone but also face significant limitations in dense urban environments due to site access constraints. To address these challenges, we developed a lab-scale CDPR platform capable of autonomously transporting building envelope components from a designated pickup zone to their target installation location, minimizing the need for human intervention. This study describes the system’s mechanical design, actuation architecture, real-time feedback system, and control strategy of the CDPR, and evaluates its performance in a laboratory environment. The robot’s actuation system uses torque control for end-effector manipulation. The robot’s real-time pose feedback comes from a construction-grade total station and a wireless inertial measurement unit (IMU), which together support precise end-effector control. Experimental results demonstrate the successful integration of the hardware, sensing, state estimation, and control subsystems. Preliminary tests showed that our lab-scale prototype can position the end effector with an error of less than 3 mm, which is a level of precision not previously achieved by existing CDPRs in construction applications. The key findings are twofold: (1) torque-only control is necessary but not sufficient for minimizing final pose error, and (2) incorporating real-time pose feedback can achieve the desired placement accuracy.

Liu, Yifang [Oak Ridge National Laboratory (ORNL),↗

Development of a Flow Stabilization Algorithm Enabling Measurement of Low Oxygen Concentrations in Liquid Sodium

To enhance the automated control of the plugging meter (PM) and thereby enhance detection fidelity in ultralow oxygen environments [≤1 parts per million by weight (wppm)], a novel proportional derivative controller has been implemented with conventional PM hardware. This ramp sign stabilized flow (RSSF) controller manipulates the sign (heating or cooling direction) at a fixed rate, enabling precise temperature adjustment around the saturation temperature of the bulk sodium. This adjustment helps maintain flow stability in a partially formed sodium oxide plug, thus greatly reducing the temperature amplitude in the plugging cycle and promoting simple and accurate oxygen determinations in addition to an increased sampling rate. Rather than relying on the subjective nature of indexing the time when the flow rate changes due to the plugging or unplugging onset to the PM temperature, a running average of the correlated oxygen concentration with time over multiple plugging events can provide oxygen readings ranging from an absolute uncertainty of 500 wppb in real time to less than 50 wppb for a 24-h sampling window. Finally, the RSSF controller was tested at 508 ± 7 wppb with measured oxygen of 542 ± 179 wppb, further reducing the variance between the saturation temperature and the plugging temperature.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗

Comparing Control Performance Between Simulation and Experiment using the Microreactor Automated Control System Testbed

In the advanced reactor domain, a flexible and scalable software/hardware infrastructure is crucial for integrating and validating various control technologies. This study used the Microreactor Automated Control System (MACS) hardware platform as a testbed. MACS was originally designed to mirror Idaho National Laboratory (INL)'s Microreactor Applications Research Validation and Evaluation (MARVEL), a 85-kW thermal fission microreactor. It features control drums for simulated reactivity control; lights that function as a surrogate reactor core, with the brightness being proportional to the reactor power; and light sensors that emulate neutron detectors. To transform MACS into a physical twin of MARVEL for evaluating control methods, the Control and Optimization Modular Modeling Application for Nuclear Deployment (COMMAND) software was employed. This software integrated the hardware with two models of the MARVEL core, based on Reactor Excursion and Leak Analysis Program (RELAP5-3D) and Monte Carlo N-Particle (MCNP) models. The study aimed to demonstrate the gap between control theory and actual practice—a gap that often necessitates empirical adjustments such as control gain retuning, filters, time discretization, and integrator anti-windup measures. Controllers were developed based on increasingly complex simulations without hardware, starting from the base MARVEL model and then introducing actuator saturation constraints and sensor noise. The final control strategy was then tested using MACS, and a comparative performance analysis was conducted.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

A Climatology and Life‐Cycle Characteristics of Atmospheric Fronts and Their Associated Precipitation

Abstract Atmospheric fronts are one of the main sources of mid‐latitude variability. We employ a novel method for identifying and tracking fronts and frontal precipitation. Thermal and dynamical variables are used to identify fronts as areal objects in space, which are tracked in time using the open‐source TempestExtremes software package. Precipitation objects are co‐located to identify frontal precipitation. The method is subjected to validation and sensitivity tests using manually curated data from the National Weather Service. Climatologies of fronts and frontal precipitation are computed from reanalysis and observations; fronts are present upwards of 14% of the time in the storm tracks, and represent the majority (up to 90%) of total and extreme precipitation. Novel aspects of the method are showcased through the lifetime characteristics of fronts across North America. Three sets of warm and cold fronts were discovered, and their duration, distance‐traveled, and translation velocity are examined. Plain Language Summary Mid‐latitude low‐pressure systems and weather fronts are important for our day‐to‐day experience of weather events, particularly in the mid‐latitudes. This work makes use of standardized atmospheric data and creates a method of automatically tracking these important atmospheric features and their precipitation to quantify their relative role in global precipitation. Weather fronts are persistent in the mid‐latitudes and are associated with the majority of precipitation–particularly the most intense precipitation. Trajectories of fronts over North America are categorized to create a set of archetypal fronts that occur in that region. The differences between these types of fronts are characterized. Key Points An automated, efficient, and skillful frontal detection algorithm is developed and validated Fronts contribute a larger fraction of extreme precipitation than all precipitation in mid‐latitude storm tracks Fronts across North America have substantial variation in characteristics depending on their origin location

extratropical cyclone↗

Traffic Control via Connected and Automated Vehicles (CAVs): An Open-Road Field Experiment with 100 CAVs

The CIRCLES project aims to reduce instabilities in traffic flow, which are naturally occurring phenomena due to human driving behavior. Also called “phantom jams” or “stop-and-go waves,” these instabilities are a significant source of wasted energy. Toward this goal, the CIRCLES project designed a control system, referred to as the MegaController by the CIRCLES team, that could be deployed in real traffic. Our field experiment, the MegaVanderTest (MVT), leveraged a heterogeneous fleet of 100 longitudinally controlled vehicles as Lagrangian traffic actuators, each of which ran a controller with the architecture described in this article. The MegaController is a hierarchical control architecture that consists of two main layers. The upper layer is called the Speed Planner and is a centralized optimal control algorithm. It assigns speed targets to the vehicles, conveyed through the LTE cellular network. The lower layer is a control layer, running on each vehicle. It performs local actuation by overriding the stock adaptive cruise controller, using the stock onboard sensors. The Speed Planner ingests live data feeds provided by third parties as well as data from our own control vehicles and uses both to perform the speed assignment. The architecture of the Speed Planner allows for the modular use of standard control techniques, such as optimal control, model predictive control (MPC), kernel methods, and others. The architecture of the local controller allows for the flexible implementation of local controllers. Corresponding techniques include deep reinforcement learning (RL), MPC, and explicit controllers. Depending on the vehicle architecture, all onboard sensing data can be accessed by the local controllers or only some. Likewise, control inputs vary across different automakers, with inputs ranging from torque or acceleration requests for some cars to electronic selection of adaptive cruise control (ACC) setpoints in others. The proposed architecture technically allows for the combination of all possible settings proposed previously, that is {Speed Planner algorithms} × {local Vehicle Controller algorithms} × {full or partial sensing} × {torque or speed control}. As a result, most configurations were tested throughout the ramp up to the MegaVandertest (MVT).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Marine and continental stratocumulus cloud microphysical properties obtained from routine ARM Cimel sunphotometer observations

This study investigates marine and continental stratocumulus (Sc) cloud properties obtained from an automated implementation of a multispectral photometer retrieval. Photometer methods simultaneously retrieve cloud optical depth (τ) and cloud droplet effective radius (r e ), with estimates for liquid water path (LWP) calculated on the availability of those quantities. These applied methods evaluate retrieved cloud properties for Sc identified during a recent 6 year period over the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) program sites in Oklahoma, USA (SGP) and in the Azores, Portugal (ENA). Modest agreement in key quantity retrievals is found between the routine photometer products and multisensor collocated profiling references. Cumulative breakdowns contingent on cloud thickness indicate increases in all retrieved quantities in thicker clouds, with larger discrepancies in the relative performance between the retrievals collected in the presence of drizzle. Under continental cloud conditions, the clouds of a similar thickness and r e to those sampled under marine conditions report a factor of 1.5 larger τ and LWP. An r 2 ≅0.65 is found between photometer τ retrievals and shadowband radiometer measurements, with photometer retrievals reporting a high (relative) bias. The τ intercomparisons indicate that variability between retrievals is a factor of three larger than errors reported from individual retrieval input perturbation tests. Photometer r e retrievals suggest a low r 2 (< 0.1) having a standard deviation ≅ 3 µm when compared to ARM baseline multi-sensor radar/radiometer references (accounting for offsets in the cloud droplet number concentration assumptions of the latter). However, photometer LWP calculations remain relatively unbiased in non-drizzling conditions, with errors O (50 g m −2 ) and r 2 ≅0.5 to collocated radiometer and interferometer references. Additional sensitivity tests for island influences on marine Sc properties suggest that while island-influenced winds may promote larger cloud LWP or thickness, the influence could be within retrieval method uncertainty and/or collocated instrument variability.

54 ENVIRONMENTAL SCIENCES↗

Modeling Oxidative Dehydrogenation of Propane with Supported Vanadia Catalysts Using Multireference Methods

The oxidative dehydrogenation of propane over supported vanadium oxide catalysts poses significant computational challenges due to complex electronic structure changes along the reaction coordinate, driven primarily by changes in the oxidation states of vanadium. To address these challenges, we systematically test quantum chemical methods, including multireference (MR) approaches, domain-based local pair natural orbital coupled cluster theory (DLPNO-CCSD(T)), and density functional theory (DFT). The initial C–H bond-breaking transition state requires MR treatment due to its multireference character, while subsequent steps permit efficient single-reference calculations. For the rate-limiting C–H activation step mediated by the vanadyl moiety, complete 1 active space second-order perturbation theory (CASPT2) yields an apparent activation barrier (E app 600K ) of 138 kJ/mol, consistent with experimental values (134 ± 4 kJ/mol; Gruene et al. Catal. Today 2010, 157, 137). In contrast, DLPNO-CCSD(T) overestimates this barrier (198 kJ/mol), whereas DFT predictions span 125–150 kJ/mol, depending on the functional. Our multireference investigation of this transition metal oxide-catalyzed process demonstrates that an active space that incorporates the C–H σ and V=O σ/π bonding orbitals, oxygen lone pairs, and their antibonding counterparts adequately captures electronic structure changes along the chemical transformation. Furthermore, these findings provide a general strategy for active space selection in transition metal oxide-catalyzed C/O–H bond activation reactions. The reference dataset from this work, which includes MR calculations with manually selected active spaces for all intermediates and transition states in the propane ODH reaction network, will serve as a benchmark for automating active space selection in similar systems.

Catalysts↗

Accelerating polyketide synthase engineering for high TRY production of biofuels and bioproducts (Final Technical Report)

Polyketide synthase (PKS) enzymes have a modular, deterministic logic that holds the potential to act as a flexible chemical factory for the biological production of a huge diversity of valuable small molecule compounds. However, engineering a custom PKS to produce a specific desired product currently requires years of trial and error, for reasons that remain poorly understood. In this project, we have developed a rapid, high throughput, Design-Build-Test-Learn (DBTL) cycle for polyketide synthases (PKSs) and demonstrate its utility for production of materials precursors. The objectives are 1) to develop a rapid, high-throughput (HT) DBTL cycle for PKSs that will enable production of a large number of unnatural, organic molecules on demand at high titer, rate, and yield (TRY); 2) to demonstrate the utility of the PKS DBTL cycle to produce three molecules: one commodity chemical (caprolactam or valerolactam) and two novel materials precursors (caprolactam or valerolactam derivatives); and 3) to demonstrate the utility of the PKS DBTL cycle to increase the TRY of one molecule (caprolactam or valerolactam). In this project, we have successfully demonstrated our high throughput PKS DBTL pipeline, and have biologically produced valerolactam and several other novel nylon monomers.

60 APPLIED LIFE SCIENCES↗

SIDDA: SInkhorn Dynamic Domain Adaptation

Modern neural networks (NNs) often do not generalize well in the presence of a "covariate shift"; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more domain-invariant features. Domain adaptation (DA) methods include a range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SIDDA, an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, and real astronomical observations. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with equivariant neural networks (ENNs). We find that SIDDA enhances the generalization capabilities of NNs, achieving up to a ≈40% improvement in classification accuracy on unlabeled target data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, we find that SIDDA enhances model calibration on both source and target data--achieving over an order of magnitude improvement in the ECE and Brier score. SIDDA's versatility, combined with its automated approach to domain alignment, has the potential to advance multi-dataset studies by enabling the development of highly generalizable models.

Pandya, Sneh [Northeastern U.]↗

From natural language to control signals: a conceptual framework for semantic channel finding in complex experimental infrastructure

Modern experimental platforms such as particle accelerators, fusion devices, telescopes, and industrial process control systems expose tens to hundreds of thousands of control and diagnostic channels, accumulated over decades of hardware evolution. Operators and AI systems alike depend on informal expert knowledge, inconsistent naming conventions, and scattered documentation to locate the signals required for monitoring, troubleshooting, and automated control, creating a persistent bottleneck for reliability, scalability, and emerging language-model-driven interfaces. We formalize semantic channel finding, the task of mapping natural-language intent to concrete control-system signals, as a general problem in complex experimental infrastructure, and introduce a four-paradigm conceptual framework to guide architecture selection based on facility-specific data regimes. The paradigms span (i) direct in-context lookup over small, curated channel dictionaries, (ii) constrained hierarchical navigation through structured trees, (iii) interactive agent exploration using iterative reasoning and tool-based database queries, and (iv) ontology-grounded semantic search that decouples channel meaning from facility-specific naming conventions. We demonstrate the practical feasibility of each paradigm through proof-of-concept implementations at four operational facilities spanning two orders of magnitude in scale: from compact free-electron lasers to large synchrotron light sources, operating under diverse control-system architectures ranging from clean hierarchical naming schemes to legacy environments with decades of heterogeneous conventions. Where evaluated against expert-curated operational queries, these instantiations achieve 90%–97% accuracy, validating the framework’s applicability across real-world deployment scenarios. To accelerate adoption across the broader scientific and industrial control-system community, we release open-source, plug-and-play implementations of all three interactive paradigms-direct lookup, hierarchical navigation, and middle-layer exploration-within the Osprey framework, together with tools for channel database generation, interactive testing, and minimal-configuration deployment. This work establishes semantic channel finding as a foundational capability for human-centric and agentic AI interfaces at large-scale facilities, providing both a systematic framework for architecture design and practical resources to enable adoption without building custom infrastructure from scratch.

channel finding↗