Search NASA⌕ Search

SEARCH · Search NASA

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Automated Label‐Free Assay for Viral Detection and Inhibitor Screening via Biomembrane‐Functionalized Microelectrode Arrays

Most virus infection assays have indirect readout such as virus number following entry (e.g., PCR, cell lysis). While effective, these technologies are labor‐intensive, require specialized environments (e.g., sterile or RNA‐free), and detect later‐stage viral events like lysis or cell death, lacking sensitivity to early fusion events. To address these limitations, we present biologically relevant 2D membrane materials, host‐cell‐derived supported lipid bilayers (hcd‐SLBs), integrated with organic microelectrode arrays (OMEAs) for detection of severe acute respiratory syndrome coronavirus 2 (SARS‐CoV‐2) fusion. By overexpressing angiotensin‐converting enzyme 2 (ACE2) receptors on the native membranes, the platform functions as a viral sensor capable of detecting virus pseudo particles (VPPs) through the late pathway. Additionally, hcd‐SLBs extracted from human lung epithelium expressing native ACE2 detect fusion events through the early pathway. The platform's utility as a drug‐screening tool is demonstrated by testing antibodies targeting either the ACE2 on the host membrane or the viral spike (S) proteins. To enhance the throughput, microfluidics are integrated for automation and OMEAs are incorporated within each channel, miniaturizing the testing units. This system supports high‐throughput data generation, automation, and scalability, providing an efficient platform for viral fusion detection that advances the study of pathogen‐host interactions and accelerates antiviral drug discovery.

Biology↗

Classifying handedness in chiral nanomaterials using label error robust deep learning

Abstract High-throughput scanning electron microscopy (SEM) coupled with classification using neural networks is an ideal method to determine the morphological handedness of large populations of chiral nanoparticles. Automated labeling removes the time-consuming manual labeling of training data, but introduces label error, and subsequently classification error in the trained neural network. Here, we evaluate methods to minimize classification error when training from automated labels of SEM datasets of chiral Tellurium nanoparticles. Using the mirror relationship between images of opposite handed particles, we artificially create populations of varying label error. We analyze the impact of label error rate and training method on the classification error of neural networks on an ideal dataset and on a practical dataset. Of the three training methods considered, we find that a pretraining approach yields the most accurate results across label error rates on ideal datasets, where size and other morphological variables are held constant, but that a co-teaching approach performs the best in practical application.

36 MATERIALS SCIENCE↗

Automated instant labeling chemistry workflow for real-time monitoring of monoclonal antibody N -glycosylation

With the transition toward continuous bioprocessing, process analytical technology (PAT) is becoming necessary for rapid and reliable in-process monitoring during biotherapeutics manufacturing. Bioprocess 4.0 is looking to build end-to-end bioprocesses that include PAT-enabled real-time process control. This is especially important for drug product quality attributes that can change during bioprocessing, such as protein N-glycosylation, a critical quality attribute for most monoclonal antibody (mAb) therapeutics. Glycosylation of mAbs is known to influence their efficacy as therapeutics and is regulated for a majority of mAb products on the market today. Currently, there is no method to truly measure N-glycosylation using on-line PAT, hence making it impractical to design upstream process control strategies. We recently described the N-GLYcanyzer workflow: an integrated PAT unit that measures mAb N-glycosylation within 3 hours of automated sampling from a bioreactor. Here, we integrated Agilent's Instant Procainamide (InstantPC) based chemistry workflow into the N-GLYcanyzer PAT unit to allow for nearly 10× faster near real-time analysis of mAb glycoforms. Furthermore, our methodology is explained in detail to allow for replication of the PAT workflow as well as present a case study demonstrating the use of this PAT to autonomously monitor a mammalian cell perfusion process at the bench scale to gain increased knowledge of mAb glycosylation dynamics during continuous biologics manufacturing using Chinese hamster ovary (CHO) cells.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AI-Enabled Operations at Fermi Complex: Multivariate Time Series Prediction for Outage Prediction and Diagnosis

The Main Control Room of the Fermilab accelerator complex continuously gathers extensive time-series data from thousands of sensors monitoring the beam. However, unplanned events such as trips or voltage fluctuations often result in beam outages, causing operational downtime. This downtime not only consumes operator effort in diagnosing and addressing the issue but also leads to unnecessary energy consumption by idle machines awaiting beam restoration. The current threshold-based alarm system is reactive and faces challenges including frequent false alarms and inconsistent outage-cause labeling. To address these limitations, we propose an AI-enabled framework that leverages predictive analytics and automated labeling. Using data from $2,703$ Linac devices and $80$ operator-labeled outages, we evaluate state-of-the-art deep learning architectures, including recurrent, attention-based, and linear models, for beam outage prediction. Additionally, we assess a Random Forest-based labeling system for providing consistent, confidence-scored outage annotations. Our findings highlight the strengths and weaknesses of these architectures for beam outage prediction and identify critical gaps that must be addressed to fully harness AI for transitioning downtime handling from reactive to predictive, ultimately reducing downtime and improving decision-making in accelerator management.

Jain, Milan [PNL, Richland] (ORCID:000000021676111↗

Automated Credibility Assessments of User Features in Scientific Software

Scientific software (SciSoft) is complex, often containing a mixture of production capabilities co-mingled with features under active research and development. Furthermore, SciSoft is often developed over decades by non-computer scientists who may not have a strong background in or prioritize software architecture design, testing, and quality (e.g., test coverage). These conditions lead to difficulty in understanding which software components or functions implement what user-facing features and therefore those features’ software quality pedigree. This lack of understanding poses challenges in assessing readiness and credibility of user features, and often relies on a SciSoft subject matter expert’s (SME) laborious investigation and assertion. This final report of a one-year Computing and Information Sciences Lab Directed Research and Development project presents a general framework for modeling SciSoft architecture as a direct relationship between user features and the software components/functions that implement them. Our approach leverages automated labeling of the SciSoft’s regression test suite and employs machine learning algorithms to construct the architecture model. We demonstrate this framework on the Solid Mechanics component of the SIERRA multi-physics engineering analysis suite developed at Sandia National Laboratories.

97 MATHEMATICS AND COMPUTING↗

Active Learning Meets Foundation Models: Fast Remote Sensing Data Annotation for Object Detection

Object detection in remote sensing demands extensive, high-quality annotations—a process that is both labor-intensive and time-consuming. In this work, we introduce a real-time active learning and semi-automated labeling framework that leverages foundation models to streamline dataset annotation for object detection in remote sensing imagery. For example, by integrating a Segment Anything Model (SAM), our approach generates mask-based bounding boxes that serve as the basis for dual sampling: (a) uncertainty estimation to pinpoint challenging samples, and (b) diversity assessment to ensure broad data coverage. Furthermore, our Dynamic Box Switching Module (DBS) addresses the well-known cold start problem for object detection models by replacing its suboptimal initial predictions with SAM-derived masks, thereby enhancing early-stage localization accuracy. Extensive evaluations on multiple remote sensing datasets plus a real-world user study, demonstrate that our framework not only reduces annotation effort, but also significantly boosts detection performance compared to traditional active learning sampling methods. The code for training and the user interface will be made available.

Burges, Marvin [ORNL] (ORCID:0000000312690769)↗

The L-CAPE Project at FNAL

The controls system at FNAL records data asynchronously from several thousand Linac devices at their respective cadences, ranging from 15Hz down to once per minute. In case of downtimes, current operations are mostly reactive, investigating the cause of an outage and labeling it after the fact. However, as one of the most upstream systems at the FNAL accelerator complex, the Linac’s foreknowledge of an impending downtime as well as its duration could prompt downstream systems to go into standby, potentially leading to energy savings. The goals of the Linac Condition Anomaly Prediction of Emergence (L-CAPE) project that started in late 2020 are (1) to apply data-analytic methods to improve the information that is available to operators in the control room, and (2) to use machine learning to automate the labeling of outage types as they occur and discover patterns in the data that could lead to the prediction of outages. We present an overview of the challenges in dealing with time-series data from 2000+ devices, our approach to developing an ML-based automated outage labeling system, and the status of augmenting operations by identifying the most likely devices predicting an outage.

43 PARTICLE ACCELERATORS↗

Curation and Dissemination of Complex Multi-Modal Datasets for Radiation Detection, Localization, and Tracking

The PANDAWN sensor network in Chicago, IL, is a state-of-the-art testbed for networked, multi-modal sensing. It integrates AI/data science methods into its operation, from data acquisition to automated data labeling and curation workflows. The curation and dissemination of diverse multi-modal datasets will enable the development of new radiological/nuclear (R/N) detection, localization, and tracking algorithms and methods relevant across the nonproliferation mission space. This article first introduces the PANDAWN sensor network and the features that make it stand out from previous multi-modal data acquisition efforts. We then review the various data streams acquired on the PANDAWN nodes and present the implementation of an automated data curation pipeline that includes the labeling of radiation and contextual data streams. Here, we finally provide a short overview of different studies that leveraged the curated datasets.

Data curation↗

Creating a Training Dataset for Semantic Segmentation of Canal Networks for Irrigation Modernization

Canal infrastructure has provided critical irrigation water to the western United States for over a century. To continue providing vital water resources to the semi-arid West, irrigation systems must undergo maintenance and modernization. Many canal companies are resource-constrained, and because funding opportunities often require detailed knowledge of existing infrastructure, they can struggle to secure financial capital. We address this problem by creating training data for a semantic segmentation deep learning model to map canal networks throughout the western United States. To create a diverse and robust training dataset, we labelled 1-m NAIP imagery with the locations of no canals, wet canals, and dry/vegetated canals. Since creating these datasets is time consuming, we first developed a preprocessing methodology to identify canals within our four study areas. We used NAIP imagery and provided canal centerline data to buffer, standardize, and cluster the imagery, automating the labeling process as much as possible. However, this still required manual cleaning and manual classification of canal type. Challenges arose when canals were interrupted (e.g., road culverts or piped sections) or when nearby features shared similar characteristics (e.g., irrigated fields, trees, and shadows). Combining automated preprocessing with manual refinement produced four detailed canal masks to be used in the semantic segmentation model developed by Richard Tapia.

13 - HYDRO ENERGY↗

The Geothermal Artificial Intelligence for geothermal exploration

Exploration of geothermal resources involves analysis and management of a large number of uncertainties, which makes investment and operations decisions challenging. Remote Sensing (RS), Machine Learning (ML) and Artificial Intelligence (AI) have potential in managing the challenges of geothermal exploration. In this paper, we present a methodology that integrates RS, ML and AI to create an initial assessment of geothermal potential, by resorting to known indicators of geothermal areas namely mineral markers, surface temperature, faults and deformation. We demonstrated the implementation of the method in two sites (Brady and Desert Peak geothermal sites) that are close to each other but have different characteristics (Brady having clear surface manifestations and Desert Peak being a blind site). Here, we processed various satellite images and geospatial data for mineral markers, temperature, faults and deformation and then implemented ML methods to obtain pattern of surface manifestation of geothermal sites. We developed an AI that uses patterns from surface manifestations to predict geothermal potential of each pixel. We tested the Geothermal AI using independent data sets obtaining accuracy of 92-95%; also tested the Geothermal AI trained on one site by executing it for the other site to predict the geothermal / non-geothermal delineation, the Geothermal AI performed quite well in prediction with 72-76% accuracy.

15 GEOTHERMAL ENERGY↗

A hybrid machine-learning approach for analysis of methane hydrate formation dynamics in porous media with synchrotron CT imaging

Fast multi-phase processes in methane hydrate bearing samples pose a challenge for quantitative micro-computed tomography study and experiment steering due to complex tomographic data analysis involving time-consuming segmentation procedures. This is because of the sample's multi-scale structure, which changes over time, low contrast between solid and fluid materials, and the large amount of data acquired during dynamic processes. Here, a hybrid approach is proposed for the automatic segmentation of tomographic data from time-resolved imaging of methane gas-hydrate formation in sandy granular media, which includes a deep-learning 3D U-Net model. To prepare a training dataset for the 3D U-Net, a technique to automate data labeling based on sample-specific information about the mineral matrix immobility and occasional fluid movement in pores is proposed. Automatic segmentation allowed for studying properties of the hydrate growth in pores, as well as dynamic processes such as incremental flow and redistribution of pore brine. Results of the quantitative analysis showed that for typical gas-hydrate stability parameters (100 bar methane pressure, 7°C temperature) the rate of formation is slow (less than 1% per hour), after which the surface area of contact between brine and gas increases, resulting in faster formation (2.5% per hour). Hydrate growth reaches the saturation point after 11 h of the experiment. Finally, the efficacy of the proposed segmentation scheme in on-the-fly automatic data analysis and experiment steering with zooming to regions of interest is demonstrated.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

End-to-end deep learning pipeline for real-time Bragg peak segmentation: from training to large-scale deployment

X-ray crystallography reconstruction, which transforms discrete X-ray diffraction patterns into three-dimensional molecular structures, relies critically on accurate Bragg peak finding for structure determination. As X-ray free electron laser (XFEL) facilities advance toward MHz data rates (1 million images per second), traditional peak finding algorithms that require manual parameter tuning or exhaustive grid searches across multiple experiments become increasingly impractical. While deep learning approaches offer promising solutions, their deployment in high-throughput environments presents significant challenges in automated dataset labeling, model scalability, edge deployment efficiency, and distributed inference capabilities. We present an end-to-end deep learning pipeline with three key components: (1) a data engine that combines traditional algorithms with our peak matching algorithm to generate high-quality training data at scale, (2) a modular architecture that scales from a few million to hundreds of million parameters, enabling us to train large expert-level models offline while deploying smaller, distilled models at the edge, and (3) a decoupled producer-consumer architecture that separates specialized data source layer from model inference, enabling flexible deployment across diverse computing environments. Using this integrated approach, our pipeline achieves accuracy comparable to traditional methods tuned by human experts while eliminating the need for experiment-specific parameter tuning. Although current throughput requires optimization for MHz facilities, our system's scalable architecture and demonstrated model compression capabilities provide a foundation for future high-throughput XFEL deployments.

Wang, Cong↗

An automated liquid jet for fluorescence dosimetry and microsecond radiolytic labeling of proteins

X-ray radiolytic labeling uses broadband X-rays for in situ hydroxyl radical labeling to map protein interactions and conformation. High flux density beams are essential to overcome radical scavengers. However, conventional sample delivery environments, such as capillary flow, limit the use of a fully unattenuated focused broadband beam. An alternative is to use a liquid jet, and we have previously demonstrated that use of this form of sample delivery can increase labeling by tenfold at an unfocused X-ray source. Here we report the first use of a liquid jet for automated inline quantitative fluorescence dosage characterization and sample exposure at a high flux density microfocused synchrotron beamline. Our approach enables exposure times in single-digit microseconds while retaining a high level of side-chain labeling. This development significantly boosts the method’s overall effectiveness and efficiency, generates high-quality data, and opens up the arena for high throughput and ultrafast time-resolved in situ hydroxyl radical labeling.

59 BASIC BIOLOGICAL SCIENCES↗

Automated 3D cytoplasm segmentation in soft X-ray tomography

Cells’ structure is key to understanding cellular function, diagnostics, and therapy development. Soft X-ray tomography (SXT) is a unique tool to image cellular structure without fixation or labeling at high spatial resolution and throughput. Fast acquisition times increase demand for accelerated image analysis, like segmentation. Currently, segmenting cellular structures is done manually and is a major bottleneck in the SXT data analysis. This paper introduces ACSeg, an automated 3D cytoplasm segmentation model. ACSeg is generated using semi-automated labels and 3D U-Net and is trained on 43 SXT tomograms of immune T cells, rapidly converging to high-accuracy segmentation, therefore reducing time and labor. Furthermore, adding only 6 SXT tomograms of other cell types diversifies the model, showing potential for optimal experimental design. ACSeg successfully segmented unseen tomograms and is published on Biomedisa, enabling high-throughput analysis of cell volume and structure of cytoplasm in diverse cell types.

59 BASIC BIOLOGICAL SCIENCES↗

Auto-Curation of Seismic Event Data for Signal Denoising

Denoising contaminated seismic signals for later processing is a fundamental problem in seismic signals analysis. Neural network approaches have shown success denoising local signals when trained on short-time Fourier transform spectrograms. One challenge of this approach is the onerous process of hand-labeling event signals for training. By leveraging the SCALODEEP seismic event detector, we develop an automated set of techniques for labeling event data. Despite region specific challenges, training the neural network denoiser on machine curated events shows comparable performance to the neural network trained on hand curated events. We showcase our technique with two experiments, one using Utah regional data and one using regional data from the Korean peninsula.

58 GEOSCIENCES↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Detecting technological maturity from bibliometric patterns

We report the capability to identify emergent technologies based upon easily accessed open-source indicators, such as publications, is important for decision-makers in industry and government. The scientific contribution of this work is the proposition of a machine learning approach to the detection of the maturity of emerging technologies based on publication counts. Time-series of publication counts have universal features that distinguish emerging and growing technologies. We train an artificial neural network classifier, a supervised machine learning algorithm, upon these features to predict the maturity (emergent vs. growth) of an arbitrary technology. With a training set comprised of 22 technologies we obtain a classification accuracy ranging from 58.3% to 100% with an average accuracy of 84.6% for six test technologies. To enhance classifier performance, we augmented the training corpus with synthetic time-series technology life cycle curves, formed by calculating weighted averages of curves in the original training set. Training the classifier on the synthetic data set resulted in improved accuracy, ranging from 83.3% to 100% with an average accuracy of 90.4% for the test technologies. The performance of our classifier exceeds that of competing machine learning approaches in the literature, which report an average classification accuracy of only 85.7% at maximum. Moreover, in contrast to current methods our approach does not require subject matter expertise to generate training labels, and it can be automated and scaled.

97 MATHEMATICS AND COMPUTING↗

Automated detection of photovoltaic cleaning events: A performance comparison of techniques as applied to a broad set of labeled photovoltaic data sets

Extracting accurate soiling loss information from photovoltaic (PV) production data first requires segmenting the time series data per natural or manually occurring cleaning events. Maintenance logs are often incomplete, rain data are often unavailable, and the debate on rain thresholds for cleaning and dew or wind cleanings is still ongoing. The present work aims to overtake these issues by improving automated methods to detect these cleaning events and therefore improve extraction of soiling loss information. Time series power production data from 22 PV inverters were labeled for natural or manually occurring cleaning events. The data sets were carefully selected to include varying degrees of soiling, cleaning events, and noise. Several algorithms, including filtering logic and change point detection, were examined for efficacy at detecting the labeled cleanings. All the methods introduced except for changepoint detection showed significant improvement at detecting the labeled cleaning events per the mean F 1 score. Furthermore, the highest performing cleaning detection algorithm achieved an absolute increase in the mean F 1 score of 43% over the default version of the RdTools stochastic rate and recovery (SRR) algorithm. The highest performing algorithm included irradiance filtering and a cleaning detection threshold, adjusted based on the 40-day centered rolling median of the absolute day-to-day deviations in the daily performance index (PI). Furthermore, these improvements are promising as cleaning detection is an essential step in the automated analysis of PV soiling.

14 SOLAR ENERGY↗