Search NASA⌕ Search

SEARCH · Search NASA

Results for “label shift”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

High dimensional binary classification under label shift: phase transition and regularization

Label Shift has been widely believed to be harmful to the generalization performance of machine learning models. Researchers have proposed many approaches to mitigate the impact of the label shift, e.g., balancing the training data. However, these methods often consider the underparametrized regime, where the sample size is much larger than the data dimension. The research under the overparametrized regime is very limited. Here, to bridge this gap, we propose a new asymptotic analysis of the Fisher Linear Discriminant classifier for binary classification with label shift. Specifically, we prove that there exists a phase transition phenomenon: Under certain overparametrized regime, the classifier trained using imbalanced data outperforms the counterpart with reduced balanced data. Moreover, we investigate the impact of regularization to the label shift: The aforementioned phase transition vanishes as the regularization becomes strong.

binary classification↗

Assessing Membership Inference Attacks under Distribution Shifts

Membership inference attacks (MIAs) exploit machine learning models to infer whether a data point was in the training set, posing significant privacy risks even with limited black-box access. These attacks rely on the attacker approximating the target model’s training distribution, yet the impact of distribution shifts between target and shadow models on MIA success remains underexplored. We systematically evaluate five types of distribution shifts —-cutout, jitter, Gaussian noise, label shift, and attribute shift —- at varying intensities. Our results reveal that these shifts affect MIA effectiveness in nuanced ways, with some reducing attack success while others exacerbate vulnerabilities, and the same shift can have opposite effects depending on the type of MIA. This highlights the complex interplay between distributional differences and attack performance, offering critical insights for improving model defenses against MIAs.

Shi, Yichuan [Massachusetts Institute of Technolog↗

Boundary-Aware Adversarial Learning Domain Adaption and Active Learning for Cross-Sensor Building Extraction

The use of convolutional neural networks (CNNs) for building extraction from remote sensing images has been widely studied and many public datasets have been made available for accelerating development of these CNN models. Yet adapting pretrained models at scale in real-world scenarios remains a challenging task. The main barrier is that certain new labels are still needed to compensate for domain shifting between the labeled data and new images that potentially cover new geographic locations or that are from a different sensor. In this article, we propose to add informatively labeled samples from a new image pool under the paradigm of active learning. To select the most useful samples based on model uncertainty, we first tackle the problem of uncalibrated uncertainty estimation due to distribution shifting by adapting feature extractors with boundary-based adversarial learning. Calibrated uncertainty is used as the query criterion in the active learning process, where the most uncertain samples are selected for annotation and included for model retraining. The proposed workflow was tested with three data pairs in which each workflow represents a scenario often encountered in real-world applications, including adapting pretrained models to new images collected with different sensors or to new geographic areas where appearances and types of buildings are very different. Compared to several baselines, including random sampling, temperature scaling (a well-known uncertainty calibration technique), different query strategies, and active domain adaptation methods, the proposed workflow shows that strategically querying a smaller set of samples for labeling achieves comparable or better building extraction performance. The proposed method reduces the number of labeled samples required to achieve sufficient model accuracy, thus significantly reducing hundreds of person-hours for labeled data creation. In addition, we include a few considerations when deploying this workflow in a GPU cluster that can be easily adapted to achieve operational building extraction model retraining.

97 MATHEMATICS AND COMPUTING↗

Domain Shift Analysis in Chest Radiographs Classification in a Veterans Healthcare Administration Population

This study aims to assess the impact of domain shift on chest X-ray classification accuracy and to analyze the influence of ground truth label quality and demographic factors such as age group, sex, and study year. We used a DenseNet121 model pre-trained MIMIC-CXR dataset for deep learning-based multi-label classification using ground truth labels from radiology reports extracted using the CheXpert and CheXbert Labeler. We compared the performance of the 14 chest X-ray labels on the MIMIC-CXR and Veterans Healthcare Administration chest X-ray dataset (VA-CXR). The validation of ground truth and the assessment of multi-label classification performance across various NLP extraction tools revealed that the VA-CXR dataset exhibited lower disagreement rates than the MIMIC-CXR datasets. Additionally, there were notable differences in AUC scores between models utilizing CheXpert and CheXbert. When evaluating multi-label classification performance across different datasets, minimal domain shift was observed in the unseen VA dataset, except for the label “Enlarged Cardiomediastinum.” The subgroup with the most significant variations in multi-label classification performance was study year. These findings underscore the importance of considering domain shift in chest X-ray classification tasks, paying particular attention to the temporality of the exam. Our study reveals the significant impact of domain shift and demographic factors on chest X-ray classification, emphasizing the need for improved transfer learning and robust model development. Addressing these challenges is crucial for advancing medical imaging research and improving patient care.

chest X-ray image classification↗

RINO: Renormalization Group Invariance with No Labels

A common challenge with supervised machine learning (ML) in high energy physics (HEP) is the reliance on simulations for labeled data, which can often mismodel the underlying collision or detector response. To help mitigate this problem of domain shift, we propose RINO (Renormalization Group Invariance with No Labels), a self-supervised learning approach that can instead pretrain models directly on collision data, learning embeddings invariant to renormalization group flow scales. In this work, we pretrain a transformer-based model on jets originating from quantum chromodynamic (QCD) interactions from the JetClass dataset, emulating real QCD-dominated experimental data, and then finetune on the JetNet dataset -- emulating simulations -- for the task of identifying jets originating from top quark decays. RINO demonstrates improved generalization from the JetNet training data to JetClass data compared to supervised training on JetNet from scratch, demonstrating the potential for RINO pretraining on real collision data followed by fine-tuning on small, high-quality MC datasets, to improve the robustness of ML models in HEP.

Hao, Zichun [Caltech] (ORCID:0000000256244907)↗

Nucleotide-Protectable Labeling of Sulfhydryl Groups in Subunit I of the ATPhase from Halobacterium Saccharovorum

A membrane-bound ATPase from the archaebacterium Halobacterium saccharovorum is inhibited by N-ethyl-maleimide in a nucleotide-protectable manner. When the enzyme was incubated with N-[C-14]jethylmaleimide, the bulk of radioactivity was as- sociated with the 87,000-Da subunit (subunit 1). ATP, ADP, or AMP reduced incorporation of the inhibitor. No charge shift of subunit I was detected following labeling with N-ethylmaleimide, indicating an electroneutral reaction. The results are consistent with the selective modification of sulfhydryl groups in subunit I at or near the catalytic site and are further evidence of a resemblance between this archaebacterial ATPase and the vacuolar-type ATPases.

Sulzner, Michael↗

Culturability as an indicator of succession in microbial communities

Successional theory predicts that opportunistic species with high investment of energy in reproduction and wide niche width will be replaced by equilibrium species with relatively higher investment of energy in maintenance and narrower niche width as communities develop. Since the ability to rapidly grow into a detectable colony on nonselective agar medium could be considered as characteristic of opportunistic types of bacteria, the percentage of culturable cells may be an indicator of successional state in microbial communities. The ratios of culturable cells (colony forming units on R2A agar) to total cells (acridine orange direct microscopic counts) and culturable cells to active cells (reduction of 5-cyano-2,3-ditolyl tetrazolium chloride) were measured over time in two types of laboratory microcosms (the rhizosphere of hydroponically grown wheat and aerobic, continuously stirred tank reactors containing plant biomass) to determine the effectiveness of culturabilty as an index of successional state. The culturable cell:total cell ratio in the rhizosphere decreased from approximately 0.25 to less than 0.05 during the first 30-50 days of plant growth, and from 0.65 to 0.14 during the first 7 days of operation of the bioreactor. The culturable cell:active cell ratio followed similar trends, but the values were consistently greater than the culturable cell:total cell ratio, and even exceeded I in early samples. Follow-up studies used a cultivation-independent method, terminal restriction fragment length polymorphisms (TRFLP) from whole community DNA, to assess community structure. The number of TRFLP peaks increased with time, while the number of culturable types did not, indicating that the general decrease in culturability is associated with a shift in community structure. The ratio of respired to assimilated C-14-labeled amino acids increased with the age of rhizosphere communities, supporting the hypothesis that a shift in resource allocation from growth to maintenance occurs with time. Results from this work indicate that the percentage of culturable cells may be a useful method for assessing the successional state of microbial communities.

NASA Center KSC↗

Quantum dots as strain- and metabolism-specific microbiological labels

Biologically conjugated quantum dots (QDs) have shown great promise as multiwavelength fluorescent labels for on-chip bioassays and eukaryotic cells. However, use of these photoluminescent nanocrystals in bacteria has not previously been reported, and their large size (3 to 10 nm) makes it unclear whether they inhibit bacterial recognition of attached molecules and whether they are able to pass through bacterial cell walls. Here we describe the use of conjugated CdSe QDs for strain- and metabolism-specific microbial labeling in a wide variety of bacteria and fungi, and our analysis was geared toward using receptors for a conjugated biomolecule that are present and active on the organism's surface. While cell surface molecules, such as glycoproteins, make excellent targets for conjugated QDs, internal labeling is inconsistent and leads to large spectral shifts compared with the original fluorescence, suggesting that there is breakup or dissolution of the QDs. Transmission electron microscopy of whole mounts and thin sections confirmed that bacteria are able to extract Cd and Se from QDs in a fashion dependent upon the QD surface conjugate.

NASA Center JPL↗

H, C, N, and O Isotopic Substitution Studies of the 2165 cm (4.62 micron) "XCN" Feature Produced by UV Photolysis of Mixed Molecular Ices

To better understand the chemical species that gives rise to the 2165/cm (4.62 micron) "XCN" absorption feature seen towards embedded protostars such as W33A, we have performed laboratory studies using deuterium (H-2) isotopic labeling. We report the observation of a small but significant deuterium isotope shift for the "XCN" peak which demonstrates that the atomic motion(s) causing the "XCN" band in the laboratory samples must involve hydrogen. We also report the results of C-13, N-15 and O-18 labeling experiments that are consistent with previously reported values.

Bernstein, Max P.↗

Leveraging Automated Fiber Placement Computer Aided Process Planning Framework for Defect Validation and Dynamic Layup Strategies

Process planning represents an essential stage of the Automated Fiber Placement (AFP) workflow. It develops useful and efficient machine processes based upon the working material, composite design, and manufacturing resources. The current state of process planning requires a high degree of interaction from the process planner and could greatly benefit from increased automation. Therefore, a list of key steps and functions are created to identify the more difficult and time-consuming phases of process planning. Additionally, a set of metrics must exist by which to evaluate the effectiveness of the manufactured laminate from the machine code created during the Process Planning stage. Layup strategies, in addition to dog ears, stagger shifts, steering constraints, and starting points, represented the group of functions labeled as process optimization and ranked the highest in terms of priority for automation. The laminates resulting from the selected parameters are evaluated through the occurrences of principal defect metrics such as fiber gaps, overlaps, angle deviation and steering violations. This document presents an automated software solution to the layup strategy and starting point selection phase of process planning. A series of ply scenarios are generated with variations of these ply parameters and evaluated according to a set of metrics entered by the Process Planner. These metrics are generated through use of the Analytical Hierarchy Process (AHP), where relative importance between each of the fiber features are defined. The ply scenarios are selected which reduce the overall fiber feature scores based on the defects the Process Planner wishes to minimize.

HiCAM↗

A Thorough Characterization of the Tellurocyanate Anion

Tellurocyanate, [TeCN] − , is the heaviest group 16 congener of the cyanate anion, [OCN] − . Due to the relative instability of the C─Te bond, tellurocyanate chemistry has seen only scarce attention. Here, we present the facile synthesis and thorough characterization of [K@crypt-222][TeCN]. The anion is essentially linear with interatomic distances C─N = 1.150(6)Å and C─Te = 2.051(4)Å, thus approximating a C≡N triple bond and for C─Te a bond order between 1 and 2. Fully 13 C and 15 N labeled [Te 13 C 15 N] − allowed for the extraction of chemical shifts and all possible coupling constants ( 13 C = 77.8 ppm, 15 N = 285.7 ppm, 125 Te = −566 ppm, 1 J 13C-15N = 8 Hz, 1 J 13C-125Te = 748 Hz, 2 J1 5N-125Te = 55 Hz), which were also determined independently by quantum chemical calculations. In the series [ChCN] − (Ch = O─Te), [TeCN] − shows the strongest spin-orbit coupling (SOC) induced heavy-atom effect on the light-atom shielding (SO-HALA-effect). In contrast, 15 N shifts are also well described without considering relativistic effects and/or SOC. Negative-ion photoelectron spectroscopy was used to extract the electron affinity (EA = 3.034 eV) and spin-orbit splitting (3807 cm −1 ) of [TeCN] • . These values continue the trends of falling EA and rising SOC in the series [ChCN] • .

Bonding analysis↗

Uncertainty-refined image segmentation under domain shift

Digital image segmentation is provided. The method comprises training a neural network for image segmentation with a labeled training dataset from a first domain, wherein a subset of nodes in the neural net are dropped out during training. The neural network receives image data from a second, different domain. A vector of N values that sum to 1 is calculated for each image element, wherein each value represents an image segmentation class. A label is assigned to each image element according to the class with the highest value in the vector. Multiple inferences are performed with active dropout layers for each image element, and an uncertainty value is generated for each image element. Uncertainty is resolved according to expected characteristics. The label of any image element with an uncertainty above a threshold is replaced with a new label corresponding to a segmentation class based on domain knowledge.

Martinez, Carianne↗

CCAAT/enhancer-binding protein delta activates insulin-like growth factor-I gene transcription in osteoblasts. Identification of a novel cyclic AMP signaling pathway in bone

Insulin-like growth factor-I (IGF-I) plays a key role in skeletal growth by stimulating bone cell replication and differentiation. We previously showed that prostaglandin E2 (PGE2) and other cAMP-activating agents enhanced IGF-I gene transcription in cultured primary rat osteoblasts through promoter 1, the major IGF-I promoter, and identified a short segment of the promoter, termed HS3D, that was essential for hormonal regulation of IGF-I gene expression. We now demonstrate that CCAAT/enhancer-binding protein (C/EBP) delta is a major component of a PGE2-stimulated DNA-protein complex involving HS3D and find that C/EBPdelta transactivates IGF-I promoter 1 through this site. Competition gel shift studies first indicated that a core C/EBP half-site (GCAAT) was required for binding of a labeled HS3D oligomer to osteoblast nuclear proteins. Southwestern blotting and UV-cross-linking studies showed that the HS3D probe recognized a approximately 35-kDa nuclear protein, and antibody supershift assays indicated that C/EBPdelta comprised most of the PGE2-activated gel-shifted complex. C/EBPdelta was detected by Western immunoblotting in osteoblast nuclear extracts after treatment of cells with PGE2. An HS3D oligonucleotide competed effectively with a high affinity C/EBP site from the rat albumin gene for binding to osteoblast nuclear proteins. Co-transfection of osteoblast cell cultures with a C/EBPdelta expression plasmid enhanced basal and PGE2-activated IGF-I promoter 1-luciferase activity but did not stimulate a reporter gene lacking an HS3D site. By contrast, an expression plasmid for the related protein, C/EBPbeta, did not alter basal IGF-I gene activity but did increase the response to PGE2. In osteoblasts and in COS-7 cells, C/EBPdelta, but not C/EBPbeta, transactivated a reporter gene containing four tandem copies of HS3D fused to a minimal promoter; neither transcription factor stimulated a gene with four copies of an HS3D mutant that was unable to bind osteoblast nuclear proteins. These results identify C/EBPdelta as a hormonally activated inducer of IGF-I gene transcription in osteoblasts and show that the HS3D element within IGF-I promoter 1 is a high affinity binding site for this protein.

NASA Discipline Musculoskeletal↗

Unraveling metabolism underpinning biomass composition shift in Scenedesmus obliquus under simulated outdoor conditions using 13 C-fluxomics

To render the resulting biomass more attractive and amenable for utilization as the basis for low-carbon intensity bioproducts, single-celled algae need to be biochemically and metabolically poised to assimilate and store the delivered carbon in the fastest and most efficient manner. Accelerating biochemical carbon storage, as primarily carbohydrates or lipids, is critical to achieve the high carbon capture potential that is assigned to algae. To guide strain optimization and engineering for maximizing carbon capture and storage, it is essential to elucidate the link between carbon metabolism and biomass composition. Most published metabolomics work in algae remains largely restricted to ideal and simplified environmental conditions in model organisms, thereby limiting their translation to outdoor implementation. In this work, we utilize 13 C isotopic labeling to characterize distinct intracellular metabolic fluxes before, during, and after nitrogen depletion-induced compositional shifts in Scenedesmus obliquus UTEX 393. The results indicate that a transition to carbohydrates is characterized by diverting flux to starch instead of replenishing the Calvin cycle for CO 2 fixation whereas the subsequent transition to lipids is fueled by NADPH produced by upregulating the phosphoenolpyruvate carboxylase (PEPC)–malic enzyme (ME) cycle flux. Our work highlights bottlenecks to carbohydrate- and lipid-rich biomass and can guide implementable strategies to control the fate of fixed carbon in S. obliquus.

09 BIOMASS FUELS↗

A VLSI decomposition of the deBruijn graph

The nth order deBruijn graph Bn is the state diagram for an n-stage binary shift register. It is a directed graph with 2 to the n vertices, each labeled with an n-bit binary string, and 2 to the n+1 edges, each labeled with an (n+1)-bit binary string. It is shown that Bn can be built by appropriately connecting together with extra edges many isomorphic copies of a fixed graph, which is called a building block for Bn. The efficiency of such a building block is refined as the fraction of the edges of Bn which are present in the copies of the building block. It is then shown that for any alpha less than 1, there exists a graph which is a building block for Bn of efficiency greater than alpha for all sufficiently large n. The results are illustrated by showing how a special hierarchical family of building blocks has been used to construct a very large Viterbi decoder which will be used on the Galileo mission.

Collins, Oliver↗

Hydra: Computer Vision for Online Data Quality Monitoring

Hydra is a system utilizing computer vision for near real-time data quality monitoring. Currently operational across all of Jefferson Lab’s experimental halls, it reduces the workload of shift takers by autonomously monitoring diagnostic plots during experiments. Hydra uses "off-the-shelf" supervised learning technologies and is supported by a comprehensive MySQL database. To simplify access, web apps have been developed to facilitate both labeling and monitoring of Hydra’s inferences. Hydra can connect with the alarm system and incorporates complete historical tracking, enabling it to identify issues that shift takers could miss. When issues are detected, a natural first question is: "Why does Hydra think there is a problem?" To answer, Hydra employs Gradient-weighted Class Activation Maps (GradCAM) to identify regions of the image that are important for the specific classification. This interpretive layer enhances transparency and trustworthiness, which is essential for integration with experiment workflows and operation. The Hydra system, results, and sociological considerations for deployment will be discussed.

Jeske, Torri↗

SIDDA: SInkhorn Dynamic Domain Adaptation for image classification with equivariant neural networks

Modern neural networks (NNs) often do not generalize well in the presence of a ‘covariate shift’; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels given the data remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more robust, domain-invariant features. Domain adaptation (DA) methods include a broad range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SInkhorn Dynamic Domain Adaptation (SIDDA), an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, real astronomical observations, and remote sensing data. These datasets exhibit covariate shifts due to noise, blurring, differences between telescopes, and variations in imaging wavelengths. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with symmetry-aware equivariant NNs (ENNs). We find that SIDDA consistently enhances the generalization capabilities of NNs, achieving up to a ${\approx}40\%$ improvement in classification accuracy on unlabeled target data, while also providing a more modest performance gain of $\lesssim 1\%$ on labeled source data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, if SIDDA achieves proper domain alignment, it also enhances model calibration on both source and target data, with the most significant gains in the unlabeled target domain—achieving over an order of magnitude improvement in the expected calibration error and Brier score. SIDDA’s versatility across various NN models and datasets, combined with its automated approach to domain alignment, has the potential to significantly advance multi-dataset studies by enabling the development of highly generalizable models.

79 ASTRONOMY AND ASTROPHYSICS↗

Opportunistic Short‐Term Water Uptake Dynamics by Subalpine Trees Observed via In Situ Water Isotope Measurements

Abstract Variations in tree water sources are important to understand in semi‐arid ecosystems because climatic shifts towards lower snowpack and increased drought affect water availability in subalpine forests of the western US. Here, we use daily in situ measurements of stable isotopes ( 2 H & 18 O) in soil and tree stem water, soil matric potential and sap flow to study tree water uptake dynamics. We instrumented three soil profiles down to 90 cm, as well as three aspen and engelmann spruce trees near Gothic, Colorado, in the East River watershed. We observed the fate of natural isotopic variations in rainfall, soil, and plants from June to October 2022, and in August 2023 we conducted a 2 H labeled irrigation experiment. Our observations showed that all studied aspen trees compensated for water scarcity in the shallow soil by shifting the dominant water source at 60(±20) cm to ⅔ of uptake from 90 cm within a few days of a dry period. Both species relied on snowmelt stored in the subsoil to sustain transpiration. Intense rainfall caused the plant water uptake to shift partially to top soil layers within 2 days. Spruce transpiration was lower and relied more on snowmelt, because rainfall infiltration was low in the spruce stand due to high canopy interception. Our findings highlight the important role of snowmelt stored in the deep soil layers for subalpine forest drought response and the dominant fate of monsoonal rainfall to become transpiration rather than recharging groundwater and streams in the Upper Colorado River. Plain Language Summary There is a need to understand how trees in mountainous regions respond to dry conditions that lead to water scarcity, because climate projections suggest that such conditions will become more frequent in the future. Here we present a novel data set of measurements of daily stable isotopes of water across soil profiles and in tree stems of aspen and spruce. Our data show that when the upper soil dried out, aspen trees shifted to using water from deeper layers (beneath 60 cm) to keep transpiring. For spruce trees the uptake pattern is less clear, but both types of trees mainly used snowmelt stored in the deeper soil layers to survive the dry summer. After heavy rain, aspen and spruce trees switched to using water from the top 20 cm of soil. However, for spruce, only some rain reached the soil because the dense tree canopy intercepted it, so spruce trees stayed more dependent on snowmelt and used less water overall. This study shows how important deep snowmelt water is for helping forests survive dry periods and suggests that most summer rain is quickly used by trees rather than replenishing streams and groundwater in the headwaters of the Colorado River. Key Points Tree water resources changed within a few days from snow dominated to higher share of rainfall as soils wetted up after a dry period Compensatory plant water uptake by aspen from the deep layer (90 cm), while uptake from soil depths that became drier (60 cm) declined Strong differences between water sources and availability beneath aspen and spruce, respectively

Sprenger, Matthias↗