Search NASA⌕ Search

SEARCH · Search NASA

Results for “data availability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Node-degree aware edge sampling mitigates inflated classification performance in biomedical random walk-based graph representation learning

Motivation: Graph representation learning is a family of related approaches that learn low-dimensional vector representations of nodes and other graph elements called embeddings. Embeddings approximate characteristics of the graph and can be used for a variety of machine-learning tasks such as novel edge prediction. For many biomedical applications, partial knowledge exists about positive edges that represent relationships between pairs of entities, but little to no knowledge is available about negative edges that represent the explicit lack of a relationship between two nodes. For this reason, classification procedures are forced to assume that the vast majority of unlabeled edges are negative. Existing approaches to sampling negative edges for training and evaluating classifiers do so by uniformly sampling pairs of nodes. Results: We show here that this sampling strategy typically leads to sets of positive and negative examples with imbalanced node degree distributions. Using representative heterogeneous biomedical knowledge graph and random walk-based graph machine learning, we show that this strategy substantially impacts classification performance. If users of graph machine-learning models apply the models to prioritize examples that are drawn from approximately the same distribution as the positive examples are, then performance of models as estimated in the validation phase may be artificially inflated. We present a degree-aware node sampling approach that mitigates this effect and is simple to implement. Availability and implementation: Our code and data are publicly available at https://github.com/monarch-initiative/negativeExampleSelection.

59 BASIC BIOLOGICAL SCIENCES↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗

Centralized Interactive Phenomics Resource: an integrated online phenomics knowledgebase for health data users

Development of clinical phenotypes from electronic health records (EHRs) can be resource intensive. Several phenotype libraries have been created to facilitate reuse of definitions. However, these platforms vary in target audience and utility. Here, we describe the development of the Centralized Interactive Phenomics Resource (CIPHER) knowledgebase, a comprehensive public-facing phenotype library, which aims to facilitate clinical and health services research. The platform was designed to collect and catalog EHR-based computable phenotype algorithms from any healthcare system, scale metadata management, facilitate phenotype discovery, and allow for integration of tools and user workflows. Phenomics experts were engaged in the development and testing of the site. The knowledgebase stores phenotype metadata using the CIPHER standard, and definitions are accessible through complex searching. Phenotypes are contributed to the knowledgebase via webform, allowing metadata validation. Data visualization tools linking to the knowledgebase enhance user interaction with content and accelerate phenotype development. The CIPHER knowledgebase was developed in the largest healthcare system in the United States and piloted with external partners. The design of the CIPHER website supports a variety of front-end tools and features to facilitate phenotype development and reuse. Health data users are encouraged to contribute their algorithms to the knowledgebase for wider dissemination to the research community, and to use the platform as a springboard for phenotyping. CIPHER is a public resource for all health data users available at https://phenomics.va.ornl.gov/ which facilitates phenotype reuse, development, and dissemination of phenotyping knowledge.

60 APPLIED LIFE SCIENCES↗

Strong nebular emissions associated with Mg ii absorptions detected in the SDSS spectra of background quasars

ABSTRACT We present long-slit spectroscopic observations of 40 Galaxy On Top of Quasars (GOTOQs) at ${0.37 \leqslant z \leqslant 1.01}$ using the South African Large Telescope. Using this and available photometric data, we measure the impact parameters of the foreground galaxies to be in the range of 3–16 kpc with a median value of 8.6 kpc. This is the largest sample of galaxies producing Mg ii absorption at such low impact parameters. These quasar–galaxy pairs are ideal for probing the disc–halo interface. At such impact parameters, we do not find any anticorrelation between rest equivalent width (REW) of Ca ii, Mn ii, Fe ii, Mg ii, and Mg i absorption and impact parameter. These sight lines are typically redder than those of strong Mg ii absorbers, with the colour excess, E(B − V) for our sample ranging from −0.191 to 0.422, with a median value of 0.058. In the E(B − V) versus W3935 plane, GOTOQs occupy the same region as Ca ii absorbers. For a given E(B − V), we find larger W3935 than what has been found in the Milky Way, probably due to a smaller dust-to-gas ratio in GOTOQs. Galaxy parameters could be measured for twelve cases, and their properties seem to follow the trends found for strong Mg ii absorbers. Measuring the host galaxy properties for the full sample using HST photometry or AO-assisted ground-based imaging is important to gain insights into the relationship between the stellar mass of galaxies and the metal line REW distributions at low impact parameters.

Astronomy & Astrophysics↗

Glauber-theory analysis of nuclear reactions on a 12 C target with variational Monte Carlo wave functions

The application of Glauber theory has been playing an increasingly important role with the study of unstable or exotic nuclei. Its adaptation to medium and high-energy nucleus-nucleus collisions is severely limited because one has to evaluate the matrix elements of multiple-scattering operators. The extraction of physical observables has been done using ‘approximate’ Glauber theory whose validity is hard to evaluate. Here, we perform a full calculation of the matrix elements using Monte Carlo integration and analyze the elastic differential cross sections and the total reaction cross sections for p+¹²C, ⁴,⁶He+¹²C, and ¹²C+¹²C collisions. We use the variational Monte Carlo wave functions for ⁴,⁶He and ¹²C obtained by using realistic two- and three-nucleon potentials. We demonstrate the performance of the Glauber-theory calculations by comparing with available experimental data. We further discuss the accuracy of the conventional approximate methods in the light of the cumulant expansion for Glauber’s phase-shift function.

Horiuchi, W. [Osaka Metropolitan University (Japan↗

Microscopic optical potentials for medium-mass isotopes derived at the first order of Watson multiple-scattering theory

We perform a first-principles calculation of optical potentials for nucleon elastic scattering off medium-mass isotopes. Fully based on a saturating chiral Hamiltonian, the optical potentials are derived by folding nuclear density distributions computed with self-consistent Green's function theory with a nucleon-nucleon t matrix computed with a consistent chiral interaction. The dependence on the folding interaction as well as the convergence of the target densities are investigated. Numerical results are presented and discussed for differential cross sections and analyzing powers, with focus on elastic proton scattering off calcium and nickel isotopes. Our optical potentials generally show a remarkable agreement with the available experimental data for laboratory energies in the range 65–200 MeV. We study the evolution of the scattering observables with increasing proton-neutron asymmetry by computing theoretical predictions of the cross section and analyzing power over the calcium and nickel isotopic chains. Published by the American Physical Society 2024

Physics↗

New constraint on the Np 237 ( n , γ ) Np 238 integral cross section using the Godiva-IV critical assembly

Accurate knowledge of the 237 Np(n, γ) 238 Np cross section at fast neutron energies is important for applied nuclear science. The presently available experimental data has large disagreements in the fast neutron region. Perform a model-independent measurement of the 237 Np(n, γ) 238 Np integral cross section using a well characterized fast neutron source and compare the result with previous measurements and current nuclear data evaluations. Provide an integral measurement that can be used as a benchmark for current evaluations. Multiple samples of 237 Np were irradiated in the Godiva-IV critical assembly. Following the irradiation, the samples placed in a γ-ray counting setup and the γ-rays emitted from the decay of 238 Np were measured over a time period of approximately 7 days. Multiple γ-ray decay branches of 238 Np were observed. The observed activity of 238 Np was used to calculate the amount of 238 Np produced during the irradiation via the 237 Np(n, γ) 238 Np reaction and an integral cross section of 342(11) mb was measured for the Godiva-IV neutron spectrum. Further, the 238 Np half-life has been measured with a result of 50.31(5) hours. The 237 Np(n, γ) 238 Np integral cross section measured in this work is in agreement with overlapping 1σ error bands to ENDF/B-VIII.0. However, the measured value is 3σ away from the calculated integral cross section using JENDL-5. This measurement offers a reliable benchmark for future 237 Np(n, γ) 238 Np cross section evaluations.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Ab initio investigation of the Li 7 ( p , e + e - ) Be 8 process and the X17 boson

Observations of anomalies in the electron-positron angular correlations in high-energy decays in 4 He, 8 Be, and 12 C have been reported recently by the ATOMKI collaboration. These could be explained by the creation and subsequent decay of a new boson with a mass of ≈ 17MeV. Theoretical understanding of pair creation in the proton capture reactions used in these experiments is important for the interpretation of the anomalies. We apply the ab initio no-core shell model with continuum (NCSMC) to the proton capture on 7 Li. The NCSMC describes both bound and unbound states in light nuclei in a unified way with chiral two- and three-nucleon interactions as the only input. We investigate the structure of 8 Be, the p+ 7 Li elastic scattering, the 7 Li(p,y)⁢ 8 Be cross section, and the internal pair creation 7 Li ⁢(p,e + ⁢e - ) 8 Be. Here we discuss the impact of a proper treatment of the initial scattering state on the electron-positron angular correlation spectrum and compare our results to available ATOMKI data sets. Finally, we calculate 7 Li ⁢(p,X)⁢ 8 Be cross sections for several proposed models of the hypothetical X17 particle.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Jet suppression and azimuthal anisotropy from RHIC to LHC

Azimuthal anisotropies of high- p T particles produced in heavy-ion collisions are understood as an effect of a geometrical selection bias. Particles oriented in the direction in which the QCD medium formed in these collisions is shorter suffer less energy loss, and thus, are over-represented in the final ensemble compared to those oriented in the direction in which the medium is longer. In this work we present the first semianalytical predictions, including propagation through a realistic, hydrodynamical background, of the elliptic azimuthal anisotropy for jets, obtaining a quantitative agreement with available experimental data as a function of the jet p T , its cone size R , and the collisions centrality. Jets are multipartonic, extended objects and their energy loss is sensitive to substructure fluctuations. This sensitivity is determined by the physics of color coherence that relates to the ability of the medium to resolve those partonic fluctuations. Specifically, color dipoles with an angular separation smaller than a critical angle, θ c , are not resolved by the medium and they effectively act as a coherent source of energy loss. We find that elliptic jet azimuthal anisotropy has a specially strong dependence on coherence physics due to the marked length dependence of θ c . By combining our predictions for the collision systems and center-of-mass energies studied at RHIC and the LHC, covering a wide range of typical values of θ c , we show that the relative size of elliptic jet azimuthal anisotropies for jets with different cone sizes R follows a universal trend that indicates a transition from a coherent regime of jet quenching to a decoherent regime. These results suggest a way forward to reveal the role played by the physics of jet color decoherence in probing deconfined QCD matter. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Quantum Monte Carlo Calculations of Magnetic Form Factors in Light Nuclei

Here, we present Quantum Monte Carlo calculations of magnetic form factors in A = 6-10 nuclei, based on Norfolk two- and three-nucleon interactions, and associated one- and two-body electromagnetic currents. Agreement with the available experimental data for 6 Li, 7 Li, 9 Be and 10 B up to values of momentum transfer q ~ 3 fm -1 is achieved when two-nucleon currents are accounted for. We present a set of predictions for the magnetic form factors of 7 Be, 8 Li, 9 Li, and 9 C. In these systems, two body currents account for ~ 40-60% of the total magnetic strength. Measurements in any of these radioactive systems would provide valuable insights on the nuclear magnetic structure emerging from the underlying many-nucleon dynamics. A particularly interesting case is that of 7 Be, as it would enable investigations of the magnetic structure of mirror nuclei

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Multiscale Physics of Atomic Nuclei from First Principles

Atomic nuclei exhibit multiple energy scales ranging from hundreds of MeV in binding energies to fractions of an MeV for low-lying collective excitations. As the limits of nuclear binding are approached near the neutron and proton drip lines, traditional shell structure starts to melt with an onset of deformation and an emergence of coexisting shapes. It is a long-standing challenge to describe this multiscale physics starting from nuclear forces with roots in quantum chromodynamics. Here, we achieve this within a unified and nonperturbative quantum many-body framework that captures both short- and long-range correlations starting from modern nucleon-nucleon and three-nucleon forces from chiral effective field theory. The short-range (dynamic) correlations which account for the bulk of the binding energy are included within a symmetry-breaking framework, while long-range (static) correlations (and fine details about the collective structure) are included by employing symmetry projection techniques. Our calculations accurately reproduce—within theoretical error bars—available experimental data for low-lying collective states and the electromagnetic quadrupole transitions in 20−30 Ne. In addition, we reveal coexisting spherical and deformed shapes in 30 Ne, which indicates the breakdown of the magic neutron number 𝑁 = 20 as the key nucleus 28 O is approached, and we predict that the drip line nuclei 32,34 Ne are strongly deformed and collective. By developing reduced-order models for symmetry-projected states, we perform a global sensitivity analysis and find that the subleading singlet 𝑆-wave contact and a pion-nucleon coupling strongly impact nuclear deformation in chiral effective field theory. The techniques developed in this work clarify how microscopic nuclear forces generate the multiscale physics of nuclei spanning collective phenomena as well as short-range correlations and allow one to capture emergent and dynamical phenomena in finite fermion systems such as atom clusters, molecules, and atomic nuclei.

74 ATOMIC AND MOLECULAR PHYSICS↗

Observing the effects of numbers of valence nucleons on 0$^{+}_{𝑔⁡𝑠}$ → 2$^{+}_{1}$ transitions in deformed nuclei by comparing proton and neutron transition matrix elements

We examined the ratios of neutron and proton transition matrix elements, 𝑀 𝑛 /𝑀 𝑝 , for the 0$^{+}_{𝑔⁡𝑠}$ → 2$^{+}_{1}$ transitions in 48 even-even stable nuclei with 𝑁 > 20 for which electromagnetic matrix elements were compiled by Pritychenko et al. and for which high-quality inelastic proton-scattering data were available. Several deformed rare-earth nuclei have (𝑀 𝑛 /𝑀 𝑝 )/(𝑁/𝑍) values significantly below 1.0, which is not consistent with a simple liquid-drop picture. However, this phenomenon can be explained using a schematic picture in which 𝑀 𝑝 reaches a maximum at proton midshell (𝑍 = 66) and 𝑀𝑛 reaches its maximum at neutron midshell (𝑁 = 104). Several midmass vibrational nuclei have 𝑀 𝑛 /𝑀 𝑝 values significantly below 𝑁/𝑍, which is not consistent with the expectation that 𝑀 𝑛 /𝑀 𝑝 = 𝑁/𝑍 in such nuclei. As a result, a shell-model investigation of these observations might yield insights about this behavior.

Collective levels↗

Fresh look at the nuclear transparency using the generalized parton distributions

Color transparency (CT) is a fundamental phenomenon in QCD in which hadrons produced in high-energy exclusive processes traverse nuclear matter with minimal interactions. Nuclear transparency, which quantifies this attenuation suppression, is a quantity with high sensitivity to CT effects and provides critical insights into QCD dynamics in nuclear environments. In this study, we revisit nuclear transparency using the framework of generalized parton distributions (GPDs). By constructing nuclear GPDs (nGPDs) through the incorporation of nuclear parton distribution functions, we calculate the nuclear transparency 𝑇⁡(𝑄 2 ) for the carbon nucleus as a function of momentum transfer 𝑄 2 considering various definitions and compare the results obtained with available experimental data. Our finding highlights the importance of choosing a physically motivated definition of nuclear transparency. Moreover, we emphasize that a more reliable determination of nGPDs requires a dedicated global analysis incorporating nuclear data. Such an approach is essential for improving the theoretical understanding of CT and for achieving consistency with experimental observations in the high-𝑄 2 regime.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Temporal Forecasting of Distributed Temperature Sensing in a Thermal Hydraulic System With Machine Learning and Statistical Models

We benchmark performance of long-short term memory (LSTM) network machine learning model and autoregressive integrated moving average (ARIMA) statistical model in temporal forecasting of distributed temperature sensing (DTS). Data in this study consists of fluid temperature transient measured with two co-located Rayleigh scattering fiber optic sensors (FOS) in a forced convection mixing zone of a thermal tee. We treat each gauge of a FOS as an independent temperature sensor. We first study prediction of DTS time series using Vanilla LSTM and ARIMA models trained on prior history of the same FOS that is used for testing. The results yield maximum absolute percentage error (MaxAPE) and root mean squared percentage error (RMSPE) of 1.58% and 0.06% for ARIMA, and 3.14% and 0.44% for LSTM, respectively. Next, we investigate zero-shot forecasting (ZSF) with LSTM and ARIMA trained on history of the co-located FOS only, which is advantageous when limited training data is available. The ZSF MaxAPE and RMSPE values for ARIMA are comparable to those of the Vanilla use case, while the error values for LSTM increase. We show that in ZSF, performance of LSTM network can be improved by training on most correlated gauges between the two FOS, which are identified by calculating the Pearson correlation coefficient. The improved ZSF MaxAPE and RMSPE for LSTM are 4.4% and 0.33%, respectively. Performance of ZSF LSTM can be further enhanced through transfer learning (TL), where LSTM is re-trained on a subset of the FOS that is the target of forecasting. We show that LSTM pre-trained on correlated dataset and re-trained on 30% of testing target dataset achieves MaxAPE and RMSPE values of 2.32% and 0.28%, respectively.

ARIMA↗

The Use of Machine Learning Models for Predicting the Dielectric Strength of Gases

Technological advancements in high voltage systems have pushed sulfur hexafluoride (SF6) to its operational limits. Furthermore, this gas has other drawbacks including a high liquefaction temperature and a high global warming potential. Therefore, there has been an urgent need to find alternative gases with high dielectric strength (DS). In this work, density functional theory (DFT) is used to calculate molecular descriptors that are fed into an artificial neural network (ANN) and a random forest (RF). These machine learning (ML) models are then used to predict the DS for hundreds of molecules. A finite element model (FEM) is also used to calculate the electric field profile of multiple simple electrode geometries as the applied voltage to the system is increased. Results indicate that the random forest model has better generalization to unseen data than the neural network. The highest DS value predicted by the RF was 2.16 relative to the experimental DS of SF6. The results also demonstrate how choosing a gas with a higher DS and a geometry with minimal edges and corners can significantly increase the operating voltage of an electrical system. Due to its superior generalization, the RF represents the most promising path toward an accurate DS predictor once sufficient experimental data are available.

Mileski, Matthew [AFIT]↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Bringing Different Views Together: A Hybrid Cooperative Perception Framework for Connected Autonomous Vehicles

Cooperative perception will be essential for connected autonomous vehicles to enhance object recognition and optimize path planning by sending data information about the surrounding environment. However, an inherent challenge in existing systems is the high bandwidth cost of transmitting information in real-time, which restricts cooperative perception’s practicality. Here, this work presents a hybrid cooperative perception fusion framework aimed at mitigating this issue by optimizing data transmission according to available bandwidth or through data reduction techniques. Our methods ensure that vehicles can rapidly transmit high-confidence data without overwhelming the network. Experimental results indicate that our methodology substantially diminishes data transmission sizes while maintaining object detection accuracy. For cooperative perception in autonomous vehicle systems, our approach provides a scalable and effective way to get past the bandwidth barrier.

Carrillo, Dominic [Univ. of North Texas, Denton, T↗