Search NASA⌕ Search

SEARCH · Search NASA

Results for “knowledge extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Rapid Adaptation of Chemical Named Entity Recognition Using Few-Shot Learning and LLM Distillation

Named entity recognition (NER) has been widely used in chemical text mining for the automatic identification and extraction of chemical entities. However, existing chemical NER systems primarily focus on scenarios with abundant training data, requiring significant human effort on annotations. This poses challenges for applications in the chemical field, such as catalysis, where many advancements have traditionally relied on trial-and-error investigations and incremental adjustment of variables. This hinders catalysis science and technology progress in addressing emerging energy and environmental crises. In this work, we propose a few-shot NER model that can quickly adapt to extract new types of chemical entities by using only a limited number of annotated examples. Our model employs a metric-learning approach to transfer entity similarity knowledge from high-resource chemical domains (with abundant annotations) to enable effective entity recognition in low-resource specialized domains (limited annotation). We validate the effectiveness of our model on a few-shot chemical NER benchmark built based on six existing chemical NER data sets. Experiments show that the proposed few-shot NER model can achieve reasonable performance with only 5 examples per entity type and shows consistent improvement as the number of examples increases. Furthermore, we demonstrate how the proposed model can be trained with large language model (LLM) annotated data, opening a new pathway for rapid adaptation of NER systems. Furthermore, our approach leverages the knowledge broadness of large language models for chemistry while distilling this knowledge into a lightweight model suitable for efficient and in-house use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Assessing the impact of uniform rotation on the structure of neutron stars

Driven by recent laboratory experiments and astronomical observations, significant advances have deepened our understanding of neutron-star physics. NICER's Pulse Profile Modeling has refined our knowledge of neutron star masses and radii, while gravitational-wave detections have revealed key insights into the structure of neutron stars. Particularly relevant is the extraction of the tidal deformability by the LIGO-Virgo collaboration and the most recent determination of stellar radii by NICER, both suggesting a relatively soft equation of state (EOS) at intermediate densities. Additionally, measurements from the PREX collaboration and from pulsar timing suggest instead that the EOS is stiff in the vicinity of saturation density and at the highest densities accessible to date. But how stiff can the EOS be at these very high densities? Recent events featuring compact objects near the “lower mass gap” have raised questions about the existence of very massive neutron stars. Motivated by this finding and in light of new refinements to theoretical models, we explore the possibility that these massive objects may indeed be rapidly rotating neutron stars. Here, we explore how rotation affects both the maximum neutron star mass and their associated radii, and discuss the implications they may have on the equation of state.

Equations of state of nuclear matter↗

Enterprise: Exploration of Concepts, Perspectives and Implications for Systems Engineering

The purpose of this paper is to explore the concept of ‘enterprise’ in the context of Systems Engineering (SE). The term ‘enterprise’ has been used extensively to generally describe large complex entities that have an extensive scope of operations. However, a deeper examination of ‘enterprise’ significance for SE can provide insights as our challenges continue with increasingly complex, uncertain, ambiguous, and integrated entities struggling to thrive in the future. The paper explores three central topics. First, the concept of enterprise is introduced as a central aspect of the future focus for SE, as recognized in the INCOSE SE Vision 2035. Second, a more detailed examination of the enterprise concept is developed in relationship to SE. The thrust of this examination is to understand the nature and role of ‘enterprise’ across a broad spectrum of literature and knowledge, ultimately providing a more informed perspective of enterprise for SE. As part of this exploration, a bibliometric analysis of the term ‘enterprise’ is performed. This exploration extracts key themes (clusters) in the ‘enterprise’ literature. Third, challenges for further development and inculcation of ‘enterprise’ within the SE discipline and support for realization of the SE 2035 Vision are suggested. These challenges point out the need to ‘think differently’ about ‘enterprise’ within the SE context. ‘Enterprise’ is proposed as a central, albeit different, perspective for the SE discipline. Finally, the paper closes with a first–generation perspective for ‘enterprise’ in pursuit of the SE Vision 2035.

42 ENGINEERING↗

Traceable random numbers from a non-local quantum advantage

The unpredictability of random numbers is fundamental to both digital security and applications that fairly distribute resources. However, existing random number generators have limitations—the generation processes cannot be fully traced, audited and certified to be unpredictable. The algorithmic steps used in pseudorandom number generators are auditable, but they cannot guarantee that their outputs were a priori unpredictable given knowledge of the initial seed. Device-independent quantum random number generators can ensure that the source of randomness was unknown beforehand, but the steps used to extract the randomness are vulnerable to tampering. Here we demonstrate a fully traceable random number generation protocol based on device-independent techniques. Our protocol extracts randomness from unpredictable non-local quantum correlations, and uses distributed intertwined hash chains to cryptographically trace and verify the extraction process. This protocol forms the basis for a public traceable and certifiable quantum randomness beacon that we have launched. Over the first 40 days of operation, we completed the protocol 7,434 out of 7,454 attempts—a success rate of 99.7%. Each time the protocol succeeded, the beacon emitted a pulse of 512 bits of traceable randomness. The bits are certified to be uniform with error multiplied by actual success probability bounded by 2−64. Further, the generation of certifiable and traceable randomness represents a public service that operates with an entanglement-derived advantage over comparable classical approaches.

97 MATHEMATICS AND COMPUTING↗

GRAPH — an readout ASIC for large MCP based detectors

We present a programmable 16 channel, mixed signal, low power readout ASIC, having the project historically named Gigasample Recorder of Analog waveforms from a PHotodetector (GRAPH). It is designed to read large aperture single photon imaging detectors using micro channel plates for charge multiplication, and measuring the detector's response on crossed strips anodes to extrapolate the incoming photon position. Each channel consists of a fast, low power and low noise charge sensitive amplifier, which provides a myriad of coarse and fine programmable options for gain and shaping settings. Further, the amplified signal is recorded using, to our knowledge novel, the Hybrid Universal sampLing Architecture (HULA) ADC. A kind of mixed signal double buffer memory, that enables concurrent waveform recording, and selected event digitized data extraction. The sampling frequency is freely adjustable between few kHz up to 125 MHz, while the chip's internal digital memory holds a history 2048 samples for each channel, with a digital headroom of 12 bits. An optimized region of interest sample-read algorithm allows to extract the information just around the event pulse peak, while selecting the next event, thus substantially reducing the operational dead time. The chip is designed in 130 nm TSMC CMOS technology, and its power consumption is around 47 mW per channel.

47 OTHER INSTRUMENTATION↗

Transfer learning nonlinear plasma dynamic transitions in low dimensional embeddings via deep neural networks

Deep learning algorithms provide a new paradigm to study high-dimensional dynamical behaviors, such as those in fusion plasma systems. Development of novel, data-driven model reduction methods, coupled with detection of abnormal modes with plasma physics, opens a unique opportunity to identify plasma instabilities through automated construction of parsimonious models that can be tuned to balance accuracy and cost. Our fusion transfer learning (FTL) model demonstrates success in rapidly reconstructing nonlinear kink mode structures by learning from a limited amount of nonlinear simulation data. The knowledge transfer process leverages a pre-trained neural encoder–decoder network, initially trained on linear simulations, to effectively capture nonlinear dynamics. The low-dimensional embeddings extract the coherent structures of interest, while preserving the inherent dynamics of the complex system. Experimental results highlight FTL’s capacity to capture transitional behaviors and dynamical features in plasma dynamics—a task often challenging for conventional methods. The model developed in this study is generalizable and can be extended broadly through transfer learning to address various magnetohydrodynamics modes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Improving the Quasi‐Biennial Oscillation via a Surrogate‐Accelerated Multi‐Objective Optimization

Accurate simulation of the quasi-biennial oscillation (QBO) is challenging due to uncertainties in representing convectively generated gravity waves. We develop an end-to-end uncertainty quantification workflow that calibrates these gravity wave processes in E3SM for a realistic QBO. Central to our approach is a domain knowledge-informed, compressed representation of high-dimensional spatio-temporal wind fields. By employing a parsimonious statistical model that learns the fundamental frequency from complex observations, we extract interpretable and physically meaningful quantities capturing key attributes. Building on this, we train a probabilistic surrogate model that approximates the fundamental characteristics of the QBO as functions of critical physics parameters governing gravity wave generation. Leveraging the Karhunen–Loève decomposition, our surrogate efficiently represents these characteristics as a set of orthogonal features, capturing cross-correlations among multiple physics quantities evaluated at different pressure levels and enabling rapid surrogate-based inference at a fraction of the computational cost of full-scale simulations. Finally, we analyze the inverse problem using a multi-objective approach. Our study reveals a tension between amplitude and period that constrains the QBO representation, precluding a single optimal solution. To navigate this, we quantify the bi-criteria trade-off and generate a set of Pareto optimal parameter values that balance the conflicting objectives. This integrated workflow improves the fidelity of QBO simulations and offers a versatile template for uncertainty quantification in complex geophysical models.

54 ENVIRONMENTAL SCIENCES↗

Titanium alloy response sensitivity to variations in spectral reconstructions of National Ignition Facility xenon line-emission x-ray sources

Thermomechanical shock experiments on the National Ignition Facility (NIF) aim to study high strain rate dynamic material response. In such experiments, the NIF laser is used to generate high fluence x-ray emission sources, which irradiate material samples of interest. Under sufficiently high x-ray energy deposition, thermomechanical impulses are generated in the materials. While it is known that the characteristics of x-ray generated impulses vary as a function of incident x-ray spectra, it remains unclear how spectral assumptions and uncertainties in NIF spectral reconstructions affect our interpretation of impulsive loading. Here, in this paper, we simulate the response of a standard titanium alloy baseline sample to synthetic analytically derived and measured NIF xenon line-emission x-ray sources with a radiation hydrodynamics code. We vary the source spectral characteristics based on different source reconstruction techniques to understand the resulting variation in baseline sample response and compare the simulated response with experimental results. We find that the response is highly sensitive to assumptions made about the spectral contents and that knowledge of spectral uncertainties bounds our understanding of the resulting material response. The results of this effort help to extend our ability to use baseline material samples to extract quantitative properties from x-ray experiments on the NIF.

Alloys↗

FTL: Transfer Learning Nonlinear Plasma Dynamic Transitions in Low Dimensional Embeddings (FTL) v1.0

Fusion Transfer Learning (FTL) model provides a new paradigm to study high-dimensional dynamical behaviors, such as those in fusion plasma systems. The knowledge transfer process leverages a pre-trained neural encoder-decoder network, initially trained on linear simulations, to effectively capture nonlinear dynamics. The low-dimensional embeddings extract the coherent structures of interest, while preserving the inherent dynamics of the complex system. Experimental results highlight FTL's capacity to capture transitional behaviors and dynamical features in plasma dynamics -- a task often challenging for conventional methods. The model developed in this study is generalizable and can be extended broadly through transfer learning to address various magnetohydrodynamics (MHD) modes.

Bai, Zhe↗

Signatures of Mollicutes-related endobacteria in publicly available Mucoromycota genomes

ABSTRACT Mucoromycota fungi and their Mollicutes-related endobacteria (MRE) are an ideal system for studying bacterial–fungal interactions and evolution due to the long-term and intimate nature of their interactions. However, methods for detecting MRE face specific challenges due to the poor representation of MRE in sequencing databases coupled with the high sequence divergence of their genomes, making traditional similarity searches unreliable. This has precluded estimations on the diversity of MRE associated with Mucoromycota. To determine the prevalence of previously undetected MRE in fungal genome sequences, we scanned 389 Mucoromycota genome assemblies available from the National Center for Biotechnology Information for the presence of MRE sequences using publicly available tools to map contigs from fungal assemblies to publicly available MRE genomes. We demonstrate a higher diversity of MRE genomes than previously described in Mucoromycota and a lack of cophylogeny between MRE and the majority of their fungal hosts. This supports the late invasion hypothesis regarding MRE acquisition across most of the examined fungal families. In contrast with other Mucoromycota lineages, MRE from the Gigasporaceae displayed some degree of cophylogeny with their hosts, which may indicate that horizontal transmission is restricted between members of this family or that transmission is strictly vertical. These results underscore the need for a refined process to capture sequencing data from potential fungal endosymbionts to discern their evolution and transmission. Screens of fungal genomes for MRE can help improve the quality of fungal genome assemblies while identifying new MRE lineages to further test hypotheses on their origin and evolution. IMPORTANCE Mollicutes-related endobacteria (MRE) are obligate intracellular bacteria found within Mucoromycota fungi. Despite their frequent detection, MRE roles in host functioning are still unknown. Comparative genomic investigations can improve our understanding of the impact of MRE on their fungal hosts by identifying similarities and differences in MRE genome evolution. However, MRE genomes have only been assembled from a small fraction of Mucoromycota hosts. Here, we demonstrate that MRE can be present yet undetected in publicly available Mucoromycota genome assemblies. We use these newfound sequences to assess the broader diversity of MRE and their phylogenetic relationships with respect to their hosts. We demonstrate that publicly available tools can be used to extract novel MRE sequences from assembled fungal genomes leading to insights on MRE evolution. This work contributes to a greater understanding of the fungal microbiome, which is crucial to improving knowledge on the dynamics and impacts of fungi in microbial ecosystems.

59 BASIC BIOLOGICAL SCIENCES↗

A model to assess Zircaloy’s mechanical property changes following a transient beyond critical heat flux

Maintaining the integrity of nuclear fuel rods is essential for ensuring public health and safety in nuclear power generation. During reactor operation, this integrity is confirmed by demonstrating compliance with established regulatory acceptance criteria. For moderate-frequency events, such as limiting transients and anticipated operational occurrences (AOOs), the current fuel integrity criterion is based on preventing boiling transition. This criterion assumes that prevention of boiling transition will prevent excessive cladding heating and, thus, fuel failure during normal operations. While conservative, this approach places significant constraints on core design, fuel cycle economics, and a plant’s ability to perform major power uprates, leading to suboptimal fuel utilization and inefficient carbon-free energy production. A more efficient approach could be achieved by revising the failure criterion to a material-specific limit rather than strictly preventing the boiling transition, since boiling transition per se is not a cause of fuel cladding failure. Here, as a result, a new licensing framework based on material properties, termed time-at-temperature (t@T), is needed. This approach would allow for brief periods of post–critical heat flux operation during an AOO without compromising safety. Implementing the t@T licensing strategy requires a robust technical foundation in material properties, which must be established through comprehensive data collection on both unirradiated and irradiated fuel and cladding materials. This foundation would enable the development of a safety basis that ensures safe operation while providing greater flexibility and efficiency for reactor operation. This paper documents a thorough review of the available data to establish a baseline knowledge that can inform the development of cladding mechanical models, as well as identify experimental data gaps that need to be addressed in future research. Machine learning and data informatics were utilized to extract the importance of parameters on the t@T parameter. Industry tools were used to perform baseline analyses to define the relevant transient conditions for data analysis. The subsequent review successfully identified applicable experimental data, as well as sufficient data to evaluate changes in cladding mechanical properties following an AOO transient. Rather than developing new models, this work coupled existing irradiation annealing and recrystallization models to calculate changes in hardness, yield stress, and ultimate tensile stress following an AOO event. The findings from this review were summarized to highlight the experimental data needs required to fill remaining gaps and support the development of future t@T licensing methodologies.

Cladding performance↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A LENS on DUNE-PRISM: Characterizing a Neutrino Beam with Off-Axis Measurements

Upcoming precision long-baseline neutrino oscillation experiments will be severely limited by the large systematic uncertainties associated with neutrino flux predictions and neutrino--nucleus cross sections. A promising remedy is the PRISM (Precision Reaction Independent Spectrum Measurement) technique, whereby the near detector measures the neutrino spectrum at different angles with respect to the beam axis. These measurements are then linearly combined into a prediction of the oscillated neutrino flux at the far detector. This prediction is data-driven, but still dependent on some theoretical knowledge about the neutrino flux. In this paper, we study to what extent off-axis measurements themselves can be used to directly constrain neutrino flux models. In particular, we use them to extract separately the fluxes and spectra of different meson species in the beam. We call this measurement LENS (Lateral Extraction of Neutrino Spectra). Second, we demonstrate how the thus improved flux model helps to further constrain the far detector flux prediction, thereby ultimately improving oscillation measurements.

FOS: Physical sciences↗

Exploiting Multi-Domain Features for Detection of Unclassified Electromagnetic Signals

Deep Learning based classification techniques have shown excellent performance in static environments, where the training and testing samples are drawn from the same distribution. However, real world scenarios often present samples that do not belong to the known set of classes chosen during training. This is quite common for electromagnetic signals, where it is impractical to assume that all possible waveforms are known a-priori, specially in scenarios like warfare. To address this problem, we propose a deep learning based adversarial model where the generator learns to generate waveform features that can deceive the discriminator model as true samples. We introduce domain knowledge of wireless signals by decomposing the signal into a lower dimensional unique feature set, which is used for classifying known versus unknown signals. We further introduce multiple domain representations of the signal to extract features and combine them together to accurately classify new waveforms as an unknown class. Our results show that combined features from multiple domains outperform any single domain representation, especially at low SNR regimes with fewer number of samples to classify.

99 - GENERAL AND MISCELLANEOUS↗

Exploiting Multi-Domain Features for Detection of Unclassified Electromagnetic Signals (Presentation)

Deep Learning based classification techniques have shown excellent performance in static environments, where the training and testing samples are drawn from the same distribution. However, real world scenarios often present samples that do not belong to the known set of classes chosen during training. This is quite common for electromagnetic signals, where it is impractical to assume that all possible waveforms are known a-priori, specially in scenarios like warfare. To address this problem, we propose a deep learning based adversarial model where the generator learns to generate waveform features that can deceive the discriminator model as true samples. We introduce domain knowledge of wireless signals by decomposing the signal into a lower dimensional unique feature set, which is used for classifying known versus unknown signals. We further introduce multiple domain representations of the signal to extract features and combine them together to accurately classify new waveforms as an unknown class. Our results show that combined features from multiple domains outperform any single domain representation, especially at low SNR regimes with fewer number of samples to classify.

99 - GENERAL AND MISCELLANEOUS↗

A semi-analytic estimate for the effective sound speed counterterm in the EFTofLSS

The Effective Field Theory of Large Scale Structure (EFTofLSS) has found tremendous success as a perturbative framework for the evolution of large scale structure, and it is now routinely used to compare theoretical predictions against cosmological observations. The model for the total matter field includes one nuisance parameter at 1-loop order, the effective sound speed, which can be extracted by matching the EFT to full N-body simulations. In this work we first leverage the Layzer-Irvine cosmic energy equation to show that the equation of state can be exactly computed with knowledge of the fully nonlinear power spectrum. When augmented with separate universe methods, we show one can estimate the effective sound speed. This estimate is in good agreement with simulation results, with errors at the few tens of percent level. Here, we apply our method to investigate the cosmology dependence of the effective sound speed and to shed light on what cosmic structures shape its value.

Cosmological perturbation theory in GR and beyond↗

Actinomycetota isolated from the sponge Hymeniacidon perlevis as a source of novel compounds with pharmacological applications: diversity, bioactivity screening, and metabolomic analysis

Abstract Aims To combat health conditions, such as multi-resistant bacterial infections, cancer, and metabolic diseases, new drugs need to be urgently found and, in this respect, marine Actinomycetota have a high potential to produce secondary metabolites with pharmacological importance. We aimed to study the cultivable Actinomycetota community associated with a marine sponge from the Portuguese coast, Hymeniacidon perlevis, and investigate the potential of the retrieved isolates to produce compounds with antimicrobial, anticancer and anti-obesity properties. Methods and results The analysis of the 16S rRNA gene revealed 79 Actinomycetota isolates affiliated with 12 genera—Brachybacterium, Dietzia, Glutamicibacter, Gordonia, Micrococcus, Micromonospora, Nocardia, Nocardiopsis, Paenoartrhobacter, Rhodococcus, Streptomyces, and Tsukamurella, most of which affiliated with the genus Streptomyces. The screening of antimicrobial activity revealed 13 strains, all belonging to the Streptomyces genus, capable of inhibiting the growth of Candida albicans, Bacillus subtilis, or Staphylococcus aureus. Forty-three extracts exhibited cytotoxic activity against at least one tested cell line (HepG2, HCT-116, and hCMEC-D3). Three extracts that were active against the two cancer cell lines tested, did not reduce the viability of the non-cancer endothelial cell line, hCMEC-D3. One Gordonia strain exhibited anti-obesity activity, revealed by its ability to reduce the neutral lipids in zebrafish larvae. Mass spectrometry-based dereplication analysis of active extracts identified several compounds associated with known Actinomycetota natural products. Nonetheless, five clusters contained metabolites that did not match any annotated natural products, suggesting they may represent new bioactive molecules. Conclusions This work contributed to increase the knowledge on the diversity and bioactive potential of Actinomycetota associated with H. perlevis.

Fonseca, Ana C.↗