Search NASA⌕ Search

SEARCH · Search NASA

Results for “database”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Machine learning enhanced predictions of ICRF heating: Overcoming numerical limitations via data curation

In this work, we present the development of robust surrogate models for Ion Cyclotron Range of Frequencies (ICRF) and High-Harmonic Fast Wave (HHFW) heating predictions in fusion plasmas. Building upon our previous efforts to achieve real-time capable models, we identify the cause of the outliers found using TORIC in certain HHFW heating scenarios. The outliers are observed to be spurious ion Bernstein wave (IBW)-like modes caused by a wavelength control algorithm designed to address challenging scenarios with high perpendicular wavenumbers. The effect arises from the modulation in the perpendicular susceptibility, which can induce sign reversal and IBW-like propagation for scenarios featuring normalized ion Larmor radius λ i ≫ 1. We use TORIC with this algorithm disabled to generate a novel HHFW-NSTX database that is free of outliers. Surrogate models trained on this database, including Random Forest Regressor (RFR), Multi-Layer Perceptrons, and Gaussian Process Regressors (GPR), demonstrate the ability to accurately predict HHFW heating profiles, with regression scores of R 2 ∈[0.93−0.99]. Additionally we demonstrate that it is possible to generalize predictions beyond training data by the use of both RFR and GPR models, enabling the prediction of scenarios previously limited to the original model. GPR models also provide uncertainty quantification, offering insights into model confidence. This work introduces a comprehensive Verification, Validation, and Uncertainty Quantification methodology for surrogate modeling, applicable not only to ICRF heating but also to other RF heating challenges and fusion physics problems. Beyond accelerated inference, these models show effective extrapolation capabilities, providing an alternative for addressing numerical challenges.

Artificial neural networks↗

Validation of the GFS model for gyrokinetic stability of NSTX pedestal data

This study presents a large database validation of the gyro fluid system (GFS) model for linear gyrokinetic stability for high-mode (H-mode) edge transport barrier conditions in the national spherical torus experiment (NSTX) tokamak. The database of linear stability calculations with the CGYRO gyrokinetic code was produced using plasma profile measurements from NSTX discharges to identify kinetic ballooning modes (KBM), trapped electron modes (TEM), and micro-tearing modes (MTM) that limit the pressure profile gradient in the H-mode barrier. A novel Bayesian optimization approach determines optimal resolution parameters for GFS specifically for spherical tokamak pedestal conditions. Our results demonstrate that GFS, with optimized resolution, can achieve accurate linear stability analysis in NSTX pedestal conditions for reduced resolution compared to CGYRO. GFS can accurately find the KBM, TEM, and MTM instability branches. Parametric analysis reveals that GFS accuracy in this extreme pedestal parameter range is degraded for low magnetic shear and near the separatrix conditions. These findings establish GFS as a fast linear eigenmode solver for spherical tokamak pedestal gyrokinetic stability and demonstrate a systematic methodology for determining the optimum resolution settings.

Yang, Minglei [Oak Ridge National Laboratory (ORNL↗

Cataloging Legacy Data from the Tritium Systems Test Assembly Program

The Tritium Systems Test Assembly (TSTA) at Los Alamos National Laboratory, operational from 1984 to 2001, was critical in advancing fusion fuel cycle technologies, including tritium storage, gas separation, and pumping. TSTA’s contributions, particularly in safe tritium operations, have influenced subsequent fusion projects. This paper discusses the ongoing effort to digitize and catalog TSTA’s historical data to create a searchable resource for the fusion research community. While the long-term objective is to develop a relational database for structured data management, the project remains in the early phase, with current efforts focused on scanning and indexing physical documents. Initial plans for database implementations are also presented, outlining key considerations for structure, query indexing, and standardization. As digitization progresses, future discussions will refine these implantation details to ensure an efficient and comprehensive system. This initiative aims to preserve critical legacy data, enhance the design of tritium system facilities, and support the next generation of fusion energy research.

42 ENGINEERING↗

Validation of prediction capability of operating space for plasma initiation in MAST-U

DYON is a plasma initiation modelling code that solves the differential equation system of the full circuit equations (plasma current, active coil currents and eddy currents in full passive structures) and 0D global energy and particle balance equations (Kim 2022 Nucl. Fusion 62 126012). In order to test the capability of the full electromagnetic plasma initiation model to predict individual discharges in experiments and thus the operating space in the device, a dedicated experimental database was built in MAST-U by scanning the prefilled gas pressure p 0 and the induced loop voltage V loop . In the experimental operating space of p 0 and V loop the lower and the upper limits of p 0 are determined by the plasma breakdown failure and the plasma burn-through failure, respectively. The lower limit of V loop is determined by the plasma burn-through failure. By directly reading the control room data used in each discharge (i.e. currents in the solenoid, poloidal field coils, and toroidal field coils, p 0 , and gas puffing rate), the full electromagnetic DYON consistently predicted the failed breakdown, failed burn-through, and successful plasma initiation discharges in the experimental database, demonstrating its capability to predict the operating space for inductive plasma initiation. The Paschen curve calculated with the effective connection length in MAST-U indicates a much higher p 0 required for plasma breakdown than the experimental data, indicating that individual field line evaluation is necessary to calculate the quantitative requirements for Townsend breakdown. The demonstration in this paper shows that the full electromagnetic DYON could be a useful simulation tool to assess the feasibility of inductive plasma initiation and to optimise operating scenarios in future devices.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Impact of ionization peak location on measured opaqueness in DIII-D H-mode plasmas

This study investigates the relationship between electron pedestal density and the location of the ionization peak on neutral penetration in DIII-D H-mode plasmas, utilizing a database of Lyman-α emission measurements. The high electron density leads to neutrals being ‘screened’ and the ionization front being pushed out into the Scrape-Off Layer (SOL). This is also referred to as the neutral opaqueness, which is heuristically expected to scale with edge plasma density and machine size. However, at lower electron pedestal density, the penetration depth of the neutrals varies, and measured opaqueness deviates from the heuristic scaling. The database reveals that at low density, when the ionization peak is located in SOL region, the linear relationship between the electron density and neutral penetration holds. However, when the peak is located inside the separatrix, the penetration of the neutrals (λ n 0 ) is much wider ~3.0–3.5 cm, breaking the heuristic opaqueness approximation. These findings provide valuable insights into fueling efficiency and plasma behavior, with implications for Fusion Pilot Plants where high pedestal densities are anticipated and where the neutral opaqueness behaves like its heuristic approximation. This analysis offers a framework to refine neutral opaqueness approximations, enhancing the predictive capability for advanced tokamak operations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Core plasma fueling by fast inward particle transport after hydrogen pellet injection in Wendelstein 7-X

A large database of more than 1000 individual cryogenic hydrogen pellets injected into Wendelstein 7-X for plasma fueling was analyzed to improve the understanding of the three phases of the process: the ablation, deposition and transport of the pellet material. Kilohertz-sampled electron density and temperature measurements revealed a more complex drift behavior than predicted by numerical code simulation. It could be explained by the poloidal plasma E r x B- drift rotation, which plays a significant role in stellarators, but was not previously considered in pellet injection codes like HPI2. The drift results in a fast poloidal rotation of the pellet material around the plasma core, leading to an almost homogeneous deposition over the involved flux surfaces regardless of magnetic high and low field side injection geometry. Additionally, a novel fast inward directed transport mechanism (‘FIT-effect’) was observed. The effect occurs on timescales of tens of milliseconds and cannot be explained by neoclassical transport or diffusion. It might be linked to the turbulence pinch recently found in Wendelstein 7-X. When the FIT-effect occurs, the pellet particles are rapidly transferred from the deposition flux surfaces to the plasma core, causing the plasma density profile to peak, which is beneficial for confinement in Wendelstein 7-X. The large pellet injection database was statistical analyzed with regard to pellet and plasma parameters, which delivered some starting points towards developing an understanding of the physics behind the FIT-effect. The results indicate, that plasma core fueling via pellet injection is largely independent of the injection geometry in stellarators under certain conditions, reducing the technical complexity of the injection system.

Wendelstein 7-X↗

Macroscopic trends of neoclassical tearing stability in high-field H-mode tokamak pilot plants

The neoclassical tearing mode (NTM) stability metric—minimum marginally stable island width $w$$^{*}_{m}$—was compared across 14651 inductive high-field tokamak pilot plant equilibria. Larger devices with reduced elongation and/or increased minor radius demonstrated an order-of-magnitude increase in $w$$^{*}_{m}$, primarily due to a reduction in bootstrap drive. This work is part of an ongoing effort to ensure passive NTM-stability in the ARC tokamak, in which the technology to achieve active tearing-suppression with localised electron cyclotron current drive does not yet exist. The equilibrium scenarios in the database were Monte Carlo generated and normalised to the same >400MW fusion power, minimum pressure scenario at a range of plasma currents, before tearing analysis using the modified Rutherford equation was applied for all resonant poloidal and toroidal m, n modes up to n = 4. Single-helicity toroidal Δ' calculations in resistive DCON set the minimum marginally stable island width, and a simple modal scaling proportional to –m 2 n –1 was identified for high-m Δ' values. The dominant correlates of $w$$^{*}_{m}$ and Δ' across the database were analysed using interpretable machine learning techniques.

NTM seeding↗

Overlooked cooling effects of albedo in terrestrial ecosystems

Radiative forcing (RF) resulting from changes in surface albedo is increasingly recognized as a significant driver of global climate change but has not been adequately estimated, including by Intergovernmental Panel on Climate Change (IPCC) assessment reports, compared with other warming agents. Here, we first present the physical foundation for modeling albedo-induced RF and the consequent global warming impact (GWI Δα ). We then highlight the shortcomings of available current databases and methodologies for calculating GWI Δα at multiple temporal scales. There is a clear lack of comprehensive in situ measurements of albedo due to sparse geographic coverage of ground-based stations, whereas estimates from satellites suffer from biases due to the limited frequency of image collection, and estimates from earth system models (ESMs) suffer from very coarse spatial resolution land cover maps and associated albedo values in pre-determined lookup tables. Field measurements of albedo show large differences by ecosystem type and large diurnal and seasonal changes. As indicated from our findings in southwest Michigan, GWI Δα is substantial, exceeding the RF Δα values of IPCC reports. Inclusion of GWI Δα to landowners and carbon credit markets for specific management practices are needed in future policies. We further identify four pressing research priorities: developing a comprehensive albedo database, pinpointing accurate reference sites within managed landscapes, refining algorithms for remote sensing of albedo by integrating geostationary and other orbital satellites, and integrating the GWI Δα component into future ESMs.

54 ENVIRONMENTAL SCIENCES↗

Coincidence anomaly detection for unsupervised locating of edge localized modes in the DIII-D tokamak dataset

Using supervised learning to train a machine learning model to predict an on-coming edge localized mode (ELM) requires a large number of labeled samples. Creating an appropriate data set from the very large database of discharges at a long-running tokamak, such as DIII-D, would be a very time-consuming process for a human. Considering this need and difficulty, we use coincidence anomaly detection, an unsupervised learning technique, to train an ELM-identifier to identify and label ELMs in the DIII-D discharge database. This ELM-identifier shows, simultaneously, a precision of 0.68 and a recall of 0.63 (AUC is 0.73) on identifying ELMs in example time series pulled from thousands of discharges spanning five years. In a test set of 50 discharges, the algorithm finds over 26 thousand ELM candidates, more than 5 times the existing catalog of ELMs labeled by humans.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

From natural language to control signals: a conceptual framework for semantic channel finding in complex experimental infrastructure

Modern experimental platforms such as particle accelerators, fusion devices, telescopes, and industrial process control systems expose tens to hundreds of thousands of control and diagnostic channels, accumulated over decades of hardware evolution. Operators and AI systems alike depend on informal expert knowledge, inconsistent naming conventions, and scattered documentation to locate the signals required for monitoring, troubleshooting, and automated control, creating a persistent bottleneck for reliability, scalability, and emerging language-model-driven interfaces. We formalize semantic channel finding, the task of mapping natural-language intent to concrete control-system signals, as a general problem in complex experimental infrastructure, and introduce a four-paradigm conceptual framework to guide architecture selection based on facility-specific data regimes. The paradigms span (i) direct in-context lookup over small, curated channel dictionaries, (ii) constrained hierarchical navigation through structured trees, (iii) interactive agent exploration using iterative reasoning and tool-based database queries, and (iv) ontology-grounded semantic search that decouples channel meaning from facility-specific naming conventions. We demonstrate the practical feasibility of each paradigm through proof-of-concept implementations at four operational facilities spanning two orders of magnitude in scale: from compact free-electron lasers to large synchrotron light sources, operating under diverse control-system architectures ranging from clean hierarchical naming schemes to legacy environments with decades of heterogeneous conventions. Where evaluated against expert-curated operational queries, these instantiations achieve 90%–97% accuracy, validating the framework’s applicability across real-world deployment scenarios. To accelerate adoption across the broader scientific and industrial control-system community, we release open-source, plug-and-play implementations of all three interactive paradigms-direct lookup, hierarchical navigation, and middle-layer exploration-within the Osprey framework, together with tools for channel database generation, interactive testing, and minimal-configuration deployment. This work establishes semantic channel finding as a foundational capability for human-centric and agentic AI interfaces at large-scale facilities, providing both a systematic framework for architecture design and practical resources to enable adoption without building custom infrastructure from scratch.

channel finding↗

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES↗

BindingDB in 2024: a FAIR knowledgebase of protein-small molecule binding data

Abstract BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models and computational chemistry methods development. This update reports significant growth and enhancements since our last review in 2016. Of note, the database now contains 2.9 million binding measurements spanning 1.3 million compounds and thousands of protein targets. This growth is largely attributable to our unique focus on curating data from US patents, which has yielded a substantial influx of novel binding data. Recent improvements include a remake of the website following responsive web design principles, enhanced search and filtering capabilities, new data download options and webservices and establishment of a long-term data archive replicated across dispersed sites. We also discuss BindingDB’s positioning relative to related resources, its open data sharing policies, insights gleaned from the dataset and plans for future growth and development.

Liu, Tiqing↗

Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes

Viruses are ubiquitous in all environments and impact host metabolism, evolution, and ecology, although our knowledge of their biodiversity is still extremely limited. Viral diversity from genomic and metagenomic datasets has led to an explosion of uncultivated virus genomes (UViGs) and the development of specialized databases to catalog this viral diversity, though many lack comprehensive integration. Here, we introduce meta-virus resource (MetaVR), the successor of the IMG/VR database, designed to overcome previous limitations such as large-scale querying and programmatic access. Drawing on the increase of publicly available genomes and metagenomes, MetaVR significantly expands viral diversity, now comprising 24,435,662 UViGs, a 57.6% increase from its predecessor, organized into over 12 million viral operational taxonomic units. Key enhancements include the integration of curated eukaryotic host information, the integration of protein clusters and predicted structures for comparative studies, and an API for programmatic data access. Furthermore, MetaVR features an updated taxonomic framework based on ICTV release 39, assignment to Baltimore classes, and enhanced host assignment through novel computational tools like iPHoP. These advancements position MetaVR as a unique resource for exploring viral diversity, evolution, and host interactions across diverse environments. MetaVR can be freely accessed at https://www.meta-virome.org/.

Fiamenghi, Mateus B↗

PSInet: a new global water potential network

Abstract Given the pressing challenges posed by climate change, it is crucial to develop a deeper understanding of the impacts of escalating drought and heat stress on terrestrial ecosystems and the vital services they offer. Soil and plant water potential play a pivotal role in governing the dynamics of water within ecosystems and exert direct control over plant function and mortality risk during periods of ecological stress. However, existing observations of water potential suffer from significant limitations, including their sporadic and discontinuous nature, inconsistent representation of relevant spatio-temporal scales and numerous methodological challenges. These limitations hinder the comprehensive and synthetic research needed to enhance our conceptual understanding and predictive models of plant function and survival under limited moisture availability. In this article, we present PSInet (PSI—for the Greek letter Ψ used to denote water potential), a novel collaborative network of researchers and data, designed to bridge the current critical information gap in water potential data. The primary objectives of PSInet are as follows. (i) Establishing the first openly accessible global database for time series of plant and soil water potential measurements, while providing important linkages with other relevant observation networks. (ii) Fostering an inclusive and diverse collaborative environment for all scientists studying water potential in various stages of their careers. (iii) Standardizing methodologies, processing and interpretation of water potential data through the engagement of a global community of scientists, facilitated by the dissemination of standardized protocols, best practices and early career training opportunities. (iv) Facilitating the use of the PSInet database for synthesizing knowledge and addressing prominent gaps in our understanding of plants’ physiological responses to various environmental stressors. The PSInet initiative is integral to meeting the fundamental research challenge of discerning which plant species will thrive and which will be vulnerable in a world undergoing rapid warming and increasing aridification.

Forestry↗

Populus VariantDB v3.2 facilitates CRISPR and functional genomics research

The success of CRISPR genome editing studies depends critically on the precision of guide RNA (gRNA) design. Sequence polymorphisms in outcrossing tree species pose design hazards that can render CRISPR genome editing ineffective. Despite recent advances in tree genome sequencing with haplotype resolution, sequence polymorphism information remains largely inaccessible to various functional genomics research efforts. The Populus VariantDB v3.2 addresses these challenges by providing a user-friendly search engine to query sequence polymorphisms of heterozygous genomes. The database accepts short sequences, such as gRNAs and primers, as input for searching against multiple poplar genomes, including hybrids, with customizable parameters. We provide examples to showcase the utilities of VariantDB in improving the precision of gRNA or primer design. The platform-agnostic nature of the probe search design makes Populus VariantDB v3.2 a versatile tool for the rapidly evolving CRISPR field and other sequence-sensitive functional genomics applications. The database schema is expandable and can accommodate additional tree genomes to broaden its user base.

59 BASIC BIOLOGICAL SCIENCES↗

Measurement of the 28 Si ⁢(𝑛,𝑛′⁢𝛾) cross section with 𝑛, 𝛾, and correlated 𝑛−𝛾 angular distributions

Silicon has become an unavoidable element in the circuitry central to everyday life. In turn, the interactions of silicon isotopes with neutrons for nuclear physics applications, among other motivations, have become increasingly important to understand. The dominant isotope of silicon, 28 Si, is thus of primary interest for enhanced understanding for neutron transport calculations and related investigations. Unfortunately, the existing measurement database for neutron scattering reactions on 28 Si is minimal, and nuclear data evaluations on this topic have not been updated for decades. This article details new measurements of the 28 Si ⁢(𝑛,𝑛′⁢𝛾) reaction utilizing multiple analysis methods available within the correlated gamma neutron array for scattering (CoGNAC). Specifically, high-precision near-threshold results and high-incident-energy results were obtained using the 𝛾-only and correlated 𝑛−𝛾 techniques. First-ever measurements of the correlated 𝑛−𝛾 angular distribution for particles emitted following population of the first excited state in 28 Si were obtained as well, which provide unique insight into theoretical descriptions of the inelastic neutron scattering reaction mechanism itself and detailed guidance for nuclear reaction models. The results agree well with literature data where they exist, and substantially expand on the current database for neutron reactions on 28 Si .

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Privacy Preserving Federated Learning for Advanced Scientific Ecosystems

We present a framework to provide privacy preserving (PP) federating learning (FL) across multiple computational and experimental facilities. This work joins the compute capabilities of National Energy Research Scientific Computing Center (NERSC) and Oak Ridge National Laboratory Research Cloud (ORC) with simulated experimental data, such as those produced at the SLAC National Accelerator Laboratory and Spallation Neutron Source (SNS). We describe the software infrastructure developed to provide privacy for computational and experimental networks. We developed algorithmic privacy across the federated system by embedding database security, computation, and communication into the federation architecture, utilizing scientific tools developed by the experimental community.

Archibald, Rick [ORNL] (ORCID:0000000245389780)↗