Search NASA⌕ Search

SEARCH · Search NASA

Results for “Databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34

Land Use Change Alters Soil Organic Carbon: Constrained Global Patterns and Predictors

Abstract Land use change (LUC) alters the global carbon (C) stock, but our estimation of the alteration remains uncertain and is a major impediment to predicting the global C cycle. The uncertainty is partly due to the limited number and geographical bias of observations, and limited exploration of its predictors. Here we generated a comprehensive global database of 5,980 observations from 790 articles. The number of sites evaluated is at least seven times larger than in previous meta‐analyses. Our constrained estimates of different LUC's effects on soil organic C (SOC) and their variations across global climates reveal underestimation/overestimation in previous estimates. Converting forests and grasslands to croplands reduced SOC by 24.5% ± 1.53% (−11.03 ± 1.06 Mg ha −1 ) and 22.7% ± 1.22% (−8.09 ± 0.67 Mg ha −1 ), while 28.0% ± 1.56% (4.46 ± 0.42 Mg ha −1 ) and 33.5% ± 1.68% (5.8 ± 0.38 Mg ha −1 ) increases, respectively, were obtained in the reverse processes. Converting forests to grasslands decreased SOC by 2.1% ± 1.22% (−1.13 ± 0.44 Mg ha −1 ), while the reverse process increased SOC by 18.6% ± 1.73% (3.31 ± 0.51 Mg ha −1 ). Modeled relative importance of 10 drivers of LUC's impact on SOC revealed that higher initial SOC (iSOC) does not solely determine SOC loss in SOC‐negative LUC scenarios as previously proposed. Across four decades, reconverting croplands to forests and grasslands recovered only 49.5% (6.1 ± 0.51 Mg ha −1 ) and 75.3% (7.0 ± 0.38 Mg ha −1 ) of the iSOC, respectively, indicating the need for protecting C‐rich ecosystems. Our global data set advances information on LUC's effect on SOC and can be valuable to constrain Earth system models to reliably estimate global SOC stocks and plan climate change mitigation strategies.

Environmental Sciences & Ecology↗

Linkages Between Mineral Element Composition of Soils and Sediments With Hyporheic Zone Dissolved Organic Matter Chemistry Across the Contiguous United States

The hyporheic zone is a hotspot for biogeochemical cycling where interactions with mineral metals preserve the release and biodegradation of organic matter (OM). A small fraction of OM can still be exchanged between localized sediments and the overlying water column, and recent evidence suggests there exists a longitudinal structuring in sediment dissolved OM (DOM) chemistry across the continental United States (CONUS). In this study, we tested a hypothesis that water extractable sediment DOM chemistry could be explained by sediment metal contents and integrative watershed scale features at the CONUS scale. Crowdsourced samples were characterized for high resolution mass spectrometry and coupled with sediment metals determined via x-ray fluorescence as well as with land cover and soil elemental information obtained from national databases. Our results highlight weak relationships between DOM chemistry and elemental composition at the CONUS scale indicating limited transferability of organo-metal linkages into multi-scale hydrobiogeochemical models.

58 GEOSCIENCES↗

Responses of Marginal and Intrinsic Water-Use Efficiency to Changing Aridity Using FLUXNET Observations

According to classic stomatal optimization theory, plant stomata are regulated to maximize carbon assimilation for a given water loss. A key component of stomatal optimization models is marginal water-use efficiency (mWUE), the ratio of the change of transpiration to the change in carbon assimilation. Although the mWUE is often assumed to be constant, variability of mWUE under changing hydrologic conditions has been reported. However, there has yet to be a consensus on the patterns of mWUE variabilities and their relations with atmospheric aridity. We investigate the dynamics of mWUE in response to vapor pressure deficit (VPD) and aridity index using carbon and water fluxes from 115 eddy covariance towers available from the global database FLUXNET. We demonstrate a non-linear mWUE-VPD relationship at a sub-daily scale in general; mWUE varies substantially at both low and high VPD levels. However, mWUE remains relatively constant within the mid-range of VPD. Despite the highly non-linear relationship between mWUE and VPD, the relationship can be informed by the strong linear relationship between ecosystem-level inherent water-use efficiency (IWUE) and mWUE using the slope, m *. We further identify site-specific m * and its variability with changing site-level aridity across six vegetation types. We suggest accurately representing the relationship between IWUE and VPD using Michaelis–Menten or quadratic functions to ensure precise estimation of mWUE variability for individual sites.

54 ENVIRONMENTAL SCIENCES↗

A Comprehensive Northern Hemisphere Particle Microphysics Data Set From the Precipitation Imaging Package

Microphysical observations of precipitating particles are critical data sources for numerical weather prediction models and remote sensing retrieval algorithms. However, obtaining coherent data sets of particle microphysics is challenging as they are often unindexed, distributed across disparate institutions, and have not undergone a uniform quality control process. This work introduces a unified, comprehensive Northern Hemisphere particle microphysical data set from the National Aeronautics and Space Administration precipitation imaging package (PIP), accessible in a standardized data format and stored in a centralized, public repository. Data is collected from 10 measurement sites spanning 34° latitude (37°N–71°N) over 10 years (2014–2023), which comprise a set of 1,070,000 precipitating minutes. The provided data set includes measurements of a suite of microphysical attributes for both rain and snow, including distributions of particle size, vertical velocity, and effective density, along with higher-order products including an approximation of volume-weighted equivalent particle densities, liquid equivalent snowfall, and rainfall rate estimates. The data underwent a rigorous standardization and quality assurance process to filter out erroneous observations to produce a self-describing, scalable, and achievable data set. Case study analyses demonstrate the capabilities of the data set in identifying physical processes like precipitation phase-changes at high temporal resolution. Bulk precipitation characteristics from a multi-site intercomparison also highlight distinct microphysical properties unique to each location. This curated PIP data set is a robust database of high-quality particle microphysical observations for constraining future precipitation retrieval algorithms, and offers new insights toward better understanding regional and seasonal differences in bulk precipitation characteristics.

54 ENVIRONMENTAL SCIENCES↗

Statistical Analysis of Trans‐Ionospheric Pulse Pairs and Inferences on Their Characteristics

Trans-ionospheric pulse pairs (TIPPs), first observed in 1993, are signatures of in-cloud lightning discharges observed by satellite-based broadband very high frequency (VHF) receivers. It has been definitively shown that TIPPs are the space-based signatures of compact intracloud discharges (CIDs), and that the associated pair of pulses that comprise a TIPP result from the direct VHF pulse from the discharge, followed by a pulse reflected from the Earth's surface. However, the ratio of the peak amplitudes of these two pulses can vary widely, with the second pulse often having considerably higher peak amplitude than the first. This observation has not been satisfactorily explained. Using data collected from geostationary orbit by the Radio Frequency Sensor (RFS) and matched to locations reported by the Global Lightning Dataset (GLD360), we assemble the largest database to date of 76,348 TIPPs with associated location, altitude, and amplitude ratio of the two pulses in the TIPP. We show that the amplitude ratio of TIPPs is strongly correlated to the altitude of the associated discharges and the geometry of the source location with respect to the Earth's surface and the receiver. These observations strongly suggest that the difference in amplitude of the two pulses is driven by a nondipole radiated beam pattern that is dependent on the polarity of the CID, velocity of the current wavefront, and viewing angle.

58 GEOSCIENCES↗

Geochemical Phosphorus Sequestration in Tundra Soils Impedes Delivery of Bioavailable Phosphorus to the Kuparuk River, Alaska, USA: Implications for the Broader Arctic Region

Long-term river monitoring of the Kuparuk River (North Slope, Alaska, USA) confirms significant increases in solutes that are indicative of active layer thickening due to thawing permafrost. However, there is no evidence of an increase in total dissolved phosphorus (TDP) or soluble reactive phosphorus (SRP), the nutrient that limits primary production in this and similar rivers in the region. Here, we show that Mehlich-3 extractable iron (Fe) and aluminum (Al) in active layer soils impart high P geochemical sorption capacities across a range of landscape features that we would expect to promote lateral movement of water and solutes to headwater streams in our study watershed. Reanalysis of a recently published pan-Arctic soils database that includes active layer and permafrost soil samples suggests that this high P sorption capacity could be common in other parts of the Arctic region. We conclude that soil minerals enhance P retention on hillslopes and propose pedogenic secondary Fe and Al minerals may continue to retain P in these soils and limit biological productivity in the adjacent river even as active layer thickening increases potential P mobility in the watershed. We suggest that similar interactions may occur in other areas of the Arctic where comparable geochemical conditions prevail.

Sutor, Frederick W. [Univ. of Vermont, Burlington,↗

A substitutional quantum defect in WS2 discovered by high-throughput computational screening and fabricated by site-selective STM manipulation

Abstract Point defects in two-dimensional materials are of key interest for quantum information science. However, the parameter space of possible defects is immense, making the identification of high-performance quantum defects very challenging. Here, we perform high-throughput (HT) first-principles computational screening to search for promising quantum defects within WS 2 , which present localized levels in the band gap that can lead to bright optical transitions in the visible or telecom regime. Our computed database spans more than 700 charged defects formed through substitution on the tungsten or sulfur site. We found that sulfur substitutions enable the most promising quantum defects. We computationally identify the neutral cobalt substitution to sulfur ( $${\rm{Co}}_{{{{{{{{\rm{S}}}}}}}}}^{0}$$ Co S 0 ) and fabricate it with scanning tunneling microscopy (STM). The $${\rm{Co}}_{{{{{{{{\rm{S}}}}}}}}}^{0}$$ Co S 0 electronic structure measured by STM agrees with first principles and showcases an attractive quantum defect. Our work shows how HT computational screening and nanoscale synthesis routes can be combined to design promising quantum defects.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Shifts in evolutionary lability underlie independent gains and losses of root-nodule symbiosis in a single clade of plants

Abstract Root nodule symbiosis (RNS) is a complex trait that enables plants to access atmospheric nitrogen converted into usable forms through a mutualistic relationship with soil bacteria. Pinpointing the evolutionary origins of RNS is critical for understanding its genetic basis, but building this evolutionary context is complicated by data limitations and the intermittent presence of RNS in a single clade of ca. 30,000 species of flowering plants, i.e., the nitrogen-fixing clade (NFC). We developed the most extensive de novo phylogeny for the NFC and an RNS trait database to reconstruct the evolution of RNS. Our analysis identifies evolutionary rate heterogeneity associated with a two-step process: An ancestral precursor state transitioned to a more labile state from which RNS was rapidly gained at multiple points in the NFC. We illustrate how a two-step process could explain multiple independent gains and losses of RNS, contrary to recent hypotheses suggesting one gain and numerous losses, and suggest a broader phylogenetic and genetic scope may be required for genome-phenome mapping.

59 BASIC BIOLOGICAL SCIENCES↗

Discovering type I cis-AT polyketides through computational mass spectrometry and genome mining with Seq2PKS

Type 1 polyketides are a major class of natural products used as antiviral, antibiotic, antifungal, antiparasitic, immunosuppressive, and antitumor drugs. Analysis of public microbial genomes leads to the discovery of over sixty thousand type 1 polyketide gene clusters. However, the molecular products of only about a hundred of these clusters are characterized, leaving most metabolites unknown. Characterizing polyketides relies on bioactivity-guided purification, which is expensive and time-consuming. To address this, we present Seq2PKS, a machine learning algorithm that predicts chemical structures derived from Type 1 polyketide synthases. Seq2PKS predicts numerous putative structures for each gene cluster to enhance accuracy. The correct structure is identified using a variable mass spectral database search. Benchmarks show that Seq2PKS outperforms existing methods. Applying Seq2PKS to Actinobacteria datasets, we discover biosynthetic gene clusters for monazomycin, oasomycin A, and 2-aminobenzamide-actiphenol.

60 APPLIED LIFE SCIENCES↗

Chemoproteogenomic stratification of the missense variant cysteinome

Abstract Cancer genomes are rife with genetic variants; one key outcome of this variation is widespread gain-of-cysteine mutations. These acquired cysteines can be both driver mutations and sites targeted by precision therapies. However, despite their ubiquity, nearly all acquired cysteines remain unidentified via chemoproteomics; identification is a critical step to enable functional analysis, including assessment of potential druggability and susceptibility to oxidation. Here, we pair cysteine chemoproteomics—a technique that enables proteome-wide pinpointing of functional, redox sensitive, and potentially druggable residues—with genomics to reveal the hidden landscape of cysteine genetic variation. Our chemoproteogenomics platform integrates chemoproteomic, whole exome, and RNA-seq data, with a customized two-stage false discovery rate (FDR) error controlled proteomic search, which is further enhanced with a user-friendly FragPipe interface. Chemoproteogenomics analysis reveals that cysteine acquisition is a ubiquitous feature of both healthy and cancer genomes that is further elevated in the context of decreased DNA repair. Reference cysteines proximal to missense variants are also found to be pervasive, supporting heretofore untapped opportunities for variant-specific chemical probe development campaigns. As chemoproteogenomics is further distinguished by sample-matched combinatorial variant databases and is compatible with redox proteomics and small molecule screening, we expect widespread utility in guiding proteoform-specific biology and therapeutic discovery.

Desai, Heta (ORCID:0000000343621707)↗

Parallel measurement of transcriptomes and proteomes from same single cells using nanodroplet splitting

Single-cell multiomics provides comprehensive insights into gene regulatory networks, cellular diversity, and temporal dynamics. Here, we introduce nanoSPLITS (nanodroplet SPlitting for Linked-multimodal Investigations of Trace Samples), an integrated platform that enables global profiling of the transcriptome and proteome from same single cells via RNA sequencing and mass spectrometry-based proteomics, respectively. Benchmarking of nanoSPLITS demonstrates high measurement precision with deep proteomic and transcriptomic profiling of single-cells. We apply nanoSPLITS to cyclin-dependent kinase 1 inhibited cells and found phospho-signaling events could be quantified alongside global protein and mRNA measurements, providing insights into cell cycle regulation. We extend nanoSPLITS to primary cells isolated from human pancreatic islets, introducing an efficient approach for facile identification of unknown cell types and their protein markers by mapping transcriptomic data to existing large-scale single-cell RNA sequencing reference databases. Accordingly, we establish nanoSPLITS as a multiomic technology incorporating global proteomics and anticipate the approach will be critical to furthering our understanding of biological systems.

59 BASIC BIOLOGICAL SCIENCES↗

BiG-SCAPE 2.0 and BiG-SLiCE 2.0: scalable, accurate and interactive sequence clustering of metabolic gene clusters

Microbial metabolic gene clusters encode the biosynthesis or catabolism of metabolites that facilitate ecological specialization, mediate microbiome interactions and constitute a major source of medicines and crop protection agents. Here, we present BiG-SCAPE and BiG-SLiCE 2.0, next-generation methods that facilitate scalable, accurate and interactive gene cluster analyses. BiG-SCAPE 2.0 updates its classification, alignment methods, and visualizations, enabling more accurate analysis, up to 8x faster runtimes and halved memory requirements. BiG-SLiCE 2.0 updates its distance metric, pHMM database, and classification logic, resulting in increased sensitivity nearing that of BiG-SCAPE. Analysis of 260,630 biosynthetic gene clusters from publicly available genomes reveals that both tools generate concurring estimates of gene cluster diversity, thus providing significantly extended methodological support for recent evidence indicating that the vast majority of natural product diversity remains unexplored. Together, these updates will facilitate global genome mining efforts for natural product discovery and microbiome analyses scalable with current data sizes.

Draisma, Arjan [Wageningen University & Research (↗

Isolation of genome-predicted Caldatribacterium ( Atribacterota ) reveals pervasive microbial cultivation problem due to folate precipitation

Most bacterial phyla have few or no pure cultures, including Atribacterota , comprised of ubiquitous anaerobes. Here, we report genome-guided enrichment and isolation of two Atribacterota species representing a new family, Caldatribacterium saccharofermentans from a hot spring, and Caldatribacterium inferamans from a deep aquifer. Both were co-enriched with sulfate-reducing bacteria and initially resisted isolation, which we link to inadvertent removal of precipitated folic acid by filter-sterilization of unbuffered Wolin’s vitamin solution. We then predict folate auxotrophy across the Atribacterota and ~29% of all bacteria, with extensive auxotrophy in 27% of phyla. Since ≥604 of 791 ( ≥ 76%) media with folic acid additions in the MediaDive database use unbuffered vitamin solutions in which folic acid is likely removed during filter-sterilization, we propose that folate auxotrophy limits culturability in defined media en masse. We also uncover unusual features of Caldatribacterium , including three lipid membrane-like layers (LMLs), with the inner LML surrounding the nucleoid, and a high percentage of secreted proteins, supporting a unique cell biology of Atribacterota .

Biological and medical sciences↗

Active learning of ternary alloy structures and energies

Abstract Machine learning models with uncertainty quantification have recently emerged as attractive tools to accelerate the navigation of catalyst design spaces in a data-efficient manner. Here, we combine active learning with a dropout graph convolutional network (dGCN) as a surrogate model to explore the complex materials space of high-entropy alloys (HEAs). We train the dGCN on the formation energies of disordered binary alloy structures in the Pd-Pt-Sn ternary alloy system and improve predictions on ternary structures by performing reduced optimization of the formation free energy, the target property that determines HEA stability, over ensembles of ternary structures constructed based on two coordinate systems: (a) a physics-informed ternary composition space, and (b) data-driven coordinates discovered by the Diffusion Maps manifold learning scheme. Both reduced optimization techniques improve predictions of the formation free energy in the ternary alloy space with a significantly reduced number of DFT calculations compared to a high-fidelity model. The physics-based scheme converges to the target property in a manner akin to a depth-first strategy, whereas the data-driven scheme appears more akin to a breadth-first approach. Both sampling schemes, coupled with our acquisition function, successfully exploit a database of DFT-calculated binary alloy structures and energies, augmented with a relatively small number of ternary alloy calculations, to identify stable ternary HEA compositions and structures. This generalized framework can be extended to incorporate more complex bulk and surface structural motifs, and the results demonstrate that significant dimensionality reduction is possible in thermodynamic sampling problems when suitable active learning schemes are employed.

Chemistry↗

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE↗

Feature engineering descriptors, transforms, and machine learning for grain boundaries and variable-sized atom clusters

Abstract Obtaining microscopic structure-property relationships for grain boundaries is challenging due to their complex atomic structures. Recent efforts use machine learning to derive these relationships, but the way the atomic grain boundary structure is represented can have a significant impact on the predictions. Key steps for property prediction common to grain boundaries and other variable-sized atom clustered structures include: (1) describing the atomic structure as a feature matrix, (2) transforming the variable-sized feature matrix to a fixed length common to all structures, and (3) applying a machine learning algorithm to predict properties from the transformed matrices. We examine how these steps and different combinations of engineered features impact the accuracy of grain boundary energy predictions using a database of over 7000 grain boundaries. Additionally, we assess how different engineered features support interpretability, offering insights into the physics of the structure-property relationships.

36 MATERIALS SCIENCE↗

Uncertainty quantification for misspecified machine learned interatomic potentials

The use of high-dimensional regression techniques from machine learning has significantly improved the quantitative accuracy of interatomic potentials. Atomic simulations can now plausibly target quantitative predictions in a variety of settings, which has brought renewed interest in robust means to quantify uncertainties. In many practical settings where model complexity is constrained (e.g., due to performance considerations), misspecification — the inability of any one choice of model parameters to exactly match all training data — is a key contributor to errors that is often disregarded. Here, we employ a recent misspecification-aware regression technique to quantify parameter uncertainties, which is then propagated to a broad range of phase and defect properties in tungsten. The propagation is performed through both brute-force resampling and implicit Taylor expansion. The propagated misspecification uncertainties robustly quantify and bound errors on a broad range of material properties. We demonstrate application to recent foundational machine learning interatomic potentials, accurately predicting and bounding errors in MACE-MPA-0 energy predictions across the diverse materials project database.

36 MATERIALS SCIENCE↗

Towards an open model intercomparison platform for integrated assessment models scenarios

The majority of scenarios in the IPCC database are generated by integrated assessment models (IAMs) and come from model intercomparison projects. However, the way in which the current model intercomparison projects are organized is not open to all IAM teams worldwide. Here we propose a transparent and inclusive platform that is open to anyone with an IAM regarding protocols development, scenario submissions and results evaluation. We discuss the challenges of this approach, particularly human resources and financial support. Here, we identify diversity in the level of model capability and quality of model output as possibly critical issues. Despite such challenges, the IAM community and its scientific activities can improve and benefit from the proposed platform, ultimately contributing to better climate policymaking.

IAM↗