Search NASA⌕ Search

SEARCH · Search NASA

Results for “data integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Solvent Transport in Disordered and Dynamic Membrane Pores: Implications for Reverse Osmosis and Nanofiltration Membranes

Pressure-driven separations with nanoporous membranes, such as reverse osmosis and nanofiltration, play a vital role in addressing water scarcity and enabling resource recovery. Understanding water or solvent transport in membrane pores is essential for advancing membrane separation technologies. A key question in transport modeling is to establish a relationship between solvent permeability and membrane porous structure properties, such as porosity or pore size. The nano- and subnanometer pores in polymeric membranes such as reverse osmosis and nanofiltration membranes are highly tortuous and dynamically connected, which challenges the conventional methods of transport modeling. This study addresses this challenge by developing a theoretical framework to describe solvent transport through membranes with dynamic and disordered porous structure. Specifically, we propose a lattice model to describe the pore network, while preserving the viscous nature of solvent permeation. We further establish a relationship between solvent permeability and membrane porosity or pore size, which is validated by molecular dynamics simulations and experimental data. By integrating this relationship into the solution-friction model, we define pore connectivity and local friction coefficient to quantify the impact of pore structure on solvent permeability. Our analysis highlights the dominant influence of pore connectivity on the permeability of reverse osmosis and nanofiltration membranes, particularly when the pore size approaches the dimensions of solvent molecules. Overall, this study provides critical insights into water and solvent transport mechanisms in nanoporous membranes, opening the door for strategies to substantially enhance membrane performance.

membranes↗

Machine-Learning-Driven Discovery of Water Splitting BaFe 2 O 4 and Human-in-the-Loop Improvement via Al-Substitution for Increased Thermal Stability

Thermochemical hydrogen (TCH) production offers a promising method for converting thermal energy into hydrogen fuel through heat-driven redox cycles of metal oxides. Here, in this work a defect graph neural network (dGNN) was used to predict oxygen vacancy formation energies ΔH V O combined with Materials Project predictions of oxygen chemical potential stability to screen candidate oxides via high-throughput database analysis. BaFe 2 O 4 was identified as a promising material for experimental validation based on its predicted ΔH V O , oxygen chemical potential stability range, and potential for tunable substitutions to improve thermal properties. Experimental validation using thermogravimetric analysis (TGA), stagnation flow reactor (SFR), X-ray diffraction (XRD), and electron microscopy confirmed positive water-splitting behavior but also revealed limitations in thermal stability under aggressive reduction conditions. To address this, a human-in-the-loop modification strategy was employed introducing Al substitution in BaFe 2–x Al x O 4 ; this modification improves thermal stability, alters the crystal structure and enhances overall performance. These results demonstrate a combined computational and experimental workflow in which machine learning accelerates identification of promising candidates, while targeted experimental design enables optimization of functional performance. This approach advances the development of robust, cost-effective TCH materials and highlights the importance of integrating data-driven discovery with human-guided materials design in paving the way for scalable hydrogen production technologies.

organic↗

Reduce Revenue Versus Increase Expenditure: Fires and Plant Invasion Drive Soil Carbon Loss With Different Mechanisms in a Mediterranean Shrubland

Fires and plant invasions threaten Mediterranean ecosystems substantially, particularly in the context of changing climate. Our study utilized a data-model integration approach to assess the response of soil organic carbon (SOC) to fires and plant invasion under three Shared Socio-Economic Pathway (SSP) scenarios (SSP1-26, SSP2-45, and SSP5-85). We parameterized the CLM-Microbe model and then investigated the individual and interactive impacts of fires and plant invasion on soil C by comparing factorial simulations of initialization (fire/no wildfire in 2021), fire module on/off, and with and without plant invasion during 2023–2100 in a Mediterranean ecosystem. The simulations indicated a marked C loss due to the 2021 wildfire, projected fires, and plant invasion across all future climate scenarios. Specifically, the 2021 wildfire, projected fires, and plant invasion reduced the SOC (0–30 cm) by 0.12, 0.26, and 0.15 kg C m −2 under SSP1-26, 0.12, 0.30, and 0.12 kg C m −2 under SSP2-45, and 0.12, 0.24, and 0.13 kg C m −2 under SSP5-85, respectively. However, fires and plant invasion decreased SOC through distinct mechanisms. The effects of the 2021 wildfire occurred due to its negative legacy on soil microbial community and, thus, litter accumulation, suppressing the formation of soil carbon via decomposition. Influences of projected fires happen via consuming fuel and suppressing carbon input to soils. In contrast, the impacts of plant invasions were due to enhanced microbial respiration, leading to C loss. In conclusion, these findings emphasize the need for tailored C sequestration strategies considering the disparate effects of fires and plant invasions in the Mediterranean climate.

microbe↗

Emerging multiscale insights on microbial carbon use efficiency in the land carbon cycle

Microbial carbon use efficiency (CUE) affects the fate and storage of carbon in terrestrial ecosystems, but its global importance remains uncertain. Accurately modeling and predicting CUE on a global scale is challenging due to inconsistencies in measurement techniques and the complex interactions of climatic, edaphic, and biological factors across scales. The link between microbial CUE and soil organic carbon relies on the stabilization of microbial necromass within soil aggregates or its association with minerals, necessitating an integration of microbial and stabilization processes in modeling approaches. In this perspective, we propose a comprehensive framework that integrates diverse data sources, ranging from genomic information to traditional soil carbon assessments, to refine carbon cycle models by incorporating variations in CUE, thereby enhancing our understanding of the microbial contribution to carbon cycling.

54 ENVIRONMENTAL SCIENCES↗

Labels as a feature: Network homophily for systematically annotating human GPCR drug-target interactions

Machine learning has revolutionized drug discovery by enabling the exploration of vast, uncharted chemical spaces essential for discovering novel patentable drugs. Despite the critical role of human G protein-coupled receptors in FDA-approved drugs, exhaustive in-distribution drug-target interaction testing across all pairs of human G protein-coupled receptors and known drugs is rare due to significant economic and technical challenges. This often leaves off-target effects unexplored, which poses a considerable risk to drug safety. In contrast to the traditional focus on out-of-distribution exploration (drug discovery), we introduce a neighborhood-to-prediction model termed Chemical Space Neural Networks that leverages network homophily and training-free graph neural networks with labels as features. We show that Chemical Space Neural Networks’ ability to make accurate predictions strongly correlates with network homophily. Thus, labels as features strongly increase a machine learning model’s capacity to enhance in-distribution prediction accuracy, which we show by integrating labeled data during inference. We validate these advancements in a high-throughput yeast biosensing system (3773 drug-target interactions, 539 compounds, 7 human G protein-coupled receptors) to discover novel drug-target interactions for FDA-approved drugs and to expand the general understanding of how to build reliable predictors to guide experimental verification.

Hansson, Frederik G↗

Unveiling a pervasive DNA adenine methylation regulatory network in the early-diverging fungus Rhizopus microsporus

Development of the DNA affinity purification and sequencing (DAP-seq) technique has allowed genome-scale studies of transcription factor (TF)-binding sites with high reproducibility. Here, we apply this technique to the human opportunistic pathogen Rhizopus microsporus, a mucoralean fungus belonging to the understudied group of early-diverging fungi. We characterize genome-wide binding sites of 58 TFs encoded by genes regulated through adenine methylation and representing major TF families. This analysis reveals their binding profiles and recognized sequences, expanding and diversifying the catalog of known fungal motifs. By integrating this data with DNA 6-methyladenine profiling, we uncover the extensive direct and indirect impact of this epigenetic modification on the regulation of gene expression. Furthermore, we use the generated data to identify TFs involved in biologically relevant processes such as zinc metabolism and light response. Our work enhances our understanding of regulatory mechanisms in R. microsporus and provides broader insights into gene regulation across the fungal kingdom.

Lax, Carlos [Universidad de Murcia (Spain)] (ORCID↗

Cracking the code of multi-layer films to promote circularity in single-use plastic packaging

Multi-layer film packaging (MLF) revolutionized food preservation by combining diverse material layers to optimize barrier properties, mechanical strength, and shelf-life. These materials are essential for transporting perishables across various climates and allow for access to fresh goods in “food deserts”, but they pose significant recycling challenges due to their structural complexity. This perspective examines key structure-property relationships governing barrier performance and highlights innovations in material design. We explore how machine learning can predict performance metrics and propose recyclable alternatives, integrating data-driven approaches with material science insights. By challenging the status quo of MLF design, we advocate for circularity in food packaging, inspiring innovation at the intersection of sustainability, material science, and artificial intelligence.

36 MATERIALS SCIENCE↗

Quantifying motion blur by imaging shock front propagation with broadband and narrowband X-ray sources

Time-integrated radiography using MeV Bremsstrahlung X-ray sources is the norm for imaging during system-level testing of components and structures under dynamic condition. One source of error in the analysis of the time-integrated radiography data sets stems from motion blur which smears out sharp interfaces to a greater degree with longer exposure times, which become necessary to provide sufficient signal-to-noise with low X-ray penetration of objects of interest. To quantify motion blur, a 1D shock wave through PMMA was investigated experimentally at The Dynamic Compression Sector at The Advanced Photon Source (DCS@APS) with tapered broadband and 25.46 ± 1.06 keV narrowband X-rays. Four cameras with different exposure times were used for each experiment to compare the effect that exposure time has on motion blur. In addition, our methodology to accurately simulate motion blur in terms of transmission and shape is presented and compared to our experimental results and quantified. There is a high level of agreement between the experimental and simulation results across the range of data sets investigated in this study with a percent difference range of 0.29–1.31% for the four shots. The methodology of this work serves as a steppingstone towards a physically validated model that could be used in conjunction with experimental results to deconvolve physical parameters, densities, and interfaces of interest in a way that would not be possible with experimental results alone.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Energetic and structural control of polyspecificity in a multidrug transporter

Multidrug efflux pumps are dynamic molecular machines that drive antibiotic resistance by harnessing ion gradients to export chemically diverse substrates. Despite their clinical importance, the molecular principles underlying multidrug promiscuity and energy efficiency remain poorly understood. Using multiparametric deep mutational scanning across eight substrates and two energy conditions, we deconvolute the contributions of substrate recognition, energetic coupling, and protein stability, providing an integrated, high-resolution view of multidrug transport. We find that substrate specificity arises from a distributed network of residues extending beyond the binding site, with mutations that reshape binding, coupling, conformational flexibility, and membrane interactions. Further, we apply a pH-based selection scheme to measure the effect of mutation on pH-dependent transport efficiency. By integrating these data, we reveal a fundamental relationship between efficiency and promiscuity: Highly efficient variants exhibit broad substrate profiles, while inefficient variants are narrower. In conclusion, these findings establish a direct link between energy coupling and polyspecificity, uncovering the biochemical logic underlying multidrug transport.

Biological Sciences↗

Explaining System-Level Prognostics with Established Machine Learning Methods

System-level prognostics is crucial for ensuring reliability and enabling predictive maintenance in complex systems with interconnected components. This study presents a framework that integrates data-driven methods to predict the remaining useful life (RUL) of a subsystem under multiple and concurrent faults within a nuclear power plant system with explainable artificial intelligence (XAI). A nuclear power plant (NPP) operation was simulated to model the degradation behavior of NPP components, and four machine learning models—Gradient Boosting Regressor (GBR), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory (LSTM)—were evaluated for prognostics with a novel system RUL parameter. The LSTM model demonstrated potential superior repeatability, while SHAP (SHapley Additive exPlanations) for explainability provided consistent and trustworthy global explanations. In contrast, LIME (Local Interpretable Model-agnostic Explanations) offered localized interpretability but showed reduced stability for sequential data. Key findings include the interplay between component-level degradation and system-wide performance, with LSTM effectively capturing these dynamics through sequence-level predictions. The XAI techniques enhanced transparency by identifying critical features influencing model predictions and aligning with domain knowledge. Furthermore, this framework has significant implications for improving trust and understanding in predictive maintenance, particularly in safety-critical industries like nuclear energy.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Validating the galaxy and quasar catalog-level blinding scheme for the DESI 2024 analysis

In the era of precision cosmology, ensuring the integrity of data analysis through blinding techniques is paramount — a challenge particularly relevant for the Dark Energy Spectroscopic Instrument (DESI). DESI represents a monumental effort to map the cosmic web, with the goal to measure the redshifts of tens of millions of galaxies and quasars. Given the data volume and the impact of the findings, the potential for confirmation bias poses a significant challenge. To address this, we implement and validate a comprehensive blind analysis strategy for DESI Data Release 1 (DR1), tailored to the specific observables DESI is most sensitive to: Baryonic Acoustic Oscillations (BAO), Redshift-Space Distortion (RSD) and primordial non-Gaussianities (PNG). We carry out the blinding at the catalog level, implementing shifts in the redshifts of the observed galaxies to blind for BAO and RSD signals and weights to blind for PNG through a scale-dependent bias. We validate the blinding technique on mocks as well as on data by applying a second blinding layer to perform a series of sanity checks; the latter allows probing complexities in real data not captured in mocks. We find that the blinding strategy alters the data vector in a controlled way, and the BAO and RSD analysis choices are robust to blinding. The successful validation of the blinding strategy paves the way for the unblinded DESI DR1 analysis, alongside future blind analyses with DESI and other surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Induced and natural variation affect traits independently in hybrid Populus

Abstract The genetic control of many plant traits can be highly complex. Both allelic variation (sequence change) and dosage variation (copy number change) contribute to a plant's phenotype. While numerous studies have investigated the effect of allelic or dosage variation, very few have documented both within the same system, leaving their relative contribution to phenotypic effects unclear. The Populus genome is highly polymorphic, and poplars are fairly tolerant of gene dosage variation. Here, using a previously established Populus hybrid F1 population, we assessed and compared the effect of natural allelic variation and induced dosage variation on biomass, phenology, and leaf morphology traits. We identified QTLs for many of these traits, but our results indicate limited overlap between the QTLs associated with natural allelic variation and induced dosage variation. Additionally, the integration of data from both allelic and dosage variation identifies a larger set of QTLs that together explain a larger percentage of the phenotypic variance. Finally, our results suggest that the effect of the large indels might mask that of allelic QTLs. Our study helps clarify the relationship between allelic and dosage variation and their effects on quantitative traits.

Guo, Weier (ORCID:0000000251789334)↗

Diaspora: Resilience-Enabling Services for Real-Time Distributed Workflows

The need for real-time processing to enable automated decision making and experimental steering has driven a shift from high-performance computing workflows on a centralized system to a distributed approach that integrates remote data sources, edge devices, and diverse compute facilities. Under this paradigm, data can be processed close to the source where it is generated, thus reducing latency and bandwidth usage. System resilience is thus a key challenge, requiring distributed workflows to survive component failures and to meet stringent quality-of-service requirements, which results in the need to mitigate anomalies such as congestion and low availability of resources. To address these challenges, we propose Diaspora, a unified resilience framework that is inspired by event-driven communication patterns used in public clouds. Specifically, we propose an event fabric that extends across sites, facilities, and computations to provide timely, reliable, and accurate information about data, application, and resource status. On top of the event fabric, we build resilience-enabling services that combine QoS-aware data streaming, resilient data views, resilient compute and data resources, and anomaly detection and prediction, all of which collectively enhance workflow resilience for these scientific cases.

Rao, Nageswara↗

Structure-aware Initialization via Numerical Continuation and Informed Priors

Scientific machine learning (SciML) often operates in ill-conditioned, weakly identifiable regimes due to limited data or indirect observations. In such settings, optimization and inference are highly sensitive to the starting point, making initialization--often under-reported--a consequential degree of freedom. Random initialization is not a neutral default as it induces an implicit prior over candidate solutions and can systematically bias the result, producing large run-to-run variability. Here, we formalize this view by treating initialization as a hidden confounder in SciML and develop a unifying theory for structure-aware initialization via numerical continuation, constructing warm starts from related problem instances. Across representative tasks, including physics-informed neural networks, maximum likelihood estimation, and variational inference, warm starts have been shown to consistently reduce optimization effort and improve reliability.

Data integrity↗

Measured Long-Term Solar Irradiance for Climate Studies

Solar radiation is primarily measured using high-quality radiometers (e.g., pyranometers and pyrheliometers). These instruments need to be calibrated regularly (every two years at a minimum). They are also susceptible to various sources of uncertainties. Therefore, rigorous data quality assessment is required to obtain high-confidence data from these radiometers. This is especially true when the data is used to understand climatic trends and/or extreme weather events. In this study, we attempted to create a continuous and reliable dataset by correcting the underlying data which we believe contains bias due to reference pyranometers swap procedures. These biases can be significant which can be up to two percentage points thereby influencing the interpretation of climate changes and/or extreme events. This study elaborates on the biases and methods to correct those biases.

data integrity↗

Autonomous Anomaly Detection For Continuous Streams

The code implements the Isolation Forest (IFML) algorithm within the digital twin (DT) of the AGN-201 nuclear reactor. The DT captures real-time operational data including control rod positions, reactor power, and temperature. The IFML model isolates anomalies by detecting patterns that deviate from expected operational behavior. The algorithm recursively partitions the data and assigns anomaly scores based on the isolation of rare and different events. By tuning parameters specific to the reactor’s operational data, the IFML identifies deviations such as unauthorized material insertions or reactor reactivity shifts. The system streams data using LabView and integrates with the DeepLynx data warehouse for anomaly processing.

Trevino, Eduardo↗

Multi-omic characterization of a soil microbial consortium reveals critical role of succinate and glutamate metabolism during calcium carbonate precipitation

Microbially induced calcium carbonate precipitation (MICP) holds potential for use in soil stabilization and carbon sequestration, with the overall efficiency of the process being a major determinant for use in many environmental and civil engineering applications. While the biogeochemical pathways and enzymes driving MICP are known, the microbial metabolic networks and community dynamics underlying such precipitation remain poorly characterized. To address this gap, we developed a four-member consortium of soil bacteria (Curtobacterium flaccumfaciens, Rhodococcus qingshengii, Microbacterium sp., and Bacillus toyonensis), termed carbon storing consortium - A (CSC-A), that is capable of MICP. Prior work shows that MICP production is higher in CSC-A compared to the sum of carbonate produced by each member, suggesting carbonate production is driven by consortium dynamics. To that end we used a multi-omic integration approach of genomics, transcriptomics, and metabolomics to investigate potential inter-species interactions that may influence the MICP phenotype. Genomic life history characterizations identified evidence of niche specialization by B. toyonensis and Microbacterium, while metatranscriptomic analysis suggests R. qingshengii is a keystone species during growth in urea. By comparing individual species’ metabolomes to the metabolic profile of a shared well of precipitated metabolites, we identified over 200 metabolites predicted to be produced or consumed by CSC-A members. Integrating both data types to search the KEGG reactome highlighted a network centered around glutamine metabolism and branched chain amino acid biosynthesis under regulation during CSC-A growth in urea. Succinate metabolism was also a major node in this network and laboratory assays confirmed that increasing the amount of succinate in the growth medium leads to increased carbonate precipitation by CSC-A, a critical confirmation of our modeling approach. By isolating and identifying the interconnected metabolic components underlying MICP in CSC-A, we identified keystone taxa, metabolites, and pathways important for future optimization of the application of this consortia to carbonate precipitation.

carbon storing consortium - A (CSC-A)↗