Search NASA⌕ Search

SEARCH · Search NASA

Results for “combined data integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The Monarch Initiative in 2024: an analytic platform integrating phenotypes, genes and diseases across species

Abstract Bridging the gap between genetic variations, environmental determinants, and phenotypic outcomes is critical for supporting clinical diagnosis and understanding mechanisms of diseases. It requires integrating open data at a global scale. The Monarch Initiative advances these goals by developing open ontologies, semantic data models, and knowledge graphs for translational research. The Monarch App is an integrated platform combining data about genes, phenotypes, and diseases across species. Monarch's APIs enable access to carefully curated datasets and advanced analysis tools that support the understanding and diagnosis of disease for diverse applications such as variant prioritization, deep phenotyping, and patient profile-matching. We have migrated our system into a scalable, cloud-based infrastructure; simplified Monarch's data ingestion and knowledge graph integration systems; enhanced data mapping and integration standards; and developed a new user interface with novel search and graph navigation features. Furthermore, we advanced Monarch's analytic tools by developing a customized plugin for OpenAI’s ChatGPT to increase the reliability of its responses about phenotypic data, allowing us to interrogate the knowledge in the Monarch graph using state-of-the-art Large Language Models. The resources of the Monarch Initiative can be found at monarchinitiative.org and its corresponding code repository at github.com/monarch-initiative/monarch-app.

60 APPLIED LIFE SCIENCES↗

Hydropower Resilience Database for Assessing Microgrid Formation Capability and Enhancing Power Grid Resilience

In the face of increasing frequency of extreme events, enhancing resilience and reliability of energy infrastructure demands innovative solutions. Hydropower, with its inherent generation flexibility and grid-forming capability, offers significant potential for enhancing power grid resilience through the establishment of microgrids. To enable informed decision-making and strategic planning, we present the Hydropower Resilience Database (HRD) that integrates data from various sources such as Oakridge National Laboratory (ORNL) HydroSource, National Inventory of Dams, and US Western Grid database. By integrating relevant information on hydropower plant characteristics, including dams, reservoirs, and connected electrical grid, this resource enables hydropower plant owners, utilities, and community stakeholders to identify and evaluate the feasibility of using hydropower resources in microgrids to support nearby communities and critical infrastructure in various challenging scenarios. Using HRD, a set of metrics are evaluated to quantify the capability of hydropower plants to support essential microgrid functions. An interactive tool is developed using ArcGIS, enabling the visualization and analysis using HRD. Our research contributes to strategic planning efforts for fortifying energy infrastructure, ensuring reliable power supply during disruptions, and advancing the development of robust and resilient energy systems in regions susceptible to grid vulnerabilities.

13 HYDRO ENERGY↗

Hydropower Resilience Database for Assessing Microgrid Formation Capability and Enhancing Power Grid Resilience

In the face of increasing frequency of extreme events, enhancing resilience and reliability of energy infrastructure demands innovative solutions. Hydropower, with its inherent generation flexibility and grid-forming capability, offers significant potential for enhancing power grid resilience through the establishment of microgrids. To enable informed decision-making and strategic planning, we present the Hydropower Resilience Database (HRD) that integrates data from various sources such as Oakridge National Laboratory (ORNL) HydroSource, National Inventory of Dams, and US Western Grid database. By integrating relevant information on hydropower plant characteristics, including dams, reservoirs, and connected electrical grid, this resource enables hydropower plant owners, utilities, and community stakeholders to identify and evaluate the feasibility of using hydropower resources in microgrids to support nearby communities and critical infrastructure in various challenging scenarios. Using HRD, a set of metrics are evaluated to quantify the capability of hydropower plants to support essential microgrid functions. An interactive tool is developed using ArcGIS, enabling the visualization and analysis using HRD. Our research contributes to strategic planning efforts for fortifying energy infrastructure, ensuring reliable power supply during disruptions, and advancing the development of robust and resilient energy systems in regions susceptible to grid vulnerabilities.

13 HYDRO ENERGY↗

Microwave-Assisted Plastic Upcycling: Dynamic Data Reconciliation, Parameter Estimation, and Kinetic Modeling

Microwave (MW)-assisted catalytic pyrolysis offers a promising pathway for efficient plastic upcycling. This work develops an integrated modeling framework combining dynamic data reconciliation, a temperature-dependent rate model, and a yield model to represent the time-varying production rate of components in MW-assisted LDPE pyrolysis conducted in a batch reactor. An Arrhenius-type rate model with a temperature-dependent reaction order is developed. A biexponential correlation is proposed for the yield of gaseous products that enables to capture the evolving product formation behavior during conversion. In the yield correlation, one term is used to represent the initial increase in yield, reflecting the rapid formation of intermediate or primary products at the early stages of the reaction when a larger fraction of the reactant remains available. As conversion progresses, the influence of this term gradually diminishes. The other term accounts for the subsequent decrease in the predicted yield, representing secondary reactions such as further cracking or coke formation that reduce the concentration of certain products at higher conversion. The model is found to accurately represent reconciled experimental flow rate profiles from an in-house MW-assisted catalytic batch reactor for major products, including ethylene, ethane, 1-butene, and benzene, across 250−350 °C. Ethylene remains the dominant product but decreases from about 41.95% at 250 °C to 30.14% at 350 °C, while heavier products increase significantly, with 1-butene rising to nearly 8.37% and benzene reaching 2.17% at intermediate temperatures. The model shows that the ethylene production rate can be maximized at around 270 °C. The models developed in this work can be utilized for process optimization, reactor design and scale-up of microwave-assisted plastic conversion technologies, and economic analysis.

Damahe, Harish [West Virginia Univ., Morgantown, W↗

BLADE: An Automated Framework for Classifying Light Curves from the Center for Near-Earth Object Studies Fireball Database

Fireballs (bolides) are high-energy luminous phenomena produced when meteoroids and small asteroids enter Earth’s atmosphere at hypersonic speeds, often resulting in fragmentation or complete disintegration accompanied by significant energy release. The resulting bolide light curves capture temporal brightness variations as these objects traverse increasingly dense atmospheric layers, providing essential information on meteoroid entry dynamics, fragmentation behavior, and atmospheric energy deposition processes. The Center for Near-Earth Object Studies’ (CNEOS) continuously expanding fireball database offers a globally comprehensive archive of bolide events, including light curves and associated metadata. Events associated with infrasound detections allow direct correlations between acoustic signatures and light curve features, therefore enabling detailed analyses of fragmentation dynamics and energy deposition. Here, we introduce Bolide Light-curve Analysis and Discrimination Explorer (BLADE), a robust and high-fidelity framework specifically designed to analyze bolide light curves for objects detected from space. BLADE incorporates a processing pipeline integrating Savitzky–Golay filtering, prominence-based peak detection, and gradient analysis, enabling systematic identification and classification of fragmentation events and their associated energy release characteristics. Preliminary results demonstrate that BLADE reliably distinguishes distinct bolide behaviors, providing an objective, scalable methodology for characterization and analysis of large bolide light curve data sets. This foundational work establishes a novel pathway for advanced bolide research, with promising applications in planetary defense and global atmospheric monitoring. Future research should adopt an integrative approach combining CNEOS optical data with complementary infrasound measurements, further clarifying relationships between bolide energy deposition and acoustic signatures, thus refining our understanding of meteoroid and asteroid atmospheric entry processes.

Asteroids↗

Illuminating the Material World: Autonomous Microscopy to Understand Order, Disorder, and Everything In Between

Artificial intelligence (AI) holds immense promise for revolutionizing microscopy, yet its widespread adoption has been hindered by challenges ranging from user inexperience to limited model transferability and difficulties in operationalizing machine learning. This presentation showcases our approach to developing practical autonomy for materials discovery, aiming to accelerate the integration of AI into everyday microscopy workflows. As shown in Fig. 1, I will focus on three key areas: understanding order-disorder transitions, quantifying point defects, and achieving truly device-scale microscopy. First, I will demonstrate the power of multi-modal knowledge graphs for integrating diverse microscopy data. By combining imaging, spectroscopy, and diffraction data, these graphs provide a holistic view of material behavior, capturing the intricate relationships between different modalities [1,2]. I will present a case study on how these models illuminate the structural and chemical changes associated with irradiation in oxide thin films, revealing critical insights for designing materials for extreme environments like spaceflight and nuclear energy. Specifically, I will show how multi-modal analysis clarifies the evolution of order-disorder transitions under irradiation, a key factor influencing material performance in these applications. Next, I will address the challenge of quantifying point defects in 2D materials. We demonstrate the application of computer vision and transfer learning to accurately identify and classify various defect types, such as vacancies and substitutional atoms, and to quantify their concentrations. This information is crucial for understanding and tailoring the properties of 2D materials for applications in electronics, optoelectronics, and catalysis. For example, I will show how our models can characterize the topological distribution of point defects in MXene transition metal carbides, providing valuable insights for optimizing their performance in energy storage and separation science. Finally, I will discuss our progress toward autonomous device-scale microscopy [3,4]. We are fundamentally redesigning electron microscopes around the principles of machine reasoning, enabling automation beyond basic tasks like sample navigation and data acquisition to include sophisticated experimental design. This approach paves the way for truly reproducible and massively scaled analysis campaigns. I will emphasize the importance of autonomous microscopy platforms for high-throughput materials discovery and characterization, facilitating the rapid screening of materials for a broad range of applications and accelerating the development of next-generation technologies.

36 MATERIALS SCIENCE↗

Predicting Trends in VOC Through Rapid, Multimodal Characterization of State-of-the-Art p-i-n Perovskite Devices

Perovskite photovoltaic technologies are approaching commercial deployment, yet single junction and tandem architectures both still have significant room to improve power conversion efficiency and stability. The ability to perform rapid screening of material quality after altering processing conditions is critical to accelerating the optimization and commercialization of perovskite-based technologies. Currently, researchers utilize a wide range of stand-alone metrology tools to isolate sources of power loss throughout a device stack, which can be slow and labor intensive. Here, we demonstrate the use of a multimodal metrology approach to rapidly determine the maximum achievable and predicted open circuit voltages of >100 perovskite devices during fabrication. Acquisition of these different data is facilitated by combining them into a single integrated measurement platform. We show that these data and automated analysis can be used to rapidly understand and ultimately predict quantitative trends in open circuit voltages of state-of-the-art device architectures. The data and automated analysis workflow presented provides a reliable approach to quickly identify absorber and charge transport layer combinations that can lead to improved open circuit voltages.

14 SOLAR ENERGY↗

Refinery Integration Analysis: Pathways, Challenges, and Opportunities

Integrating biomass-derived intermediates into traditional petroleum refineries presents unique challenges and opportunities, requiring innovative analysis approaches that account for biofuel producers and refiners. Consequently, teams within the National Renewable Energy Laboratory (NREL) have developed a comprehensive refinery integration analysis framework that combines experimental data, detailed techno-economic analyses of bio-conversion pathways, and economic projections within refinery linear programming (LP) optimization models. These models enable the identification of promising bio-integration strategies tailored to specific refinery configurations, economic conditions, and production goals. They also capture upstream and downstream impacts, highlighting critical bottlenecks and research opportunities for increasing the share of biogenic feedstocks in traditional refining operations. Refinery models are also packaged into a broader bio-economy optimization framework which enables biofuel supply chain optimization with standalone biofuel production and refinery co-processing/repurposing options. This presentation discusses promising refinery bio-integration strategies along with key challenges and opportunities identified using NREL's refinery and bio-economy optimization frameworks.

09 BIOMASS FUELS↗

Status Report of Joint RPI/ORNL NCSP Task for Thermal Neutron Total Cross Section Measurements

A new cold neutron moderator system was developed and constructed at the Rensselaer Polytechnic Institute (RPI) Gaerttner Linear Accelerator. This system was then used to measure the neutron total cross section in the thermal energy range for polyethylene, polystyrene, Lucite, and yttrium hydride. In tandem with these measurements, a fitting procedure was developed at Oak Ridge National Laboratory to vary computationally simulated phonon density of states until a fit with differential and integral data was achieved. This model, combined with the new RPI total cross section data, was used to generate thermal scattering law files for the materials measured at RPI.

36 MATERIALS SCIENCE↗

EvoNet: A phylogenomic and systems biology approach to identify genes underlying plant survival in marginal, low‐N soils

The DOE‐BER “EvoNet” project investigates the genetic and molecular basis of plant resilience in extreme environments. We do this by identifying key genes that enable “extreme survivor” species to thrive in the nitrogen-poor soils of Chile’s hyper-arid Atacama Desert. Our collections focus on 32 Atacama extremophile species, including seven grass species with potential biofuel applications. To identify genes-of-importance to survival we compared genomic and transcriptomic profiles of extremophile species that thrive in the Atacama to those of closely related “sister” species from nitrogen-rich arid and mesic regions of California. Deep RNA sequencing and de novo transcriptome assembly across these triplet species sets supported a phylogenomic framework for identifying positively selected genes associated with adaptive divergence. Our integrative analysis combined ecological and environmental data, metagenomics, evolutionary and systems biology, and metabolomics. This enabled us to create an unprecedented framework for systematically understanding how non-model plants have adapted to survive in extreme conditions. Our resulting database of positively selected ortholog groups in the extremophile plants offers promising targets for engineering crop and biofuel species with enhanced resilience to drought and extreme weather. Additionally, our newest dataset explores and exploits a complementary metabolomic approach. This new aspect provides innovative strategies to manipulate plant cell metabolism, further supporting efforts to improve agricultural productivity in the face of extreme climates. Importantly, our combined evolutionary- and metabolomic-based strategies focused on convergent patterns of adaptation, providing a genetic and metabolomic toolkit for improving crop and biofuel resilience across diverse plant species. Finally, our novel exploration of ecological and evolutionary dynamics delivered to the community a phylogenomic computational pipeline called “PhyloGeneious.” Our continued adaptations of this pipeline are publicly available to expedite evolutionary genomic research for future scientific discoveries. In total, our DOE-BER has provided genomic, metabolomic, and computational strategies to understand how extremophile plants provide evolutionary and physiological targets for improving agricultural and biofuel production.

59 BASIC BIOLOGICAL SCIENCES↗

Evidence for the charge asymmetry in $pp$ → $t\bar{t}$ production at $\sqrt{s}$ = 13 TeV with the ATLAS detector

Inclusive and differential measurements of the top–antitop ($t\bar{t}$) charge asymmetry ${A}_{C}^{t\bar{t}}$ and the leptonic asymmetry ${A}_{C}^{\ell\bar{\ell}}$ are presented in proton–proton collisions at $\sqrt{s}$ = 13 TeV recorded by the ATLAS experiment at the CERN Large Hadron Collider. The measurement uses the complete Run 2 dataset, corresponding to an integrated luminosity of 139 fb –1 , combines data in the single-lepton and dilepton channels, and employs reconstruction techniques adapted to both the resolved and boosted topologies. A Bayesian unfolding procedure is performed to correct for detector resolution and acceptance effects. The combined inclusive $t\bar{t}$ charge asymmetry is measured to be ${A}_{C}^{t\bar{t}}$ = 0.0068 ± 0.0015, which differs from zero by 4.7 standard deviations. Differential measurements are performed as a function of the invariant mass, transverse momentum and longitudinal boost of the $t\bar{t}$ system. Both the inclusive and differential measurements are found to be compatible with the Standard Model predictions, at next-to-next-to-leading order in quantum chromodynamics perturbation theory with next-to-leading-order electroweak corrections. The measurements are interpreted in the framework of the Standard Model effective field theory, placing competitive bounds on several Wilson coefficients.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Integrating Machine Learning Potential and X-ray Absorption Spectroscopy for Predicting the Chemical Speciation of Disordered Carbon Nitrides

Precise determination of atomic structural information in functional materials holds transformative potential and broad implications for emerging technologies. Spectroscopic techniques, such as X-ray absorption near-edge structure (XANES), have been widely used for material characterization; however, extracting chemical information from experimental probes remains a significant challenge, particularly for disordered materials. We present an integrated approach that combines atomic simulations, data-driven techniques, and experimental measurements to investigate chemical speciation of amorphous carbon nitride systems as a case study. Here, we discuss the development of machine learning potentials that can efficiently explore the vast configuration space of amorphous carbon nitrides. By employing statistical methods, this structural database enables the elucidation of the most representative local structures and how they evolve with chemical compositions and density. Density functional theory simulations are used to establish a correlation between the local structure and spectroscopic signatures, which then serve as the basis for interpreting and extracting chemical content from experimental data. Although our framework is specifically demonstrated for XANES and carbon nitrides, the approach described herein is readily adaptable as applied to other experimental characterization probes and materials classes.

36 MATERIALS SCIENCE↗

On-the-fly data set combinations with RNTuple

With the expected data volume increase for HL-LHC and the even more complex computing challenges set by future colliders, the need for efficient data storage and processing becomes more pressing. ROOT’s next-generation data format and I/O subsystem, RNTTuple, is designed to address these challenges. RNTTuple already demonstrates a clear improvement in storage and I/O efficiency, as well as overall stability and robustness with respect to its predecessor, TTTree. These improvements provide a solid baseline to introduce novel extensions to common high-energy and nuclear physics (HENP) workflows. Notably, many workflows could benefit from the ability to arbitrarily join and chain data set samples at runtime, which could reduce overall storage requirements and improve application runtime and ergonomics. In this paper, we present the RNTupleProcessor, which enables HENP data set combinations with RNTuple. We will discuss the main design considerations, present the interfaces to support data set combinations and show how they integrate in typical workflows.

de Geus, Florine Willemijn [CERN; Twente U., Ensch↗

CrossMP: Enabling Cross-Modality Translation between Single-Cell RNA-Seq and Single-Cell ATAC-Seq through Web-Based Portal

In recent years, there has been a growing interest in profiling multiomic modalities within individual cells simultaneously. One such example is integrating combined single-cell RNA sequencing (scRNA-seq) data and single-cell transposase-accessible chromatin sequencing (scATAC-seq) data. Integrated analysis of diverse modalities has helped researchers make more accurate predictions and gain a more comprehensive understanding than with single-modality analysis. However, generating such multimodal data is technically challenging and expensive, leading to limited availability of single-cell co-assay data. Here, we propose a model for cross-modal prediction between the transcriptome and chromatin profiles in single cells. Our model is based on a deep neural network architecture that learns the latent representations from the source modality and then predicts the target modality. It demonstrates reliable performance in accurately translating between these modalities across multiple paired human scATAC-seq and scRNA-seq datasets. Additionally, we developed CrossMP, a web-based portal allowing researchers to upload their single-cell modality data through an interactive web interface and predict the other type of modality data, using high-performance computing resources plugged at the backend.

59 BASIC BIOLOGICAL SCIENCES↗

PDB-IHM: A System for Deposition, Curation, Validation, and Dissemination of Integrative Structures

Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.

IHMCIF↗

Near-Real-Time Material Tracking: Combining Vis–NIR Spectroscopy with Flow Sensing for Accurate Nd(III) Quantification

A fiber-optic visible–near-infrared (vis–NIR) absorption spectroscopy and flow sensor system has been developed for near-real-time tracking of Nd mass in the effluent stream from a column in a fume hood. The approach leverages two unique data streams and a partial least-squares regression (PLSR) model trained on vis–NIR absorption spectra of Nd(III) (0–1.5 M) in 1 M HNO 3 . In-line volumetric flow rate and vis–NIR spectra are measured in sequence after a chromatography column. The time stamps from each data stream are then synchronized, which allows integrated volumes to be combined with Nd(III) molarities predicted by a PLSR model to accurately calculate the Nd mass flowing through the column. This integrated measurement provides instantaneous mass flow and accumulates these data over time to obtain the total mass processed. The methodology developed in this study contributes critical technical infrastructure to improve monitoring capabilities to support chemical separations and the production of strategic materials and isotopes.

Irvine, Sawyer B. [Oak Ridge National Laboratory (↗

Machine Learned Empirical Numerical Integrator from Simulated Data

Recently, a number of state-of-the-art surrogate machine learning (ML) models have been designed for global weather and climate prediction, which have been trained using reanalysis data products. Reanalysis data products are constructed using numerical model simulations that combine numerical integration of partial differential equations and parameterization schemes. These products are typically only archived and made available using coarsened spatial and temporal resolutions. This study explores the impact of the numerical generation methods used to produce the training datasets and the temporal resolution of those datasets on machine learning surrogate models. Using the nonlinear vector autoregression (NVAR) machine as an explainable ML technique, simple dynamical systems are emulated with ML models trained on data produced by three classical numerical integration schemes. NVAR is validated as a skillful ML method, capable of producing accurate predictions and, more importantly, reconstructing both the underlying dynamics and the numerical integration scheme used to generate the training data. However, the machine fails to generalize predictions on unseen test data generated by different numerical integration schemes, despite the underlying dynamical system being the same. This result provides a word of caution for the growing field of machine learning emulation of weather and climate dynamics. Furthermore, we illustrate using NVAR that training on temporally coarsened data may increase the required complexity of ML models and potentially introduce new numerical challenges. Finally, we discover that empirical integration schemes with arbitrary time-stepping sizes can be constructed directly from the data, which implies a potential for the development of empirical numerical integration schemes.

54 ENVIRONMENTAL SCIENCES↗