Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data engineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data for A Generalized Platform for Artificial Intelligence-powered Autonomous Protein Engineering

Proteins are the molecular machines of life with numerous applications in energy, health, and sustainability. However, engineering proteins with desired functions for practical applications remains slow, expensive, and specialist-dependent. Here we report a generally applicable platform for autonomous enzyme engineering that integrates machine learning and large language models with biofoundry automation to eliminate the need for human intervention, judgement, and domain expertise. Requiring only an input protein sequence and a quantifiable way to measure fitness, this automated platform can be applied to engineer a wide array of proteins. As a proof of concept, we engineer Arabidopsis thaliana halide methyltransferase (AtHMT) for a 90-foldimprovement in substrate preference and 16-fold improvement in ethyl-transferase activity, along with developing a Yersinia mollaretii phytase (YmPhytase) variant with 26-fold improvement in activity at neutral pH. This is accomplished in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme. This platform for autonomous experimentation paves the way for rapid advancements across diverse industries, from medicine and biotechnology to renewable energy and sustainable chemistry.

AI/ML↗

Comment on “Thermodynamic Models for the (HClO 4 + NaClO 4 ){aq} and (HBr + NaBr){aq} Systems at 298.15 K and 0.1 MPa” Authored by Oakes, C. S., Ward, A. L., Chugunov, N. Journal of Chemical & Engineering Data , 68 , 2554–2562

Oakes et al. (2023) published a review article in this journal. In that paper, Oakes et al. (2023) developed thermodynamic models to describe electrolyte solutions for HClO 4 –NaClO 4 –H 2 O and HBr–NaBr–H 2 O systems, based on literature data. In their paper, previously published work from researchers in the field was criticized; some of it is ours. Here, in this brief Comment, we first comment on their models, and then we briefly provide a technical response to that criticism.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improved biomass feedstock materials handling and feeding engineering data sets, design methods, and modeling/simulation tools

Forest Concepts led a project to develop a set of tools which enable feedstock handing equipment designers to better understand and model the flowability of bulk particulate biomass materials. The work products included improved flowability mathematical models, data sets used to populate the models, and two new laboratory devices – a biomass-scale true cubical triaxial tester and a biomass-scale gas pycnometer. This work was funded in part by the US Department of Energy under contact DE-EE0008254.

09 BIOMASS FUELS↗

Data for Engineering and Evolution of Yarrowia lipolytica for Producing Lipids from Lignocellulosic Hydrolysates

Yarrowia lipolytica , an oleaginous yeast, shows promise for industrial fermentation due to its robust acetyl-CoA flux and well-developed genetic engineering tools. However, its lack of an active xylose metabolism restricts the conversion of cellulosic sugars to valuable products. To address this, metabolic engineering, and adaptive laboratory evolution (ALE) were applied to the Y. lipolytica PO1f strain, resulting in an efficient xylose-assimilating strain (XEV). Whole-genome sequencing (WGS) of the XEV followed by reverse engineering revealed that the amplification of the heterologous oxidoreductase pathway and a mutation in the GTPase-activating protein gene (YALI0B12100g) might be the primary reasons for improved xylose assimilation in the XEV strain. When a sorghum hydrolysate was used, the XEV strain showed superior xylose consumption and lipid production compared to its parental strain (X123). This study advances our understanding of xylose metabolism in Y. lipolytica and proposes effective metabolic engineering strategies for optimizing lignocellulosic hydrolysates.

Hydrolysate↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

AI-Driven Crack Detection for Remanufacturing Cylinder Heads Using Deep Learning and Engineering-Informed Data Augmentation

Detecting cracks in cylinder heads traditionally relies on manual inspection, which is time-consuming and susceptible to human error. As an alternative, automated object detection utilizing computer vision and machine learning models has been explored. However, these methods often face challenges due to a lack of sufficiently annotated training data, limited image diversity, and the inherently small size of cracks. Addressing these constraints, this paper introduces a novel automated crack-detection method that enhances data availability through a synthetic data generation technique. Unlike general data augmentation practices, our method involves copying cracks from one location to another, guided by both random and informed engineering decisions about likely crack formations due to cyclic thermomechanical loads. The innovative aspect of our approach lies in the integration of domain-specific engineering knowledge into the synthetic generation process, which substantially improves detection accuracy. We evaluate our method’s effectiveness using two metrics: the F2 score, which emphasizes recall to prioritize detecting all potential cracks, and mean average precision (MAP), a standard measure in object detection. Experimental results demonstrate that, without engineering insights, our method increases the F2 score from 0.40 to 0.65, while maintaining a stable MAP. Incorporating detailed engineering knowledge further enhances the F2 score to 0.70 and improves MAP to 0.57, representing increases of 63% and 43%, respectively. These results confirm that our approach not only mitigates the limitations of traditional data augmentation but also significantly advances the reliability and precision of crack detection in industrial settings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Evaluation of Spray and Combustion Models for Simulating Dilute Combustion in a Direct-Injection Spark-Ignition Engine

Dilute combustion in spark-ignition engines has the potential to improve thermal efficiency by mitigating knock and by reducing throttling and wall heat losses. However, ignition and combustion processes can become unstable for dilute operation due to a lowered laminar flame speed, resulting in excessive cycle-to-cycle variability (CCV) of the combustion process. To compensate for the slower combustion in less reactive mixtures, a modified intake port geometry can be employed to generate a strong tumble flow in the cylinder and elevate turbulence levels around the spark plug, thereby promoting a faster transition to turbulent deflagration. Consequently, optimizing combustion chamber geometry and operating strategy is crucial to maximizing the benefits of using dilute combustion with enhanced in-cylinder turbulence across a wide range of operating conditions. Computational fluid dynamics (CFD) simulations can be utilized for virtual engine optimization tasks, but this would require the models to be truly predictive regarding the impact of changes to the engine design and operational parameters.In this study, multicycle large-eddy simulations (LES) are performed for a direct-injection spark-ignition engine to investigate the model performance in predicting engine combustion characteristics with respect to changes in the intake configuration. A tumble plate that blocks the lower part of the intake port inlet is used to vary the tumble. A set of CFD models that have been recently developed are employed, which takes into account the drag of nonspherical droplets, flash-boiling behavior of liquid sprays, spray-wall interaction, surrogate formulation of a research-grade E10 gasoline, and fast chemical kinetic solvers. Simulation results are compared to experimental engine data in terms of cylinder pressure, apparent heat release rate, mass fraction burned timing, and flame images. It is found that LES employing the state-of-the-art CFD models are capable of properly predicting the spray processes and reproducing the measured mean cylinder pressure for the case with the tumble plate. On the other hand, the LES over-predicts the combustion rate during the early combustion stage and under-estimates the CCV, and these discrepancies become larger when the tumble plate is removed.

computational fluid dynamics simulation↗

Data‐Driven Engineering of Thermostable Collagen‐Mimetic Peptoid Triple Helices

Collagen-mimetic peptides (CMPs) are engineered molecules designed to replicate the triple-helical structure of natural collagen. A repeating x–y-Gly sequence is the defining motif of CMPs and is critical to their triple-helical structure and stability. Substitutions to the residues occupying the x and y positions present a means to modulate the CMP structure and properties. Peptoid residues—N-substituted glycine derivatives—present an attractive potential substitution due to their thermal stability, proteolytic resistance, biocompatibility, and diverse palette of non-natural side chains, but also tend to introduce a high degree of backbone flexibility that can diminish the stability of the triple helix. In this work, we report a computational active learning cycle comprising molecular dynamics simulation, Gaussian process regression, and Bayesian optimization to computationally identify a number of promising peptoid substitutions predicted to stabilize the desired quaternary structure through side chain interactions and produce stable peptoid-based collagen-like triple helices. To experimentally test the computational predictions, a top candidate identified by the screen was synthesized and imaged using scanning electron microscopy to resolve fibril-like bundles consistent with collagen-like triple helices. This work predicts a number of CMP peptoid substitutions capable of forming stable triple-helical structures, presents a generalizable design strategy for engineering desired peptoid structures, and opens new avenues for the design of peptoid-based biomimetic materials.

active learning↗

Data for Metabolic Engineering of Nonmodel Yeast Issatchenkia orientalis SD108 for 5-Aminolevulinic Acid Production

Biological production of 5‐aminolevulinic acid (5‐ALA) has received growing attentionover theyears.However, thereis the tradeoff between 5‐ALA biosynthesis and cell growth because the fermentation broth will become acidic due to the production of 5‐ALA. To address this limitation, we engineered an acid‐tolerant yeast, Issatchenkia orientalis SD108, for 5‐ALA production. We first discovered that the cell growth rate of I. orientalis SD108 was boosted by 5‐ALA and its endogenous ALA synthetase (ALAS) showed higher activity than those homologs from other yeasts. The titer of 5‐ALA was improved from 28mg/L to 120‐, 150‐, and 300mg/L, by optimizing plasmid design, overexpressing a transporter, and increasing gene copy number, respectively. After redirecting the metabolic flux using the pyruvate decarboxylase (PDC) knockout strain (SD108ΔPDC) and culturing with urea, we increased the titer of 5‐ALA to 510mg/L, a 13‐fold enhancement, proving the importance of the newly identified IoALAS with higher activity and the strategic selection of nitrogen sources for knockout strains. This study demonstrates the acid‐tolerant I. orientalis SD108ΔPDC has a high potential for 5‐ALA production at a large scale in the future.

Bioproducts↗

End-to-end deep learning pipeline for real-time Bragg peak segmentation: from training to large-scale deployment

X-ray crystallography reconstruction, which transforms discrete X-ray diffraction patterns into three-dimensional molecular structures, relies critically on accurate Bragg peak finding for structure determination. As X-ray free electron laser (XFEL) facilities advance toward MHz data rates (1 million images per second), traditional peak finding algorithms that require manual parameter tuning or exhaustive grid searches across multiple experiments become increasingly impractical. While deep learning approaches offer promising solutions, their deployment in high-throughput environments presents significant challenges in automated dataset labeling, model scalability, edge deployment efficiency, and distributed inference capabilities. We present an end-to-end deep learning pipeline with three key components: (1) a data engine that combines traditional algorithms with our peak matching algorithm to generate high-quality training data at scale, (2) a modular architecture that scales from a few million to hundreds of million parameters, enabling us to train large expert-level models offline while deploying smaller, distilled models at the edge, and (3) a decoupled producer-consumer architecture that separates specialized data source layer from model inference, enabling flexible deployment across diverse computing environments. Using this integrated approach, our pipeline achieves accuracy comparable to traditional methods tuned by human experts while eliminating the need for experiment-specific parameter tuning. Although current throughput requires optimization for MHz facilities, our system's scalable architecture and demonstrated model compression capabilities provide a foundation for future high-throughput XFEL deployments.

Wang, Cong↗

ARM Cloud and Precipitation Measurements and Science Group (CPMSG) 2024 Workshop Report

The mission of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility is to improve the understanding and representation of cloud and aerosol processes and their interaction with the Earth's surface in Earth system models (ESMs) by providing comprehensive field observations and supporting advanced data analytics. The ARM Cloud and Precipitation Measurements and Science Group (CPMSG) was chartered in March 2019 to help improve the performance and scientific impact of ARM measurements of clouds and precipitation. The group aims to identify and address gaps in measurement capabilities, maximize the scientific impact of ARM data, and effectively serve the scientific community. To achieve these goals, the group includes experts in cloud and precipitation science, as well as representatives from ARM infrastructure, including instrument mentors, engineers, data quality officers, and data product translators. Prior to CPMSG, early discussions on cloud and precipitation measurements primarily focused on improving radar systems, but have since evolved to include a broader scope involving radiometers and other instruments. Since its formation, the CPMSG has gathered feedback using science traceability matrices. CPMSG aims to keep these as living documents to show the measurement needs, scientific drivers, roadblocks, maturity of measurements and retrievals, and pathways to model improvements. The group meets quarterly to discuss and prioritize measurement and operational improvements.

54 ENVIRONMENTAL SCIENCES↗

ARM Cloud and Precipitation Measurements and Science Group (CPMSG) 2024 Workshop Report

The mission of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility is to improve the understanding and representation of cloud and aerosol processes and their interaction with the Earth's surface in Earth system models (ESMs) by providing comprehensive field observations and supporting advanced data analytics. The ARM Cloud and Precipitation Measurements and Science Group (CPMSG) was chartered in March 2019 to help improve the performance and scientific impact of ARM measurements of clouds and precipitation. The group aims to identify and address gaps in measurement capabilities, maximize the scientific impact of ARM data, and effectively serve the scientific community. To achieve these goals, the group includes experts in cloud and precipitation science, as well as representatives from ARM infrastructure, including instrument mentors, engineers, data quality officers, and data product translators. Prior to CPMSG, early discussions on cloud and precipitation measurements primarily focused on improving radar systems, but have since evolved to include a broader scope involving radiometers and other instruments. Since its formation, the CPMSG has gathered feedback using science traceability matrices. CPMSG aims to keep these as living documents to show the measurement needs, scientific drivers, roadblocks, maturity of measurements and retrievals, and pathways to model improvements. The group meets quarterly to discuss and prioritize measurement and operational improvements.

54 ENVIRONMENTAL SCIENCES↗

ARM Cloud and Precipitation Measurements and Science Group (CPMSG) 2024 Workshop Report

The mission of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility is to improve the understanding and representation of cloud and aerosol processes and their interaction with the Earth's surface in Earth system models (ESMs) by providing comprehensive field observations and supporting advanced data analytics. The ARM Cloud and Precipitation Measurements and Science Group (CPMSG) was chartered in March 2019 to help improve the performance and scientific impact of ARM measurements of clouds and precipitation. The group aims to identify and address gaps in measurement capabilities, maximize the scientific impact of ARM data, and effectively serve the scientific community. To achieve these goals, the group includes experts in cloud and precipitation science, as well as representatives from ARM infrastructure, including instrument mentors, engineers, data quality officers, and data product translators. Prior to CPMSG, early discussions on cloud and precipitation measurements primarily focused on improving radar systems, but have since evolved to include a broader scope involving radiometers and other instruments. Since its formation, the CPMSG has gathered feedback using science traceability matrices. CPMSG aims to keep these as living documents to show the measurement needs, scientific drivers, roadblocks, maturity of measurements and retrievals, and pathways to model improvements. The group meets quarterly to discuss and prioritize measurement and operational improvements.

54 ENVIRONMENTAL SCIENCES↗

Mechanisms of Waterflood Inefficiency: Analysis of Geological, Petrophysical and Reservoir History, a Field Case Study of FWU (East Section)

The petroleum reservoir represents a complex heterogeneous system that requires thorough characterization prior to the implementation of any incremental recovery technique. One of the most commonly utilized and successful secondary recovery techniques is waterflooding. However, a lack of sufficient investigation into the inherent behavior and characteristics of the reservoir formation in situ can result in failure or suboptimal performance of waterflood operations. Therefore, a comprehensive understanding of the geological history, static and dynamic reservoir characteristics, and petrophysical data is essential for analyzing the mechanisms and causes of waterflood inefficiency and failure. In this study, waterflood inefficiency was observed in the Morrow B reservoir located in the Farnsworth Unit, situated in the northwestern shelf of the Anadarko Basin, Texas. To assess the potential mechanisms behind the inefficiency of waterflooding in the east half, geological, petrophysical, and reservoir engineering data, along with historical information, were integrated, reviewed, and analyzed. The integration and analysis of these datasets revealed that several factors contributed to the waterflood inefficiency. Firstly, the presence of abundant dispersed authigenic clays within the reservoir, worsened by low reservoir quality and high heterogeneity, led to unfavorable conditions for waterflood operations. The use of freshwater for flooding exacerbated the adverse effects of sensitive and migratory clays, further hampering the effectiveness of the waterflood. In addition to these factors, several reservoir engineering issues played a significant role in the inefficiency of waterflooding. These issues included inadequate perforation strategies due to the absence of detailed hydraulic flow units (HFUs) and rock typing, random placement of injectors, and uncontrolled injected fresh water. These external controlling parameters further contributed to the overall inefficiencies observed during waterflood operations in the east half of the reservoir. A detailed understanding of the mechanistic factors of inefficient waterflood operation will provide adequate insights into the development of the improved recovery technique for the field.

Morgan, Anthony (ORCID:0000000211519153)↗

CFD modeling of near-wall combustion and unburned methane prediction in natural gas spark ignition engines

Natural gas-powered engines play a critical role in gas drilling, compression, and transmission sectors, but methane (CH 4 ) from engine combustion slip can be significant over their lifespan, contributing to atmospheric pollution and signaling reduced engine efficiency. Here, to address this challenge, computational fluid dynamics (CFD) simulations offer valuable insights into the in-cylinder combustion process, enabling the optimization of combustion strategies and engine designs to minimize unburned CH 4 slip. This study aims to evaluate and improve combustion models for simulating the combustion process and predicting unburned CH 4 concentrations in natural gas spark-ignition (SI) engines, including engines that are part of combined reformer-engine systems. Specifically, the performance of two flamelet-based combustion models—the Extended Coherent Flame Model (ECFM) and the G-equation model—was assessed using experimental engine data collected under varying excess-air ratio (λ) conditions and fuel compositions, including natural gas and syngas blends. In addition, to enhance the predictive capabilities of the G-equation model, a flame-wall interaction (FWI) sub-model was integrated into its framework. The effects of its model parameters, such as quenching and influence distance, on combustion behavior and unburned methane predictions were analyzed in detail. The ECFM tended to predict delayed combustion phasing under diluted mixture conditions, resulting in overprediction of unburned CH 4 concentrations. In contrast, the G-equation model provided reasonable predictions of combustion pressure, while representing higher the CH 4 reduction rate across the operating condition compared to experimental data. Incorporating the FWI sub-model—with the quenching distance calculated based on a pressure-dependent relation (P -0.48 ) and a fixed influence distance of 1.5 mm—further improved the G-equation model’s accuracy in predicting CH 4 reduction rates without compromising its ability to simulate the combustion process.

Combustion model↗

A chemical kinetic analysis of knock propensity of methanol-to-gasoline fuel

Production of low carbon gasoline-like fuels such as methanol-to-gasoline (MTG) is a promising approach to achieve rapid greenhouse gas emission reduction of the transportation sector. Despite the fact that gasoline that meets the ASTM D4814 standard for automotive spark-ignition engine fuel can be readily produced from these processes, it is unclear how the composition of MTG may affect engine performance and emissions. Here, in this paper, a surrogate for an MTG is used to numerically study the effects of gasoline composition on knock propensity and on the sensitivity of knock to thermal and fuel stratification, to oxygen dilution and to nitric oxide from exhaust gas recirculation of residual gases. Simulations were performed in ANSYS CHEMKIN-PRO using a comprehensive chemical kinetic mechanism for gasoline surrogates, and results of the MTG surrogate were compared against those of a petroleum-based regular E10 gasoline, termed PACE-20. A premium-grade MTG fuel was also formulated by adding ethanol to the MTG surrogate, and results were compared against those of four premium-grade, gasoline-like fuels representative of future alternative gasoline formulations. Surrogates and mechanism were evaluated by comparison against experimental engine data, and the model showed high accuracy at stoichiometric conditions (mean absolute error of ignition timing equal to 1.46 crank angle degrees) but larger deviations at lean conditions (mean absolute error of ignition timing equal to 5.52 crank angle degrees). Despite the fact that the MTG surrogate has a RON 1.1 units higher than that of PACE-20, it may show higher knock propensity at medium temperature conditions due to a less intense NTC behavior. MTG autoignition was more temperature- and equivalence ratio-sensitive than that of PACE20, suggesting that MTG can benefit more from naturally-occurring thermal stratification or from induced fuel stratification of the end gas to mitigate knock intensity. The sensitivity of autoignition reactivity to oxygen dilution and to NO concentration was higher for MTG than for regular gasoline at medium loads, but the opposite trend was observed at high loads due to the effect of pressure on the low-temperature chemistry of regular gasoline. Approximately 14 % vol ethanol content was required to upgrade the octane rating of MTG from regular grade to premium grade. Adding 13.6 % vol ethanol made the fuel autoignition less sensitive to both oxygen dilution and NO content (ignition time varies approx. 17 % and 50 % less with oxygen dilution and NO addition, respectively, when adding ethanol at high engine loads).

02 PETROLEUM↗