Search NASA⌕ Search

SEARCH · Search NASA

Results for “data engineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A data engineering approach for sustainable chemical end-of-life management

The presence of chemicals causing significant adverse human health and environmental effects during end-of-life (EoL) stages is a challenge for implementing sustainable management efforts and transitioning towards a safer circular life cycle. Conducting chemical risk evaluation and exposure assessment of potential EoL scenarios can help understand the chemical EoL management chain for its safer utilization in a circular life-cycle environment. However, the first step is to track the chemical flows, estimate releases, and potential exposure pathways. Hence, this work proposes an EoL data engineering approach to perform chemical flow analysis and screening to support risk evaluation and exposure assessment for designing a safer circular life cycle of chemicals. This work uses publicly-available data to identify potential post-recycling scenarios (e.g., industrial processing/use operations), estimate inter-industry chemical transfers, and exposure pathways to chemicals of interest. Furthermore, a case study demonstration shows how the data engineering framework identifies, estimates, and tracks chemical flow transfers from EoL stage facilities (e.g., recycling and recovery) to upstream chemical life cycle stage facilities (e.g., manufacturing). Also, the proposed framework considers current regulatory constraints on closing the recycling loop operations and provides a range of values for the flow allocated to post-recycling uses associated with occupational exposure and fugitive air releases from EoL operations.

54 ENVIRONMENTAL SCIENCES↗

Data for A Generalized Platform for Artificial Intelligence-powered Autonomous Protein Engineering

Proteins are the molecular machines of life with numerous applications in energy, health, and sustainability. However, engineering proteins with desired functions for practical applications remains slow, expensive, and specialist-dependent. Here we report a generally applicable platform for autonomous enzyme engineering that integrates machine learning and large language models with biofoundry automation to eliminate the need for human intervention, judgement, and domain expertise. Requiring only an input protein sequence and a quantifiable way to measure fitness, this automated platform can be applied to engineer a wide array of proteins. As a proof of concept, we engineer Arabidopsis thaliana halide methyltransferase (AtHMT) for a 90-foldimprovement in substrate preference and 16-fold improvement in ethyl-transferase activity, along with developing a Yersinia mollaretii phytase (YmPhytase) variant with 26-fold improvement in activity at neutral pH. This is accomplished in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme. This platform for autonomous experimentation paves the way for rapid advancements across diverse industries, from medicine and biotechnology to renewable energy and sustainable chemistry.

AI/ML↗

Digital Engineering & Data Acquisition

This is an Intern Poster submission regarding my research in the digital engineering team. My task was to create a software adapter to communicate between a data acquisition system (LabVIEW) and a data warehouse (Deep Lynx) as described in the poster.

99 GENERAL AND MISCELLANEOUS↗

Comment on “Thermodynamic Models for the (HClO 4 + NaClO 4 ){aq} and (HBr + NaBr){aq} Systems at 298.15 K and 0.1 MPa” Authored by Oakes, C. S., Ward, A. L., Chugunov, N. Journal of Chemical & Engineering Data , 68 , 2554–2562

Oakes et al. (2023) published a review article in this journal. In that paper, Oakes et al. (2023) developed thermodynamic models to describe electrolyte solutions for HClO 4 –NaClO 4 –H 2 O and HBr–NaBr–H 2 O systems, based on literature data. In their paper, previously published work from researchers in the field was criticized; some of it is ours. Here, in this brief Comment, we first comment on their models, and then we briefly provide a technical response to that criticism.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improved biomass feedstock materials handling and feeding engineering data sets, design methods, and modeling/simulation tools

Forest Concepts led a project to develop a set of tools which enable feedstock handing equipment designers to better understand and model the flowability of bulk particulate biomass materials. The work products included improved flowability mathematical models, data sets used to populate the models, and two new laboratory devices – a biomass-scale true cubical triaxial tester and a biomass-scale gas pycnometer. This work was funded in part by the US Department of Energy under contact DE-EE0008254.

09 BIOMASS FUELS↗

Data for Engineering and Evolution of Yarrowia lipolytica for Producing Lipids from Lignocellulosic Hydrolysates

Yarrowia lipolytica , an oleaginous yeast, shows promise for industrial fermentation due to its robust acetyl-CoA flux and well-developed genetic engineering tools. However, its lack of an active xylose metabolism restricts the conversion of cellulosic sugars to valuable products. To address this, metabolic engineering, and adaptive laboratory evolution (ALE) were applied to the Y. lipolytica PO1f strain, resulting in an efficient xylose-assimilating strain (XEV). Whole-genome sequencing (WGS) of the XEV followed by reverse engineering revealed that the amplification of the heterologous oxidoreductase pathway and a mutation in the GTPase-activating protein gene (YALI0B12100g) might be the primary reasons for improved xylose assimilation in the XEV strain. When a sorghum hydrolysate was used, the XEV strain showed superior xylose consumption and lipid production compared to its parental strain (X123). This study advances our understanding of xylose metabolism in Y. lipolytica and proposes effective metabolic engineering strategies for optimizing lignocellulosic hydrolysates.

Hydrolysate↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

HFTS-1 Natural Joint Data and Engineering Summary

This report provides an analysis of engineering and geologic data collected from the Hydraulic Fracturing Test Site #1 (HTSF-1) project in the southern Midland Basin, Reagan County, Texas. The site is being studied as part of the Science-informed Machine Learning for Accelerating Real-Time Decisions in Subsurface Applications (SMART) Initiative at the National Energy Technology Laboratory (NETL). The data collected is intended to provide a basis of understanding of the site and to construct simulation models. This report provides a summary analysis of the engineering and geologic data collected from the project to provide a basis of understanding of the site and to construct simulation models. The report examines various possible correlations in engineering properties based on natural fracture data from the program, which consisted of 11 horizontal wells, one vertical well, and one slant well. The horizontal wells are in two horizons: 1) Upper and 2) Middle Wolfcamp formations. Program data include: fracture frequency and fracture orientation data from four core runs in a slant well; fracture spacing and orientation from a vertical pilot well; laboratory triaxial testing and mineralogical determinations; and porosity results from magnetic resonance analyses from various wells, together with observations based on the data collection. In addition, available references were reviewed on the site for additional insights. As the focus of the report is on the natural system, hydraulic fracture data from the site were not examined in detail in this report. Data variability is the chief observation in examination of the database. Fracture frequency in the slant well can range from sections with values as high as five fractures per ft to sections up to 100+ ft in length with no natural observed fractures. Fracture spacing across is typically less than 10 ft, but can range up to hundreds of feet. Apparent fracturing shows the trends in two predominate orientations, E-SW and WNW-ESE, but minor variations exist. The rock units vary across the site from siliceous mudstones to calcareous mudstones, showing a general layering with depth. The laboratory properties such as strength and modulus show no apparent trend with depth, but appear to correlate with rock mineralogy with high strength and modulus values where calcium content is high. In addition, an attempt to examine variability and mineralogy on a larger scale was made using a color-coded system based on gamma ray measurements. A staged colored approach was adopted, presuming that lower gamma ray values indicate higher value of calcium content (blue scale) and that higher gamma ray values indicate higher clay mineral content (orange scale). As provided in report appendices, the system correlated well with visual examination of the slant core and the petrofabric analyses of the vertical pilot well. The results showed large variability in mineral content along the horizontal plane across the site. The change in mineral content was also rapid, on a scale less than that of the average hydraulic fracture stage length of about 180 ft.

42 ENGINEERING↗

AI-Driven Crack Detection for Remanufacturing Cylinder Heads Using Deep Learning and Engineering-Informed Data Augmentation

Detecting cracks in cylinder heads traditionally relies on manual inspection, which is time-consuming and susceptible to human error. As an alternative, automated object detection utilizing computer vision and machine learning models has been explored. However, these methods often face challenges due to a lack of sufficiently annotated training data, limited image diversity, and the inherently small size of cracks. Addressing these constraints, this paper introduces a novel automated crack-detection method that enhances data availability through a synthetic data generation technique. Unlike general data augmentation practices, our method involves copying cracks from one location to another, guided by both random and informed engineering decisions about likely crack formations due to cyclic thermomechanical loads. The innovative aspect of our approach lies in the integration of domain-specific engineering knowledge into the synthetic generation process, which substantially improves detection accuracy. We evaluate our method’s effectiveness using two metrics: the F2 score, which emphasizes recall to prioritize detecting all potential cracks, and mean average precision (MAP), a standard measure in object detection. Experimental results demonstrate that, without engineering insights, our method increases the F2 score from 0.40 to 0.65, while maintaining a stable MAP. Incorporating detailed engineering knowledge further enhances the F2 score to 0.70 and improves MAP to 0.57, representing increases of 63% and 43%, respectively. These results confirm that our approach not only mitigates the limitations of traditional data augmentation but also significantly advances the reliability and precision of crack detection in industrial settings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A fusion relevant data-driven engineering void swelling model for 9Cr tempered martensitic steels

The UCSB database on cavity evolution in 9-12Cr tempered martensitic steels (TMS), includes the results for both dual heavy and helium ion (DII), and High Flux Isotope Reactor (HFIR) in situ helium injection (ISHI) neutron irradiations at 500°C. These results were combined with literature single ion and fission neutron irradiation data to derive a model for the void volume fraction, f v , as a function of displacements per atom (dpa) and transmutant helium concentrations in atomic parts per million (appm). The scientific foundation for the paper is described in a companion paper entitled “Cavity Evolution and Void Swelling in Dual Ion Irradiated Tempered Martensitic Steels”. Here, in this study, we show that f v (dpa, He/dpa) is described by the incubation dose, dpa i , for the onset of void growth, and the post-incubation growth rate, f v ’(%/dpa). Both dpa i and f v ’ decrease with increasing He/dpa at > ~ 5. The dpa i is also lower for the ISHI neutron irradiations at the same He/dpa. Single heavy ion and fission reactor neutron irradiations, with low He/dpa ratios, have a much larger dpa i . Based on a combined analysis of DII, single ion, ISHI and fission neutron data, we further show that the post-incubation f v data analyzed here have a common empirical curve shape, with f v ’ reaching up to ~ 0.2%/dpa at very high dpa. We also show that f v ’ can be predicted based on a physical model of defect partitioning between evolving sinks. At 500°C and fusion relevant He/dpa ≈ 10, the best-fit model predicts nominal swelling, S = f v /(1-f v ), of ~ 1.1, 4.9 and 16% at 50, 100 and 200 dpa, respectively. The physically motivated, data-driven model includes estimated uncertainties for both dpa i and f v ’.

36 MATERIALS SCIENCE↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Evaluation of Spray and Combustion Models for Simulating Dilute Combustion in a Direct-Injection Spark-Ignition Engine

Dilute combustion in spark-ignition engines has the potential to improve thermal efficiency by mitigating knock and by reducing throttling and wall heat losses. However, ignition and combustion processes can become unstable for dilute operation due to a lowered laminar flame speed, resulting in excessive cycle-to-cycle variability (CCV) of the combustion process. To compensate for the slower combustion in less reactive mixtures, a modified intake port geometry can be employed to generate a strong tumble flow in the cylinder and elevate turbulence levels around the spark plug, thereby promoting a faster transition to turbulent deflagration. Consequently, optimizing combustion chamber geometry and operating strategy is crucial to maximizing the benefits of using dilute combustion with enhanced in-cylinder turbulence across a wide range of operating conditions. Computational fluid dynamics (CFD) simulations can be utilized for virtual engine optimization tasks, but this would require the models to be truly predictive regarding the impact of changes to the engine design and operational parameters.In this study, multicycle large-eddy simulations (LES) are performed for a direct-injection spark-ignition engine to investigate the model performance in predicting engine combustion characteristics with respect to changes in the intake configuration. A tumble plate that blocks the lower part of the intake port inlet is used to vary the tumble. A set of CFD models that have been recently developed are employed, which takes into account the drag of nonspherical droplets, flash-boiling behavior of liquid sprays, spray-wall interaction, surrogate formulation of a research-grade E10 gasoline, and fast chemical kinetic solvers. Simulation results are compared to experimental engine data in terms of cylinder pressure, apparent heat release rate, mass fraction burned timing, and flame images. It is found that LES employing the state-of-the-art CFD models are capable of properly predicting the spray processes and reproducing the measured mean cylinder pressure for the case with the tumble plate. On the other hand, the LES over-predicts the combustion rate during the early combustion stage and under-estimates the CCV, and these discrepancies become larger when the tumble plate is removed.

computational fluid dynamics simulation↗

Data‐Driven Engineering of Thermostable Collagen‐Mimetic Peptoid Triple Helices

Collagen-mimetic peptides (CMPs) are engineered molecules designed to replicate the triple-helical structure of natural collagen. A repeating x–y-Gly sequence is the defining motif of CMPs and is critical to their triple-helical structure and stability. Substitutions to the residues occupying the x and y positions present a means to modulate the CMP structure and properties. Peptoid residues—N-substituted glycine derivatives—present an attractive potential substitution due to their thermal stability, proteolytic resistance, biocompatibility, and diverse palette of non-natural side chains, but also tend to introduce a high degree of backbone flexibility that can diminish the stability of the triple helix. In this work, we report a computational active learning cycle comprising molecular dynamics simulation, Gaussian process regression, and Bayesian optimization to computationally identify a number of promising peptoid substitutions predicted to stabilize the desired quaternary structure through side chain interactions and produce stable peptoid-based collagen-like triple helices. To experimentally test the computational predictions, a top candidate identified by the screen was synthesized and imaged using scanning electron microscopy to resolve fibril-like bundles consistent with collagen-like triple helices. This work predicts a number of CMP peptoid substitutions capable of forming stable triple-helical structures, presents a generalizable design strategy for engineering desired peptoid structures, and opens new avenues for the design of peptoid-based biomimetic materials.

active learning↗

Data for Metabolic Engineering of Nonmodel Yeast Issatchenkia orientalis SD108 for 5-Aminolevulinic Acid Production

Biological production of 5‐aminolevulinic acid (5‐ALA) has received growing attentionover theyears.However, thereis the tradeoff between 5‐ALA biosynthesis and cell growth because the fermentation broth will become acidic due to the production of 5‐ALA. To address this limitation, we engineered an acid‐tolerant yeast, Issatchenkia orientalis SD108, for 5‐ALA production. We first discovered that the cell growth rate of I. orientalis SD108 was boosted by 5‐ALA and its endogenous ALA synthetase (ALAS) showed higher activity than those homologs from other yeasts. The titer of 5‐ALA was improved from 28mg/L to 120‐, 150‐, and 300mg/L, by optimizing plasmid design, overexpressing a transporter, and increasing gene copy number, respectively. After redirecting the metabolic flux using the pyruvate decarboxylase (PDC) knockout strain (SD108ΔPDC) and culturing with urea, we increased the titer of 5‐ALA to 510mg/L, a 13‐fold enhancement, proving the importance of the newly identified IoALAS with higher activity and the strategic selection of nitrogen sources for knockout strains. This study demonstrates the acid‐tolerant I. orientalis SD108ΔPDC has a high potential for 5‐ALA production at a large scale in the future.

Bioproducts↗

Data for "Metabolic Engineering Strategies to Produce Medium-Chain Oleochemicals via Acyl-ACP:CoA Transacylase Activity"

Microbial lipid metabolism is an attractive route for producing oleochemicals. The predominant strategy centers on heterologous thioesterases to synthesize desired chain-length fatty acids. To convert acids to oleochemicals (e.g., fatty alcohols, ketones), the narrowed fatty acid pool needs to be reactivated as coenzyme A thioesters at cost of one ATP per reactivation – an expense that could be saved if the acyl-chain was directly transferred from ACP- to CoA-thioester. Here, we demonstrate such an alternative acyl-transferase strategy by heterologous expression of PhaG, an enzyme first identified in Pseudomonads, that transfers 3-hydroxy acyl-chains between acyl-carrier protein and coenzyme A thioester forms for creating polyhydroxyalkanoate monomers. We use it to create a pool of acyl-CoA’s that can be redirected to oleochemical products. Through bioprospecting, mutagenesis, and metabolic engineering, we develop three strains of Escherichia coli capable of producing over 1 g/L of medium-chain free fatty acids, fatty alcohols, and methyl ketones.

Bioproducts↗