Search NASA⌕ Search

SEARCH · Search NASA

Results for “data processing methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa↗

Assimilating partial observation to enhance feedback control of stochastic dynamical systems

Here, in this paper, we present a novel methodology to tackle feedback optimal control problems in scenarios where the exact state of the controlled process is unknown. It integrates data assimilation techniques and optimal control solvers to manage partial observation of the state process, a common occurrence in practical scenarios. Traditional stochastic optimal control methods assume full state observation, which is often not feasible in real-world fluid dynamics control problems. Our approach underscores the significance of utilizing observational data to inform control policy design. Specifically, we introduce a kernel learning backward stochastic differential equation (SDE) filter to enhance data assimilation efficiency and propose a sample-wise stochastic optimization method within the stochastic maximum principle framework. We demonstrate the efficacy and accuracy of our method in the control of advection-diffusion-reaction flow problem and the Dubins airplane maneuvering problem with model uncertainty.

data driven↗

Correlating processing variables to material properties in recycled polypropylene: A data‐driven approach

Abstract Polypropylene (PP) is one of the most widely used plastics, yet its recycling remains limited, with less than 1% of solid waste PP being reprocessed. Mechanical recycling through extrusion is the most practical method, but inconsistent reprocessing conditions introduce variability in material properties. While temperature, screw speed, and residence time influence the thermomechanical stress applied during reprocessing, there are no standardized guidelines for optimizing these parameters. This study examines how these factors shape the properties of recycled PP, using conditions designed to mimic post‐industrial recycled (PIR) scrap. Residence time was measured using colorimetric tracking and correlated with molecular weight, viscosity, and mechanical properties over multiple extrusion cycles. Data‐driven modeling, including response surface methodology, support vector machines, and artificial neural networks, identified processing temperature as the dominant factor in material degradation, followed by residence time. Mechanical properties remained stable, while viscosity decreased predictably with increasing residence time. By linking reprocessing conditions to property evolution, this study provides a method to optimize processing parameters and reduce variability in recycled PP. These findings help manufacturers improve process control, making recycled PP more predictable for reuse in manufacturing. Highlights Study of PIR‐quality PP without additives or compatibilizers. Residence time analysis shows processing temperature drives PP property changes. Mark‐Houwink enables quick molecular weight checks for quality control. Models predict mechanical and rheological shifts in reprocessing. Optimized processing parameters minimize property degradation in recycling.

Estela‐García, John E. [Polymer Engineering Center↗

Accelerating catalytic advancements through the precision of high-throughput experiments & calculations

The growing demand for energy-efficient processes to support a sustainable future drives the need for research to rapidly explore chemical and material space through accelerated catalyst discovery initiatives. Recent breakthroughs in high-throughput experimental and computational methods are transforming the catalysis field, surpassing traditional approaches to manipulating variables in catalytic processes. Key advancements in innovation include the integration of machine learning for efficient catalyst screening, high-throughput experimentation, data-driven methodologies employing comprehensive databases, and in situ and in operando techniques for realistic observations. This progress has undoubtedly been intertwined with a collaborative framework across disciplines, reshaping catalyst discovery methods in both industry and academia. This Opinion article presents a multifaceted perspective from coauthors with expertise spanning various stages of the Technology Readiness Level spectrum, highlighting both opportunities and persistent challenges in integrating computational and experimental approaches in catalysis. These challenges span from obtaining high-quality experimental data, scaling simulations to industrially relevant materials and process conditions to navigating the complexity and predictive accuracy of computational models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Block Island Acoustic Data Analysis and Peak Detection Tools

This dataset contains acoustic signal analysis tools and processed datasets for pile driving noise characterization during Block Island Wind Farm construction (October 25, 2015), including peak detection algorithms, signal extraction methods, and visualization products for multi-channel hydrophone array data.

17 WIND ENERGY↗

Self‐Assembly Methods Induce Different Individual‐Polymer‐Block Solvation Responses

An understanding of block-specific responses to stimuli in self-assembled copolymers is essential for the effective design of responsive materials. Here, materials formed from different post-reaction processing methods of the same ROMP block copolymer exhibit different block-specific solvent responses, as characterized in situ by fluorescence lifetime imaging microscopy (FLIM). FLIM data were acquired by tagging the polar or nonpolar block separately with a covalently incorporated viscosity-sensitive fluorescent molecular rotor to measure the changing tightness or looseness of assembly upon changes in solvent composition. Rapid precipitation from an insoluble solvent to form a film results in tighter assembly and less difference in tightness/looseness between the two blocks. Further, it interrupts individual block-solvent response behavior. Both results indicate an apparent disruption of the core–shell assembly. This model from FLIM is further supported by differential scanning calorimetry (DSC) and small-angle X-ray scattering (SAXS) data from isolated solids, which show tighter long-range interactions and more long-range order in these kinetically trapped states. Here, the experiments develop FLIM as a method for pinpointing stimuli-responsive behaviors from different processing methods to individual blocks. Together these outcomes provide a future handle for tailoring solvent-triggered assembly/disassembly behavior.

López, Pía A. [University of California, Irvine, C↗

Power System Feature-Based Event Classification by Means of Multiple PMU Data

Abstract—Phasor Measurement Units (PMUs) provide time synchronized measurements across the power grid, enabling data driven event detection and classification for enhanced system monitoring and situational awareness. However, variations in event duration, spatial extent, and severity, along with coincident events, pose challenges for conventional classification models that require fixed-size inputs. This paper presents a feature-based framework that aggregates diverse attributes from all available PMUs for each event into a fixed-length vector, facilitating the application of standard machine learning classifiers, including Random Forest, XGBoost, and Multilayer Perceptron. A probabilistic post-processing scheme is further introduced to enable multi-label classification in the presence of overlapping events. Experiments using real-world PMU data demonstrate that the Random Forest model achieves 95% accuracy, while the proposed post-processing method yields an additional 3% improvement.

Nematirad, Reza↗

In‐situ Analysis of Paste Properties in Resonant Acoustic Mixers for Quality Monitoring

Formulation control is key to achieving consistent target properties of energetic materials, as feedstock variations and slight deviations in the ratios of different ingredients can have major effects on final product properties, particularly in dense pastes with high particle loading >65 vol.%. In large‐scale operations, it is imperative to either correct or remove batches of material that perform outside baseline property specifications as early as possible to avoid unnecessary processing of suboptimal material. Quality monitoring is the practice of measuring material properties during processing using process analytical technologies as opposed to only testing the properties of the final product; it is a key principle in the quality‐by‐design frameworks used for designing formulations and manufacturing processes. Herein, a process analytical technology method for correlating material properties of dense pastes directly after mixing in a Resonant Acoustic Mixer to motor data is developed and used to detect differences in the particle content of dense paste formulations. This method was also capable of detecting variations in powder feedstock properties, such as particle packing efficiency, and is sensitive enough to detect changes of 2 wt.% in the total solids content of the formulation. The techniques presented herein show excellent promise for use as a process analytical technology capable of quantifying formulation effects on material movement modes during resonant acoustic mixing.

Materials science↗

Autonomous hybrid optimization of a SiO 2 plasma etching mechanism

Computational modeling of plasma etching processes at the feature scale relevant to the fabrication of nanometer semiconductor devices is critically dependent on the reaction mechanism representing the physical processes occurring between plasma produced reactant fluxes and the surface, reaction probabilities, yields, rate coefficients, and threshold energies that characterize these processes. The increasing complexity of the structures being fabricated, new materials, and novel gas mixtures increase the complexity of the reaction mechanism used in feature scale models and increase the difficulty in developing the fundamental data required for the mechanism. This challenge is further exacerbated by the fact that acquiring these fundamental data through more complex computational models or experiments is often limited by cost, technical complexity, or inadequate models. In this paper, we discuss a method to automate the selection of fundamental data in a reduced reaction mechanism for feature scale plasma etching of SiO 2 using a fluorocarbon gas mixture by matching predictions of etch profiles to experimental data using a gradient descent (GD)/Nelder–Mead (NM) method hybrid optimization scheme. These methods produce a reaction mechanism that replicates the experimental training data as well as experimental data using related but different etch processes.

36 MATERIALS SCIENCE↗

Adsorptive denitrogenation of model aviation fuel using mesoporous silica in a packed bed adsorption system

This study aims to understand the effects of system process parameters such as flow rate, adsorbent particle size, and use of recycled adsorbent on denitrogenation performance of a model fuel using mesoporous silica gel. The goal is to reduce the nitrogen content of the model fuel from 1500 parts per million to single-digit ppm to meet ASTM specifications for drop-in fuels. This work was done with the intent of applying adsorptive denitrogenation to sustainable aviation fuel (SAF) product fractions produced via hydrothermal liquefaction (HTL). The adsorption performance of the silica is evaluated via packed column breakthrough data, with data generated from collecting from the column outlet and quantifying nitrogen content via gas chromatography. Select experiments use a significantly larger (2.5x column diameter and length) column to demonstrate linear scalability of the process. Thermogravimetric analysis data is collected to evaluate the effects of thermal calcination as a sorbent regeneration method. Effects of a more complex feed are also investigated using a known reference fuel with additional added nitrogen containing compounds. The results presented in this work successfully demonstrate up to 99.8 % removal of NCCs from a model fuel fraction at an original NCC concentration of approximately 1500 ppm to single-digit parts per million after treatment. We also examine calcination of sorbent materials to remove the adsorbed species to enable sorbent reuse and minimize waste generation and show that the calcined material can be reused up to 5 cycles with reduced adsorption capacity. Overall, this work indicates that adsorptive denitrogenation using silica gel is a viable solution to enable the integration of HTL-derived aviation fuels into existing fuel infrastructure.

Adsorption techniques↗

CoCoMET v1.0: a unified open-source toolkit for atmospheric object tracking and analysis

Advances in performance and analysis capabilities have accelerated the development of object tracking algorithms for atmospheric research. This has resulted in a growing number of studies using Lagrangian tracking techniques to analyze the evolution of atmospheric phenomena and the underlying processes. However, the increasing complexity and variety of tracking algorithms present a steep learning curve for new users and make it difficult for existing users to compare algorithm performance. We introduce CoCoMET (Community Cloud Model Evaluation Toolkit), an open-source toolkit that addresses these issues. CoCoMET simplifies the process of running multiple tracking algorithms simultaneously and analyzing objects in both model and observational datasets by specifying parameters in a single configuration file. It standardizes input data from different sources into a consistent format and unifies the tracking output across algorithms. CoCoMET enhances the functionality of existing tracking methods by calculating additional properties such as cell growth and dissipation rates, perimeter, surface area, convexity, and irregularity. In addition, CoCoMET includes a novel method for identifying mergers and splits in 2D and 3D tracks and supports the integration of Eulerian/stationary datasets external to the tracking data for process studies. Its potential utility is demonstrated through examples of model intercomparison, model evaluation against observations, and comparisons between tracking algorithms. Designed for open-source environments, CoCoMET will continue to expand with future releases, incorporating more input data types and tracking algorithms.

54 ENVIRONMENTAL SCIENCES↗

2025 Workshop on Envisioning Frontiers in AI and Computing for Biological Research: Position Papers

This workshop aims to identify key research directions for transforming biology using artificial intelligence (AI), machine learning (ML) and computational methods to facilitate the discovery of new behaviors, mechanisms, and designs of biological processes relevant to DOE missions, underpinning a broader U.S. bioeconomy. By developing novel AI/ML technologies to analyze and interpret complex biological data, researchers can organize and simulate biological processes at various scales as well as advance predictive understanding and manipulation of biological systems. This integration of computation, experimentation, and next-generation experimental technologies can lead to discoveries in new biological behaviors and mechanisms relevant to DOE missions. The focus is on how advanced computational and mathematical methods can impact this mission by exploring digital twins, foundation models, automated laboratory experiments, modeling of complex living systems, and data-driven approaches for the biodesign of plants and microbial systems. While data management is important, it is not the primary focus of this workshop, which will assess the current state, trends, and AI/ML challenges at the interface between biology and computational science to identify opportunities for high-impact research at their intersection. The goal is to define research needs and opportunities that align with biological sciences, computational sciences, and applied mathematics research.

59 BASIC BIOLOGICAL SCIENCES↗

Materials Characterization: A Primer for Solid Phase Processing Applications

The Pacific Northwest National Laboratory (PNNL) undertook the Materials Characterization, Prediction, and Control (MCPC) Laboratory Directed Research and Development (LDRD) Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems produced via advanced manufacturing methods, such as solid phase processing, for use in national security and advanced energy applications (Smith 2021). As a two-year LDRD investment requiring focused research, the MCPC project applied only a subset of the wide range of available destructive and nondestructive characterization methods to provide data to the predictive modeling and data analytics tasks. The purpose of this report is to review a wide range of destructive and nondestructive characterization methods that are relevant in solid-phase processing (SPP) applications, but not necessarily applied in the MCPC Project as a guide to the planning of characterization activities in future research. Particular attention is given to measured characteristics that can correlate to other material characteristics, with a particular interest in nondestructive evaluation (NDE) that can be applied to samples obtained in the MCPC Project. Destructive examinations include tensile tests, optical and electron microscopy, micro-hardness, and residual stress tests. NDE tests include surface visual inspection, eddy current examination for cracks, 4-point potential drop, ultrasound, x-ray, and computed tomography.

36 MATERIALS SCIENCE↗

A New Reduced Order Model For The Mechanistic Creep Behavior Of UO 2

This manuscript describes an ongoing NEAMS effort to better determine the performance of advanced nuclear fuels, in particular the creep behavior of doped UO$_2$ for light water reactors. In our previous work, we outlined a method to utilize data generated from lower length scale simulations and implement it into the engineering scale fuel performance analysis. This process has been further refined, and in addition, new data has been used to train the surrogate model which has also been substantially improved since the previous iteration. The new model is compared against the current empirical model used in BISON using both scoping calculations to define the performance over the parameter space and using integral instrumented fuel assessment cases to determine the impact of these models on the overall fuel performance. Suggestions and guidance for future improvements to this method are provided to ensure the model covers relevant parameter space and phenomena.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Development of systematic uncertainty-aware neural network trainings for binned-likelihood analyses at the LHC

We propose a neural network training method capable of accounting for the effects of systematic variations of the data model in the training process and describe its extension towards neural network multiclass classification. The procedure is evaluated on the realistic case of the measurement of Higgs boson production via gluon fusion and vector boson fusion in the τ τ decay channel at the CMS experiment. The neural network output functions are used to infer the signal strengths for inclusive production of Higgs bosons as well as for their production via gluon fusion and vector boson fusion. We observe improvements of 12 and 16% in the uncertainty in the signal strengths for gluon and vector-boson fusion, respectively, compared with a conventional neural network training based on cross-entropy.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Enabling the Broader Use of MOOSE for Nuclear Energy and Other Simulation

This Final Scientific and Technical Report summarizes work performed under the Phase IIA SBIR project “Enabling the Broader Use of MOOSE for Nuclear Energy and Other Simulation” (DE-SC0020906) from August 2023 through August 2025. The objective of the Phase IIA effort was to mature and harden capabilities developed during Phase II, with the goal of enabling practical interoperability between Coreform’s isogeometric analysis (IGA) technologies and the Multiphysics Object-Oriented Simulation Environment (MOOSE), while improving robustness, performance, and scalability for complex, nuclear-relevant geometries. Over the course of Phase IIA, the project established and validated an extraction-based interoperability pathway between Coreform tools and MOOSE. A combined mesh and matrix format was defined collaboratively with MOOSE developers and integrated into the solver, enabling standard MOOSE workflows to operate on data exported from Coreform’s IGA and Flex Representation Method (FRM) pipelines. Early demonstrations validated architectural compatibility using linear solid mechanics problems, while later efforts focused on benchmark testing and external use. By the end of the project period, engineers at BWXT were able to independently set up and execute a simulation using the Coreform–MOOSE workflow and provide direct feedback that informed further refinement. In parallel, substantial effort was devoted to improving the robustness of trimmed U-spline construction for complex CAD geometries. A growing test suite of nuclear-relevant models was compiled through collaboration with multiple stakeholders and used to drive extensive bug fixing and reliability improvements. These efforts resulted in improved robustness and performance, including the addition of fallback capabilities that enhance reliability when the underlying commercial CAD kernel fails. Performance-oriented work progressed later in the project, with the development and demonstration of methods to decompose complex geometries into structured subregions and updated data representations to support more efficient solver processing. Additionally, extensive enhancements to threadsafe parallel data structures and trimming operations established a foundation for scalable processing of large assemblies. Collaboration with Sandia National Laboratories on the SGM geometric modeling kernel advanced to a functioning interface test case, positioning the workflow for future kernel integration. Overall, the Phase IIA effort successfully transitioned the project from architectural proof-of-concept to externally exercised, solver-integrated capability, while clarifying remaining technical challenges related to standardization, performance optimization, and kernel integration.

42 ENGINEERING↗

A machine learning framework for accurate and robust analysis of radiation detector pulses

The microscopic properties of atomic nuclei are used to study various scientific questions. They are essential for understanding the fundamental forces of nature and the chemical evolution of the universe. Detecting decay radiation from radioactive nuclei makes it possible to probe these fundamental nuclear properties. Detector waveform traces may contain additional information about the radiation. Generally, advanced signal processing techniques are needed to extract this additional information, often involving fitting the waveform with model response functions using non-linear least-squares optimization with second-order gradient methods. While this is a powerful technique, it is also computationally expensive, leading to slow processing time, which scales with the volume of data. To address this problem, we have developed a machine learning (ML) approach that infers the characteristics of traces from a model detector response function. In particular, we are interested in classifying whether a single recorded trace consists of one or two pulse constituents and estimating the pulse parameters. Furthermore, our proposed ML method can precisely extract the pulses’ parameters, such as energy and timing information, and accurately classify the pulse multiplicity of a trace. Unlike non-learning-based approaches, our ML approach uses neural networks that are significantly faster at inference, as they do not require any optimization during this stage.

Curve fitting↗

Enhancing Unknown Waveform Detection by Learning Intra and Inter-domain Dependencies with Advanced Attention Fusion Mechanisms

Detection of unknown waveforms in mission-critical communications is a crucial area of interest for the Department of Energy (DoE). Traditional methods and recent deep learning-based approaches often assume that the training set includes all possible classes, which is impractical for detecting new waveforms. This limitation gives rise to the problem of open-set recognition (OSR), which involves correctly identifying known classes while detecting and rejecting unknown or unseen classes. To address this limitation, we propose a novel dual-domain complex-valued neural architecture that jointly processes time-domain and frequency-domain signal representations using transformer mechanisms. A transformer model is a deep learning architecture that uses self-attention mechanisms to process and learn relationships in sequential data. Our model employs a cosine similarity loss to extract domain-specific features and incorporates a transformer architecture in the latent space to weigh the importance of different features from the time and frequency domains. The transformer layer includes stacked self-attention and cross-attention modules to learn intra-domain and inter-domain dependencies, creating a more holistic signal representation. An attention-based fusion module intelligently combines the time and frequency-domain features using multi-head attention, enabling the network to learn the optimal feature for each domain in each input signal. Quantitative results demonstrate the impact of these architectural choices on overall performance, showing significant improvement after incorporating self and cross-attention modules and using complex attention fusion over simple weighted fusion. Our ongoing work will focus on addressing the limitations of threshold-based OSR methods by developing a novel generative framework that integrates a conditional diffusion probabilistic model (DPM). DPM is a generative framework that learns to synthesize complex data by reversing a gradual noising process using a neural network trained to denoise step-by-step. Our goal is to leverage the inherent strengths of DPMs for identifying unknown signals more robustly. One primary advantage of using a DPM is its ability to provide a more reliable anomaly score based on the model's reconstruction error, rather than relying solely on classifier confidence. Additionally, the iterative denoising process of DPMs makes this approach naturally resilient to low Signal-to-Noise Ratio (SNR) conditions, where traditional methods often fail. By implementing this generative framework, we aim to enhance the model's capability to accurately detect unknown waveforms and maintain performance in challenging environments.

99 - GENERAL AND MISCELLANEOUS↗