Search NASASearch

SEARCH · Search NASA

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Preprocessing for Eddy Dissipation Rate and TKE Profile Generation

The Aircraft Vortex Spacing System (AVOSS), a set of algorithms to determine aircraft spacing according to wake vortex behavior prediction, requires turbulence profiles to appropriately determine arrival and departure aircraft spacing. The ambient atmospheric turbulence profile must always be produced, even if the result is an arbitrary (canned) profile. The original turbulence profile code was generated By North Carolina State University and used in a non-real-time environment in the past. All the input parameters could be carefully selected and screened prior to input. Since this code must run in real-time using actual measurements in the field as input, it became imperative to begin a data checking and screening process as part of the real-time implementation. The process described herein is a step towards ensuring that the best possible turbulence profile is always provided to AVOSS. Data fill-ins, constant profiles and arbitrary profiles are used only as a last resort, but are essential to ensure uninterrupted application of AVOSS.

Zak, J. Allen

Intelligent Text Retrieval and Knowledge Acquisition from Texts for NASA Applications: Preprocessing Issues

A system that retrieves problem reports from a NASA database is described. The database is queried with natural language questions. Part-of-speech tags are first assigned to each word in the question using a rule based tagger. A partial parse of the question is then produced with independent sets of deterministic finite state a utomata. Using partial parse information, a look up strategy searches the database for problem reports relevant to the question. A bigram stemmer and irregular verb conjugates have been incorporated into the system to improve accuracy. The system is evaluated by a set of fifty five questions posed by NASA engineers. A discussion of future research is also presented.

Source record

Preprocessing Inconsistent Linear System for a Meaningful Least Squares Solution

Mathematical models of many physical/statistical problems are systems of linear equations. Due to measurement and possible human errors/mistakes in modeling/data, as well as due to certain assumptions to reduce complexity, inconsistency (contradiction) is injected into the model, viz. the linear system. While any inconsistent system irrespective of the degree of inconsistency has always a least-squares solution, one needs to check whether an equation is too much inconsistent or, equivalently too much contradictory. Such an equation will affect/distort the least-squares solution to such an extent that renders it unacceptable/unfit to be used in a real-world application. We propose an algorithm which (i) prunes numerically redundant linear equations from the system as these do not add any new information to the model, (ii) detects contradictory linear equations along with their degree of contradiction (inconsistency index), (iii) removes those equations presumed to be too contradictory, and then (iv) obtain the minimum norm least-squares solution of the acceptably inconsistent reduced linear system. The algorithm presented in Matlab reduces the computational and storage complexities and also improves the accuracy of the solution. It also provides the necessary warning about the existence of too much contradiction in the model. In addition, we suggest a thorough relook into the mathematical modeling to determine the reason why unacceptable contradiction has occurred thus prompting us to make necessary corrections/modifications to the models - both mathematical and, if necessary, physical.

Sen, Syamal K.

Preprocessing in Matlab Inconsistent Linear System for a Meaningful Least Squares Solution

Mathematical models of many physical/statistical problems are systems of linear equations~ Due to measurement and possible human errors/mistakes in modeling/data, as well as due to certain assumptions to reduce complexity, inconsistency (contradiction) is injected into the model, viz. the linear system. While any inconsistent system irrespective of the degree of inconsistency has always a least-squares solution, one needs to check whether an equation is too much inconsistent or, equivalently too much contradictory. Such an equation will affect/distort the least-squares solution to such an extent that renders it unacceptable/unfit to be used in a real-world application. We propose an algorithm which (i) prunes numerically redundant linear equations from the system as these do not add any new information to the model, (ii) detects contradictory linear equations along with their degree of contradiction (inconsistency index), (iii) removes those equations presumed to be too contradictory, and then (iv) obtain the . minimum norm least-squares solution of the acceptably inconsistent reduced linear system. The algorithm presented in Matlab reduces the computational and storage complexities and also improves the accuracy of the solution. It also provides the necessary warning about the existence of too much contradiction in the model. In addition, we suggest a thorough relook into the mathematical modeling to determine the reason why unacceptable contradiction has occurred thus prompting us to make necessary corrections/modifications to the models - both mathematical and, if necessary, physical.

Sen, Symal K.

Independent Component Analysis for Preprocessing Optical Signals in Support of Multi-User Communication

In an effort to construct an optical transceiver for supporting multi-point communication, researchers developed large field-of-view FSO transceivers with wide apertures for facilitating rapid acquisition and optical signal tracking. The design is constructed from fiber-bundle(s) intelligently arranged to provide multiple optical pathways within a single transceiver, with each pathway limited to a particular optical directionality. Because designs with wide apertures are susceptible to receiving several optical signals simultaneously from multiple transmitters, accurately decoding individual signals and achieving true multi-user communication can be difficult. The work detailed in this paper investigated the use of blind source separation techniques, mainly independent component analysis (ICA), to estimate transmitted signals from their observed, combined mixtures sans information. Effects of signal power, format, data rate, wavelength, and turbulence severity on signal separation, as well as signal demodulation accuracy was also analyzed using experimental optical signals and corresponding mixtures. Signal generators, laser sources, turbulence emulator, beam profiler, photodiodes, and oscilloscopes comprised the experiment setup. Optical signals with varying power, format, and wavelength were generated and propagated through turbulence with varying degree severity, and finally received by the fiber-bundle-based optical transceiver. Results demonstrate that an ICA technique can accurately process the combined signals into corresponding individual components. The investigation further identified signal characteristics (e.g., power levels, format, data rate) and turbulence severity under which ICA is unable to accurately perform signal separations.

Optics

Summary of the 1st AIAA Geometry and Mesh Generation Workshop (GMGW-1) and Future Plans

The 1st AIAA Geometry and Mesh Generation Workshop (GMGW-1) was held in conjunction with the AIAA Aviation Forum and Exposition 2017 and in collaboration with the 3rd AIAA Computational Fluid Dynamics (CFD) High Lift Prediction Workshop (HiLiftPW-3). As the first AIAA workshop on these topics, GMGW-1's broad objectives were to assess the current state-of-the art in geometry preprocessing and mesh generation technology as well as software as applied to aircraft and spacecraft systems. The workshop was intended to identify and develop understanding of areas of needed improvement in terms of performance, accuracy, and applicability. It was also to provide a foundation for documenting best practices for geometry preprocessing and mesh generation. The genesis of GMGW-1 is found in the indictments levied against geometry preprocessing and mesh generation - not undeservedly - by the NASA CFD Vision 2030 Study. In order to create a reference against which future progress in geometry preprocessing and mesh generation can be measured, the organizers of GMGW-1, with the assistance of the organizers of HiLiftPW- 3, focused GMGW-1 on generation of meshes of the NASA High Lift Common Research Model (HL-CRM). Some of the generated meshes were provided for use by the participants in HiLiftPW-3. All meshes and the processes by which they were generated were analyzed by GMGW-1 as a first assessment of state of the art practices. The results of GMGW-1 added quantitative detail to known problem areas including geometry modeling, data interoperability, and amount of human intervention. They do provide a clear path toward a vision of geometry preprocessing and mesh generation in the year 2030. The next milepost along this path will be a second workshop.

Chawner, John R.

Analysis Ready Data in Analytics Optimized Data Stores for Analysis of Big Earth Data in the Cloud

Cloud computing offers the possibility of making the analysis of Big Data approachable for a wider community due to affordable access to computing power, an ecosystem of usable tools for parallel processing, and migration of many large datasets to archives in the cloud, allowing data-proximal computing. Generally, data analysis acceleration in the cloud comes from running multiple nodes in a split-combine-apply strategy. Data systems such as the Earth Observing System Data and Information System are in a position to "pre-split" the data by storing them in a data store that is optimized for data parallel computing, i.e., an Analytics-Optimized Data Store (AODS). A variety of approaches to AODS are possible, from highly scalable databases to scalable filesystems to data formats optimized for cloud access (e.g., zarr and cloud-optimized datasets), with the optimal choice dependent on both the types of analysis and the geospatial structure of the data. A key question is how much preprocessing of the data to do, both before splitting and as the first part of the apply step. Again, the geospatial structure of the data and the analysis type influence the decision, with the added complexity of the user type. Trans-disciplinary users who are not well-versed in the nuances of quality-filtering and georeferencing of remote sensing orbit/swath/scene data tend to ask for more highly processed data, relying on the data provider to make sensible decisions on preprocessing parameters. (This accounts for the popularity of "Level 3" gridded data, despite the lower spatial resolution it provides.) In this case, data can be preprocessed before the split, resulting in higher performance in the rest of the "apply" step, which can be transformative for use cases such as interactive data exploration at scale. Discipline researchers who are experienced with remote sensing data often prefer more flexibility in customizing the preprocessing data into Analysis Ready Data, resulting in more need for on-the-fly preprocessing.

Lynnes, Christopher

A new approach to telemetry data processing

An approach for a preprocessing system for telemetry data processing was developed. The philosophy of the approach is the development of a preprocessing system to interface with the main processor and relieve it of the burden of stripping information from a telemetry data stream. To accomplish this task, a telemetry preprocessing language was developed. Also, a hardware device for implementing the operation of this language was designed using a cellular logic module concept. In the development of the hardware device and the cellular logic module, a distributed form of control was implemented. This is accomplished by a technique of one-to-one intermodule communications and a set of privileged communication operations. By transferring this control state from module to module, the control function is dispersed through the system. A compiler for translating the preprocessing language statements into an operations table for the hardware device was also developed. Finally, to complete the system design and verify it, a simulator for the collular logic module was written using the APL/360 system.

Broglio, C. J.

Are the U.S. Biorefineries Over the Hurdle of 2000 Ton Daily Throughput Yet?

The efficient utilization of lignocellulosic biomass for biofuel and biochemical production is hindered by material handling issues such as clogging and segregation among other challenges. Preprocessing methods such as drying, screening, and milling have improved conversion yield but have not sufficiently enhanced flowability, especially herbaceous biomass. The poor flowability of herbaceous biomass is rooted in some particle attributes that remain less altered by those methods, e.g., irregular particle shape, high roughness, and high compressibility, making it hard to scale up throughput to a key benchmark for a biorefinery – 2000 ton per day. Applying additional preprocessing methods like pelletization and torrefaction to drastically change those particle attributes can improve flow and handling but has not been comprehensively verified through test. The flowability of herbaceous biomass feedstock formats generated by three different preprocessing methods was recently assessed at Idaho National Laboratory’s Biomass Feedstock National User Facility: first, loose particles size reduced from as-received materials; second, pellets produced from an efficient densification process; and third, powders milled from torrefied pellets. Benchmarking tests including static angle of repose, basic flow energy measured in a powder rheometer, and discharge flow in an adjustable hopper, were conducted to evaluate those feedstock formats. Beyond the capacity of existing experimental apparatuses, a digital engineering approach involving flow simulations and AI models were used to identity the material attributes and processing parameters that have dominant influences on flow throughput. Techno-economic analysis focusing on hopper flow as a typical material handling operation was conducted for those feedstock formats. Perspectives will be discussed on whether the 2000-ton daily throughput for a biorefinery is achievable at an acceptable cost by using any of the tested preprocessing methods.

09 - BIOMASS FUELS

Techno-economic and life-cycle analysis of strategies for improving operability and biomass quality in catalytic fast pyrolysis of forest residues

Many of the challenges faced by the first commercial biorefineries were associated with feedstock handling, quality, and cost. Strategies are needed to enable further expansion of biorefineries and meet the growing demand for bio-based fuels and products. Here, we examine 2 key feedstock challenges and mitigation strategies in the context of a catalytic fast pyrolysis (CFP) biorefinery: (1) the operability of the feed system, which may be improved by modifying the minimum particle size fed to the reactor, and (2) the quality of the biomass, which may be improved by employing air classification to remove undesirable material and increase fuel yields. We conduct techno-economic analysis (TEA) and life-cycle analysis for these strategies, employing a discrete event simulation model for biomass preprocessing combined with a series of correlations developed from literature data and a rigorous CFP conversion model. Our results highlight the importance of balancing increased cost and material losses from preprocessing against improved operability and fuel yields. Economics and sustainability were optimized when operating at the lowest minimum particle size, emphasizing the importance of minimizing material losses while maintaining the operability of the process. Economically, additional costs and material losses from air classification could be acceptable due to improved biomass conversion, and an optimum air classification speed was identified; however, the fuel GHG emissions were minimized when air classification was not used. Valorizing material removed during preprocessing as a coproduct could improve economics and sustainability, decreasing the burden of material losses.

09 - BIOMASS FUELS

Accurate and Data‐Efficient Micro X‐ray Diffraction Phase Identification Using Multitask Learning: Application to Hydrothermal Fluids

Traditional analysis of highly distorted micro X‐ray diffraction (μ‐XRD) patterns from hydrothermal fluid environments is a time‐consuming process, often requiring substantial data preprocessing and labeled experimental data. Herein, the potential of deep learning with a multitask learning (MTL) architecture to overcome these limitations is demonstrated. MTL models are trained to identify phase information in μ‐XRD patterns, minimizing the need for labeled experimental data and masking preprocessing steps. Notably, MTL models show superior accuracy compared to binary classification convolutional neural networks. Additionally, introducing a tailored cross‐entropy loss function improves MTL model performance. Most significantly, MTL models tuned to analyze raw and unmasked XRD patterns achieve close performance to models analyzing preprocessed data, with minimal accuracy differences. This work indicates that advanced deep learning architectures like MTL can automate arduous data handling tasks, streamline the analysis of distorted XRD patterns, and reduce the reliance on labor‐intensive experimental datasets.

97 MATHEMATICS AND COMPUTING

Comparative life cycle assessment of woody biomass processing: air classification, drying, and size reduction powered by bioelectricity versus grid electricity

Sulfur accumulation during biofuel production is pollutive and toxic to conversion catalysts and causes the premature breakdown of processing equipment. Air classification is an effective preprocessing technology for ash and sulfur reduction from biomass feedstocks. Here, a life cycle assessment (LCA) sought to understand the environmental impact of implementing air classification as a sulfur-mitigation technique to improve feedstock quality for pine residues using a grid electricity scenario (GES) versus a bioelectricity scenario (BES). Global warming potential (GWP) for preprocessing was simulated using inventory databases embedded in SimaPro and the Argonne National Laboratory’s GREET model, specifically focusing on comparing the GWP of a GES versus a BES. Overall, the GES had a GWP impact over seven times that of the BES (136 versus 18 kg CO 2 equivalent per tonne of usable feedstock), with steam generation during rotary drying accounting for 57% of the GES’s GWP. Air classification represents 0.4% and 1.6% of the total GWP impact for the GES and BES, respectively. Therefore, air classification can facilitate a 30% reduction in feedstock sulfur content to improve feedstock quality for biofuel conversion and lessen corrosion of equipment while contributing minimal GWP impact during preprocessing.

Air classification

Subject-specific modeling framework for particle deposition using computational fluid dynamics

Quantifying particle deposition and dose in the respiratory tract requires a physiologically realistic representation and reproducible computational workflows. However, existing modeling frameworks, such as the International Commission on Radiological Protection (ICRP) compartmental models and the Multiple Path Particle Dosimetry (MPPD) tool, lack detailed deposition profiles and subject-specific capabilities. The combination of advances in computer vision algorithms applied to the respiratory tract and Computational Fluid and Particle Dynamics (CFPD) allows high-fidelity simulations of particle behavior in anatomically accurate geometries derived from individual CT scans. The segmentation, preprocessing, and file preparation task for a CFPD simulation was often time-consuming, and no prior studies to-date have yet presented a fully automated framework. This work presents a fully automated workflow to obtain individualized particle deposition profiles in the human respiratory tract. The pipeline starts with segmenting upper and lower airway geometries using morphological and deep learning-based methods, generating three-dimensional (3D) models from CT imaging data. Next, a series of algorithms are presented to quality check and prepare the 3D geometry for a CFD or CFPD simulation. The preprocessing step includes correcting geometric artifacts, enforcing a physically consistent mesh, and automatically identifying and capping multiple outlets, which is required for CFD/CFPD simulations. These processed models are then input into open-source (OpenFOAM) or commercial (StarCCM+) CFD solvers, where flow and transient particle transport equations — including turbulence and particle–wall interactions are solved under realistic breathing conditions. Finally, the resulting particle deposition profiles can be integrated with Monte Carlo radiation transport codes and state-of-the-art computational phantoms to assess organ-specific absorbed doses in scenarios of radioactive aerosol inhalation. The presented work streamlines respiratory tract segmentation, preprocessing for CFD/CFPD simulations, and integration with dose assessment workflows, reducing manual intervention and improving access to high-fidelity, subject-specific modeling. The high precision in predicted particle deposition and dose distributions can improve personalized treatment strategies in respiratory medicine and refine dose estimates for radiation protection.

AI

Performance Comparison of Machine Learning Models for Ultrasonic Nondestructive Evaluation of Alkali-Silica Reaction in Concrete

Alkali-silica reaction (ASR) causes concrete degradation, leading to cracking, rebar corrosion, and reduced structural integrity, which raises safety concerns. Ultrasonic nondestructive evaluation (NDE) effectively assesses concrete properties and monitors ASR progression. However, its deployment and analysis require specialized expertise and subjective interpretation. As computational power increases, artificial intelligence (AI) and machine learning (ML) algorithms are increasingly being used to automate NDE data analysis across various industries for AI-assisted automation. Regulatory agencies are adapting to this technological shift, prompting a need to evaluate current ML technologies’ capabilities and limitations in assessing concrete material properties and damage. This report presents a comparative analysis of four ML regression models for predicting concrete material damage induced by ASR expansion using long-term ultrasonic data monitoring. The models investigated include linear regression (LR), support vector regression (SVR), shallow neural networks (NN), and deep neural networks (DNN). LR, SVR, and shallow NN models use features extracted from ultrasonic signals, whereas the DNN model processes time-domain ultrasonic signals and frequency spectra directly. The study systematically compared the models’ performance from various perspectives, including model input, prediction performance, and generalization ability. The findings indicate significant variability in model performance, with some ML algorithms achieving very high or very low prediction accuracy depending on the preprocessing and feature engineering (extraction and selection) applied. Key insights include the observation that shallow ML models (LR, SVR, and shallow NNs) require meticulous preprocessing and feature extraction to achieve high accuracy. In contrast, the DNN model, although it bypasses the need for feature engineering, necessitates extensive preprocessing to mitigate noise and computational demands. The SVR model emerged as the top performer among the shallow models, and the DNN model exhibited superior performance on specific datasets but struggled with generalization across specimens from different batches. Additionally, the SVR model is sensitive to temperature variations, whereas the DNN model is robust in this regard. Using recurrent neural networks is recommended for future ASR expansion prediction studies. Recurrent neural networks’ inherent ability to capture temporal dependencies and long-term patterns makes them well suited for analyzing sequential ultrasonic monitoring data. Overall, the results and conclusions of this study could provide insights into the capabilities and effectiveness of ML when applied to ultrasonic NDE data and help identify best practices for using ML for ultrasonic NDE of concrete material properties.

36 MATERIALS SCIENCE

Layered recognition networks that pre-process, classify, and describe

A brief overview is presented of six types of pattern recognition programs that: (1) preprocess, then characterize; (2) preprocess and characterize together; (3) preprocess and characterize into a recognition cone; (4) describe as well as name; (5) compose interrelated descriptions; and (6) converse. A computer program (of types 3 through 6) is presented that transforms and characterizes the input scene through the successive layers of a recognition cone, and then engages in a stylized conversation to describe the scene.

Uhr, L.