Search NASA⌕ Search

SEARCH · Search NASA

Results for “data processing methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Estimation of backgrounds from jets misidentified as τ-leptons using the Universal Fake Factor method with the ATLAS detector

Processes with τ$$\tau $$-leptons in the final state are important for Standard Model measurements and searches for physics beyond the Standard Model. The ATLAS experiment at the Large Hadron Collider observes τ$$\tau $$-leptons produced in proton–proton collisions only through their decay products. Data analyses involving hadronically decaying τ$$\tau $$-leptons face challenges due to backgrounds from jets misidentified as τ$$\tau $$-leptons that are not modelled reliably by Monte Carlo simulations. Data-driven methods such as the fake-factor method allow such misidentified backgrounds to be predicted by measuring transfer factors, known as fake factors, in data from dedicated regions. This paper describes a refined technique for determining the fake factors, the Universal Fake Factor method. It evaluates the fake factors for a signal region by using fake factors from samples enriched in different sources of jets misidentified as τ$$\tau $$-leptons (light-quark, gluon, b-quark, and pile-up jets). Each fake factor is calculated as a linear combination of fake factors measured in these different enriched samples. For the full Run 2 data set, the systematic uncertainty of the calculated fake factors, evaluated using W(μν)$$W(\mu u )$$ enriched event sample, ranges from 15 to 35% depending on the τ$$\tau $$-lepton’s transverse momentum and charged-particle decay multiplicity.

Aad, G↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Validation Data for Benchmarking Wire Arc Additive Manufacturing Process Simulations

Residual stresses cause geometric distortion and affect mechanical performance of additively manufactured structures, yet they are notoriously difficult to assess and predict. Distortion (warpage) can drive parts outside dimensional tolerance limits, leading to part rejection or rework. For parts that meet tolerance, locked-in residual stress fields can affect structural integrity during operation, particularly subcritical cracking by fatigue, creep, or corrosion. This work develops benchmark data for a common additive manufacturing process (Wire Arc Additive Manufacturing) that can be applied for calibration and validation of physical process models that predict residual stress fields. The work includes design of two different samples of differing geometry, detailed manufacturing records for a set of physical samples, and an extensive set of residual stress measurement data developed using two diverse techniques (the contour method and neutron diffraction). An initial application of the work is also reported, where a modeling challenge was issued to secure residual stress model predictions from two independent laboratories that were blind to residual stress measurement data. These initial blind residual stress predictions show significant discrepancies relative to the measurement data, illustrating the potential value of the underlying validation data. An open repository for this work, including the sample designs, manufacturing process records, and the residual stress data, is also provided for future application in non-blind validation efforts.

36 MATERIALS SCIENCE↗

Adaptive anomaly detection for identifying attacks in cyber-physical systems: A systematic literature review

Modern cyberattacks in cyber-physical systems (CPS) rapidly evolve and cannot be deterred effectively with most current methods, which focus on characterizing past threats. Adaptive anomaly detection (AAD) is among the most promising techniques to detect evolving cyberattacks, with an emphasis on fast data processing and model adaptation. AAD has been researched extensively; however, to the best of our knowledge, our work is the first systematic literature review (SLR) on current research in this field. We present a comprehensive SLR, gathering 397 relevant papers and systematically analyzing 65 of them (47 research and 18 survey papers) on AAD in CPS from 2013 to November 2023. We introduce a novel taxonomy considering attack types, CPS application, learning paradigm, data management, and algorithms. Our findings show that most studies addressed either model adaptation or data processing, but rarely both simultaneously. This indicates a research gap in fully adaptive solutions. We also categorize algorithms, datasets, and attack characteristics, and summarize strengths and weaknesses across the literature. Our review provides a structured and accessible reference for researchers and practitioners, offering insights into key trends and highlighting limitations in current approaches. Finally, we outline several future research directions, including the need for integrated real-time processing and adaptive learning, explainability, and uncertainty quantification in AAD for CPS.

Adaptation↗

Optimization of direct air capture processes using reactive transport models of adsorption-desorption cycles

In this study, we develop and implement a reactive transport model in COMSOL Multiphysics® to address the challenges of direct air carbon capture. The model is validated against experimental data and used to simulate the cyclic steady state of the adsorption-desorption process. The optimization of this model is achieved through advanced trust-region methods integrated with Gaussian Processes. Key decision variables, including adsorption and desorption times, desorption temperature and pressure, input velocity, bed porosity, column length, and radius were optimized to minimize the capture cost. After optimization, a sensitivity analysis revealed the complex interplay between the decision variables and their effect on the specific energy and cost of removing the CO 2 . We optimized the capture cost while taking into account the trade-off between energy consumption and productivity. The resulting minimum capture cost was determined to be 265.2 $/t-CO 2 , which aligns with expected values reported in the literature. Numerical results suggest the effectiveness of the optimization strategies applied, and underscore the importance of simultaneous decision variable selection in improving the performance in direct air capture processes. We also extend the modeling approach to a 2D axisymmetric model to better visualize CO₂ uptake and temperature profiles, revealing significant radial gradients during the regeneration step. As a main drawback, this enhanced model comes with a computational cost approximately 40 times higher than that of the 1D model.

Adsorption-desorption process↗

A critical review on additive manufacturing of refractory alloys from a data analytics perspective- beyond nickel-based superalloys

Refractory alloys (RAs) are promising materials due to their exceptional physicochemical properties, but most research remains at the laboratory scale. For broader adoption, advancements in manufacturing are essential. Because their high stability makes conventional methods like machining and casting difficult, additive manufacturing (AM) is emerging as an effective approach for fabricating refractory alloy components. However, AM's repeated non-equilibrium thermal cycles introduce undesired features (e.g. defects, anisotropic microstructures, and residual stresses), which are magnified due to RAs’ unique properties. This paper comprehensively reviews the state-of-the-art methods of AM for refractory alloys. It explores data analytics techniques to establish design rules based on multi-fidelity experimental and computational methods. Furthermore, it investigates integrated, collaborative efforts to harmonise standalone databases, information, knowledge, and predictive models at multi-physics, multi-stage, and multi-scale. Unlike the existing literature that focuses primarily on material systems or process fundamentals, this work provides an integrated perspective on AM of refractory alloys from a data analytics standpoint, highlighting the roles of integrated computational materials engineering (ICME), verification, validation, and uncertainty quantification (VV&UQ), and digital twin-driven qualification in overcoming data scarcity and accelerating rapid qualification.

Additive manufacturing↗

sourcePy

Pollutant source identification techniques (of which there are many variations) are either locked behind researchers writing their own code for each use case or GUI platforms that are easy to use but inflexible and opaque. The Python package sourcePy brings together many of the pollutant source identification algorithms, giving the user full control out of the box. It aims to create a platform for source identification experiments where the full analysis from beginning to end can be done in Python, with a level of specificity in design that isn't available in the GUI options. sourcePy provides users with a few key features: -A Python interface with HYSPLIT, which can be used to generate trajectories and concentration plumes -Several Python classes which standardize the preparation and processing of data related to source identification experiments -Example scripts and notebooks that allow even new python users to get started with their own experiments quickly -Visualization methods

Arseneau, Isaac [Oak Ridge National Laboratory (OR↗

Modulated Thermomechanical Analysis of Compression-Molded High-Density Polyethylene

Thermomechanical analysis (TMA) experiments conducted on high-density polyethylene (HDPE) show both reversible and irreversible dimensional changes. To further explore these reversible and irreversible processes, modulated thermomechanical analysis (MTMA) was used. Before reliable data on compression-molded HDPE was collected, a parameter optimization was performed to obtain a suitable MTMA method. Once a suitable method was obtained, several MTMA experiments were conducted on compression-molded HDPE. This work highlights the steps taken during the MTMA parameter optimization and the results obtained from MTMA experiments conducted on pristine compression-molded HDPE samples.

36 MATERIALS SCIENCE↗

Efficient Signal Processing in BOTDA: Utilizing PCA and PCA-Based Neural Networks for Temperature Monitoring

This work presents a comparative analysis of the various signal processing techniques used in the Brillouin gain spectrum (BGS) peak estimation. Traditional fitting methods such as Lorentzian curve fitting (LCF) are slow and less effective in noisy data. PCA-based methods were tested on the experimental data: A Euclidian distance-based approach, and a probabilistic deep neural network (PDNN) based approach, both using 5 principal components to represent a single BGS. Both methods significantly reduce computational time with respect to LCF, whereas PDNN offers uncertainty insights along with the parameter value. Measuring a range of temperatures, analyzing accuracy, and speed, it can be concluded that PCA trained PDNN outperforms other methods, and appears to be helpful in scenario where large datasets are generated.

Brillouin optical time domain analysis↗

Compressed sensing methods with applications to advanced air sampling

Environmental sampling methods developed by the Savannah River National Laboratory (SRNL) employ collectors with sorbent media tubes set at various locations to collect airborne emissions. Laboratory analyses of these tubes results in one-dimensional signals regarding what chemicals are being released and transported within the atmosphere. The analysis process is time consuming especially when analyzing a full year’s worth of tubes (hourly sample collection results in nearly 9,000 tubes per year). Using a signal processing method such as compressed sensing allows for recreation of the full signal while greatly reducing the number of analyzed samples required. Due to the sparsity of data retrieved from the air tubes, it is possible to use measurements a fraction of the size of the original data to gain much of the same information. This would improve the overall time and cost of analysis when modeling one-dimensional sampling signals.

54 ENVIRONMENTAL SCIENCES↗

Compressed Sensing Methods with Applications to Advanced Air Sampling [Poster]

Environmental sampling methods developed by the Savannah River National Laboratory (SRNL) employ collectors with sorbent media tubes set at various locations to collect airborne emissions. Laboratory analyses of these tubes results in one-dimensional signals regarding what chemicals are being released and transported within the atmosphere. The analysis process is time consuming especially when analyzing a full year’s worth of tubes (hourly sample collection results in nearly 9,000 tubes per year). Using a signal processing method such as compressed sensing allows for recreation of the full signal while greatly reducing the number of analyzed samples required. Due to the sparsity of data retrieved from the air tubes, it is possible to use measurements a fraction of the size of the original data to gain much of the same information. This would improve the overall time and cost of analysis when modeling one-dimensional sampling signals.

Campbell, Cassidy [Savannah River National Laborat↗

Data-driven organic solubility prediction at the limit of aleatoric uncertainty

Abstract Small molecule solubility is a critically important property which affects the efficiency, environmental impact, and phase behavior of synthetic processes. Experimental determination of solubility is a time- and resource-intensive process and existing methods for in silico estimation of solubility are limited by their generality, speed, and accuracy. This work presents two models derived from the FASTPROP and CHEMPROP architectures and trained on BigSolDB which are capable of predicting solubility at arbitrary temperatures for a wide range of small molecules in organic solvent. Both extrapolate to unseen solutes 2–3 times more accurately than the current state-of-the-art model and we demonstrate that they are approaching the aleatoric limit (0.5–1$$\log S$$ log S ) of available test data, suggesting that further improvements in prediction accuracy require more accurate datasets. The FASTPROP-derived model (called FASTSOLV) and the CHEMPROP-based model are open source, freely accessible via a Python package and web interface, highly reproducible, and up to 2 orders of magnitude faster than current alternatives.

Science & Technology - Other Topics↗

Semi-automatic image annotation using 3D LiDAR projections and depth camera data

Efficient image annotation is necessary to utilize deep learning object recognition neural networks in nuclear safeguards, such as for the detection and localization of target objects like nuclear material containers (NMCs). This capability can help automate the inventory accounting of different types of NMCs within nuclear storage facilities. The conventional manual annotation process is labor-intensive and time-consuming, hindering the rapid deployment of deep learning models for NMC identifications. This paper introduces a novel semi-automatic method for annotating 2D images of nuclear material containers (NMCs) by combining 3D light detection and ranging (LiDAR) data with color and depth camera images collected from a handheld scan system. The annotation pipeline involves an operator manually marking new target objects on a LiDAR-generated map, and projecting these 3D locations to images, thereby automatically creating annotations from the projections. The semi-automatic approach significantly reduces manual efforts and the expertise in image annotation that is required to perform the task, allowing deep learning models to be trained on-site within a few hours. The paper compares the performance of models trained on datasets annotated through various methods, including semi-automatic, manual, and commercial annotation services. The evaluation demonstrates that the semi-automatic annotation method achieves comparable or superior results, with a mean average precision (mAP) above 0.9, showcasing its efficiency in training object recognition models. Additionally, the paper explores the application of the proposed method to instance segmentation, achieving promising results in detecting multiple types of NMCs in various formations.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Safe Reinforcement Learning-Based Transient Stability Control for Islanded Microgrids With Topology Reconfiguration

This paper proposes a safe reinforcement learning (RL)-based transient stability emergency control (TSEC) method for islanded microgrids. RL requires extensive interaction with the environment to learn control strategies, hence, a data-driven approach is used as a substitute for time-consuming time-domain simulation calculations. Deep sigma point processes (DSPP), which is a Gaussian process model, is utilized to predict the normal distribution of transient stability of microgrids and to construct a transient stability chance constraint. Reward-constrained policy optimization (RCPO) can simultaneously achieve objective prediction, policy learning, and constraint cost coefficient update across multiple timescales. RCPO interacts with the DSPP-based microgrid environment through a multi-process parallel manner, greatly increasing the training speed. Case studies on a real islanded microgrid demonstrate that the proposed method can efficiently and quickly obtain the optimal emergency control strategy while adhering to all hard constraints.

14 SOLAR ENERGY↗

Part-scale evolution of fine-scale microstructural heterogeneity in solid-state additive manufacturing

Current solid-state additive manufacturing methods, refined through costly and time-consuming trial and error, have spurred interest in computational models that replicate material behavior under typical thermomechanical conditions (e.g., strain-rate ~ 102 s−1). These models, however, struggle to capture time-dependent microstructural evolution. In this work, Additive Friction-Stir Deposition (AFSD) is used as a representative case study for part-scale quantification of microstructure evolution at a fine spatial resolution (200 μm) by examining a liquid-nitrogen-cooled stop-action build via energy-dispersive X-ray diffraction coupled with a multi-channel detector. These results inform modeling efforts by linking process asymmetry to stored plastic strain, residual elastic strain, and texture development, and unlike current state-of-the-art characterization methods (e.g., EBSD or neutron diffraction), this approach provides both the spatial resolution and collection efficiency necessary to quantify fine-scale microstructural heterogeneity over large component volumes. As such, this technique provides essential validation data for computational models, e.g., crystal plasticity, enabling future prediction of heterogeneous behavior in AFSD and other additive manufacturing processes.

Franz, Cole [ORNL] (ORCID:0000000213465881)↗

Incorporating Elevation in Traffic-Vehicle CO-Simulation: Issues, Impacts, and Solutions

Traffic-vehicle co-simulation couples microscopic traffic simulation with full-body vehicle dynamics to assess system-level impacts on mobility, energy, and safety with greater realism. Incorporating elevation is critical for accurately modeling vehicle behavior and energy use, especially for gradient-sensitive vehicles such as electric and heavy-duty trucks. However, raw elevation data often contain noise, discontinuities, and inconsistencies. While such issues may be negligible in traditional traffic simulations, they significantly affect traffic-vehicle co-simulations where vehicle dynamics are sensitive to road grade variations. This paper investigates the impact of unprocessed elevation data on vehicle behavior and energy consumption using a 42-mile simulation along Interstate 81. We propose an elevation processing workflow that can mitigate the effects stem from elevation data issues, improving the realism and stability of traffic-vehicle co-simulation. Results show that the method effectively removes noise and abrupt elevation transitions while preserving roadway geometry.

Xu, Guanhao [ORNL] (ORCID:0000000214326357)↗

SARS-CoV-2 wastewater variant surveillance: pandemic response leveraging FDA’s GenomeTrakr network

ABSTRACT Wastewater surveillance has emerged as a crucial public health tool for population-level pathogen surveillance. Supported by funding from the American Rescue Plan Act of 2021, the FDA‘s genomic epidemiology program, GenomeTrakr, was leveraged to sequence SARS-CoV-2 from wastewater sites across the United States. This initiative required the evaluation, optimization, development, and publication of new methods and analytical tools spanning sample collection through variant analyses. Version-controlled protocols for each step of the process were developed and published on protocols.io. A custom data analysis tool and a publicly accessible dashboard were built to facilitate real-time visualization of the collected data, focusing on the relative abundance of SARS-CoV-2 variants and sub-lineages across different samples and sites throughout the project. From September 2021 through June 2023, a total of 3,389 wastewater samples were collected, with 2,517 undergoing sequencing and submission to NCBI under the umbrella BioProject,PRJNA757291. Sequence data were released with explicit quality control (QC) tags on all sequence records, communicating our confidence in the quality of data. Variant analysis revealed wide circulation of Delta in the fall of 2021 and captured the sweep of Omicron and subsequent diversification of this lineage through the end of the sampling period. This project successfully achieved two important goals for the FDA’s GenomeTrakr program: first, contributing timely genomic data for the SARS-CoV-2 pandemic response, and second, establishing both capacity and best practices for culture-independent, population-level environmental surveillance for other pathogens of interest to the FDA. IMPORTANCE This paper serves two primary objectives. First, it summarizes the genomic and contextual data collected during a Covid-19 pandemic response project, which utilized the FDA’s laboratory network, traditionally employed for sequencing foodborne pathogens, for sequencing SARS-CoV-2 from wastewater samples. Second, it outlines best practices for gathering and organizing population-level next generation sequencing (NGS) data collected for culture-free, surveillance of pathogens sourced from environmental samples.

Microbiology↗

Discovering the Unknowns: A First Step

This article aims at discovering the unknown variables in the system through data analysis. The main idea is to use the time of data collection as a surrogate variable and try to identify the unknown variables by modeling gradual and sudden changes in the data. We use Gaussian process modeling and a sparse representation of the sudden changes to efficiently estimate the large number of parameters in the proposed statistical model. The method is tested on a realistic dataset generated using a one-dimensional implementation of a Magnetized Liner Inertial Fusion (MagLIF) simulation model, and encouraging results are obtained.

42 ENGINEERING↗