Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam↗

Cooperative Education

Los Alamos National Laboratory (LANL) is a multidisciplinary national laboratory that conducts research and development in national security, engineering, materials science, computational modeling, and advanced manufacturing. The laboratory develops innovative technologies to address complex scientific and engineering challenges. This project focuses on the development and evaluation of high-performance absorbing structures through computational design, simulation, and engineering analysis. Absorbing structures are used in applications where damage mitigation, structural protection, and material efficiency are critical performance requirements. The increasing demand for lightweight, high-strength, and highly efficient structural systems has created a need for improved design methodologies capable of maximizing absorption while minimizing weight and material usage. The project utilizes advanced engineering software, including 3D CAD software and FEA, to generate and optimize structural concepts. Computational simulations are performed to evaluate structural behavior under loading conditions, while mathematical analyses are conducted using Python-based tools as well as established analytical equations from material and structural mechanics. The project benefits LANL by supporting the development of advanced design methodologies and improving the understanding of material and structural performance. During the internship term, a significant portion of the design development, simulation, and data analysis activities will be completed. Success of the project depends on collaboration among engineering mentors and technical staff members. Work will be conducted at Los Alamos National Laboratory using laboratory computing resources and engineering software.

42 ENGINEERING↗

Randomized Algorithms for Symmetric Nonnegative Matrix Factorization

Symmetric Nonnegative Matrix Factorization (SymNMF) is a technique in data analysis and machine learning that approximates a matrix with a product of a nonnegative, low-rank matrix and it transpose. To design faster and more scalable algorithms for SymNMF we develop two randomized algorithms for its computation. The first method uses randomized matrix sketching to compute an initial low-rank approximation to the input matrix and proceeds to uses this as a low-rank input to rapidly compute a SymNMF. The second methods uses randomized leverage score sampling to approximately solve constrained least squares problems. Many successful methods for SymNMF rely on (approximately) solving sequences of constrained least squares problems. Here, we prove theoretically that leverage score sampling can approximately solve constrained least squares problems to e-accuracy. Finally we demonstrate both methods work in practice by applying them to graph clustering tasks on large real world data sets. These experiments show that our methods approximately maintain solution quality and achieve significant speed ups for both large dense and large sparse problems.

97 MATHEMATICS AND COMPUTING↗

PET and polyolefin plastics supply chains in Michigan: present and future systems analysis of environmental and socio-economic impacts

Many actions are underway at global, national, and local levels to increase plastics circularity. However, studies evaluating the environmental and socio-economic impacts of such a transition are lacking at regional levels in the United States. In this work, the existing polyethylene terephthalate and polyolefin plastics supply chains in Michigan were compared to a potential future (‘NextCycle’) scenario that looks at increasing Michigan’s overall recycling rate to 45%. Material flow analysis data was combined with environmental and socio-economic metrics to evaluate the sustainability of these supply chains for the modeled scenarios. Overall, the NextCycle scenario for these supply chains achieved a net 14% and 34% savings of greenhouse gas (GHG) emissions and energy impacts, when compared with their respective baseline values. Additionally, the NextCycle scenario showed a net gain in employment and wages, however, it showed a net loss of revenue generation outside of Michigan due to the avoided use of virgin resins in Michigan.

54 ENVIRONMENTAL SCIENCES↗

Improving the User Interface of the DeepLynx Data Warehouse

DeepLynx is an open-source ontology-based data warehouse created by INL to support the creation and life cycle of digital engineering projects, with a particular emphasis on digital twins [1]. Digital twins are systems that represent physical assets and process in a real-time digital environment [1]. Most well-known commercial data warehouses use Graphical User Interfaces (GUIs) for users to interact with their systems [3]. Limited publications have addressed the design of these interfaces and understanding of their target users. The current users and development team acknowledge the need to improve the current UI, not just for aesthetics but to improve functionality and workflow of DeepLynx. Traditional data warehouse users are developers, data scientists and business analysts [2]. DeepLynx users have a vast range of experience using data warehouses, and diverse roles, including engineers, scientists and management positions. Because there is a broader audience of target users for DeepLynx than a typical data warehouse, it is essential that DeepLynx has a useable and intuitive user interface. To achieve this the team performed human-computer interaction methods, including a Heuristic Evaluation of current UI using Neilsen’s Usability Heuristic, create personas based on current users by designing a user survey, data analysis and develop of personas. Followed by a redesign of the UI following using Neilsen’s Usability Heuristic and Norman’s Principles of Interactive Design in industry standard software Figma. Lastly a Heuristic Evaluation of new UI design, using Neilsen’s Usability Heuristic and User testing of redesign UI and have a group of users complete a Thinking Aloud Test of the new UI. Preliminary results of the Heuristic Evaluation of current UI arise issue with Consistency and Standards, Visibility of System Status, Match System and Real World and Recognition Rather than Recall. These issues were addressed in the proposed redesign by applying Neilsen’s Usability Heuristic and Norman’s Principles of Interactive Design. Next steps include formalized list of lessons learned and design implications for future publications.

97 MATHEMATICS AND COMPUTING↗

INSPIRED: Inelastic neutron scattering prediction for instantaneous results and experimental design

Inelastic neutron scattering (INS) has unique advantages in probing how atoms vibrate and how the vibrations propagate and interact. Such dynamic information is crucial in understanding various material properties, from heat capacity, thermal conductivity, phase transitions, and chemical reactions to more exotic quantum behavior. The analysis and interpretation of the INS spectra often start from a model structure of the sample, followed by a series of calculations to obtain the simulated spectra to compare with experiments. The conventional way to perform such calculations usually requires significant time, computing resources, and specialized expertise. Here, we present a new program named INSPIRED (Inelastic Neutron Scattering Prediction for Instantaneous Results and Experimental Design), which enables users to perform rapid INS simulations in several different ways on their personal computers in just a few clicks, with the crystal structure as the only input file. Specifically, the users can choose a pre-trained symmetry-aware neural network (coupled with an autoencoder) to predict the phonon density of states (DOS), 1D S(E) and 2D S(|Q|,E) spectra for any given structure. One can also choose an existing density functional theory (DFT) calculation from a database (containing over 12,000 crystals), and quickly obtain the simulated INS spectra for single crystals and powders. It is also possible to use pre-trained universal machine learning force fields to relax a given crystal structure, calculate the phonon dispersion and DOS, and, subsequently, the INS spectra. All these functions are implemented with a PyQt graphic user interface. Finally, we expect these new tools will benefit broad user communities and significantly improve the efficiency of experiment design, execution, and data analysis for INS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

L-PBF High-Throughput Data Pipeline Approach for Multi-modal Integration

Abstract Metal-based additive manufacturing requires active monitoring solutions for assessing part quality. Multiple sensors and data streams, however, generate large heterogeneous data sets that are impractical for manual assessment and characterization. In this work, an automated pipeline is developed that enables feature extraction from high-speed camera video and multi-modal data analysis. The framework removes the need for manual assessment through the utilization of deep learning techniques and training models in a weakly supervised paradigm. We demonstrate this pipeline’s capability over 700,000 high-speed camera frames. The pipeline successfully extracts melt pool and spatter geometries and links them to corresponding pyrometry, radiography, and processparameter information. 715 individual prints are examined to reveal melt pool areas that exceeds 0.07 mm 2 and pyrometry signal over a threshold (375 pyrometry units) were more likely to have defects. These automated processes enable massive throughput of characterization techniques.

36 MATERIALS SCIENCE↗

Optimizing Hydronic Heating for Comfort and Performance in Multifamily Housing

Inefficient control settings in multifamily boilers often lead to substantial energy and cost penalties. To address this, a Fault Detection and Diagnostic (FDD) tool was developed to automate data analysis and identify operational faults such as suboptimal outdoor temperature sensor placement, misconfigured outdoor air reset (OAR) curves, excess boiler cycling, and domestic hot water (DHW) setpoint errors. By comparing pre- and post-implementation periods and applying engineering models, the tool quantifies energy savings and reduces manual analysis time by over 90%. Testing on over 100 monitored sites and a targeted subset of 12 buildings showed an average 11% energy savings from remote optimization; further validation across 19 OAR curve changes confirmed the tool’s accuracy, predicting actual savings within ±5% for most cases. Simple payback can be under three years for many multifamily buildings, though rising hardware, labor, and fuel costs create uncertainties, and decarbonization goals increasingly shift focus to electrification. The FDD tool remains invaluable for optimizing existing boilers, enhancing future electrification measures, and adapting to new technologies by refining building load estimates. In doing so, it supports both near-term efficiency and long-term transitions to low-carbon alternatives, ensuring buildings achieve substantial cost and energy benefits throughout their system lifecycles.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Unveiling the nanoscale architectures and dynamics of protein assembly with in situ atomic force microscopy

Proteins play a vital role in different biological processes by forming complexes through precise folding with exclusive inter- and intra-molecular interactions. Understanding the structural and regulatory mechanisms underlying protein complex formation provides insights into biophysical processes. Furthermore, the principle of protein assembly gives guidelines for new biomimetic materials with potential applications in medicine, energy, and nanotechnology. Atomic force microscopy (AFM) is a powerful tool for investigating protein assembly and interactions across spatial scales (single molecules to cells) and temporal scales (milliseconds to days). It has significantly contributed to understanding nanoscale architectures, inter- and intra-molecular interactions, and regulatory elements that determine protein structures, assemblies, and functions. This review describes recent advancements in elucidating protein assemblies with in situ AFM. We discuss the structures, diffusions, interactions, and assembly dynamics of proteins captured by conventional and high-speed AFM in near-native environments and recent AFM developments in the multimodal high-resolution imaging, bimodal imaging, live cell imaging, and machine-learning-enhanced data analysis. These approaches show the significance of broadening the horizons of AFM and enable unprecedented explorations of protein assembly for biomaterial design and biomedical research.

36 MATERIALS SCIENCE↗

Open‐Source Anaerobic Digestion Modeling Platform, Anaerobic Digestion Model No. 1 Fast (ADM1F)

An open‐source modeling platform, called Anaerobic Digestion Model No. 1 Fast (ADM1F), is introduced to achieve fast and numerically stable simulations of anaerobic digestion processes. ADM1F is compatible with an iPython interface to facilitate model configuration, simulation, data analysis, and visualization. Faster simulations and more stable results are accomplished by implementing an advanced open‐source library of numerical methods called Portable Extensive Toolkit for Scientific Computation (PETSc) to solve the ADM1 system of equations. Leveraging PETSc, ADM1F can consistently complete a steady‐state simulation under 0.2 s, over 99% faster than a benchmark ADM1 model implemented with MATLAB while achieving agreement of model outputs within 1% of those obtained with the benchmark model. For dynamic simulations, however, ADM1F has a computational speed advantage only when the influent characteristics update more frequently than every 4 h. The ability of ADM1F to be useful as a tool to study anaerobic digestion systems is demonstrated through two example implementations of ADM1F: (1) a two‐phase co‐digestion scenario evaluating the impact of the organic loading rate and the substrate composition on reactor performance and stability, and (2) a conventional digester scenario assessing the effectiveness of recovery strategies after disruptions that led to instability. These examples demonstrate how the high simulation speed and the convenience of the iPython interface allow ADM1F to complete complex analyses within minutes, much faster than computational strategies currently reported in the literature.

anaerobic co-digestion↗

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan↗

The microbiologist's guide to metaproteomics

Metaproteomics is an emerging approach for studying microbiomes, offering the ability to characterize proteins that underpin microbial functionality within diverse ecosystems. As the primary catalytic and structural components of microbiomes, proteins provide unique insights into the active processes and ecological roles of microbial communities. By integrating metaproteomics with other omics disciplines, researchers can gain a comprehensive understanding of microbial ecology, interactions, and functional dynamics. This review, developed by the Metaproteomics Initiative (www.metaproteomics.org), serves as a practical guide for both microbiome and proteomics researchers, presenting key principles, state-of-the-art methodologies, and analytical workflows essential to metaproteomics. Topics covered include experimental design, sample preparation, mass spectrometry techniques, data analysis strategies, and statistical approaches.

bioinformatics↗

Enabling Seamless Transitions from Experimental to Production HPC for Interactive Workflows

The evolving landscape of scientific computing requires seamless transitions from experimental to production HPC environments for interactive workflows. This paper presents a structured transition pathway developed at OLCF that bridges the gap between development testbeds and production systems. We address both technological and policy challenges, introducing frameworks for data streaming architectures, secure service interfaces, and adaptive resource scheduling for time-sensitive workloads and improved HPC interactivity. Our approach transforms traditional batch-oriented HPC into a more dynamic ecosystem capable of supporting modern scientific workflows that require near real-time data analysis, experimental steering, and cross-facility integration.

Etz, Brian [ORNL] (ORCID:0000000208554863)↗

Using pile-up collisions as an abundant source of low-energy hadronic physics processes in ATLAS and an extraction of the jet energy resolution

During the 2015–2018 data-taking period, the Large Hadron Collider delivered proton-proton bunch crossings at a centre-of-mass energy of 13 TeV to the ATLAS experiment at a rate of roughly 30 MHz, where each bunch crossing contained an average of 34 independent inelastic proton-proton collisions. The ATLAS trigger system selected roughly 1 kHz of these bunch crossings to be recorded to disk. Offline algorithms then identify one of the recorded collisions as the collision of interest for subsequent data analysis, and the remaining collisions are referred to as pile-up. Pile-up collisions represent a trigger-unbiased dataset, which is evaluated to have an integrated luminosity of 1.33 pb -1 in 2015–2018. This is small compared with the normal trigger-based ATLAS dataset, but when combined with vertex-by-vertex jet reconstruction it provides up to 50 times more dijet events than the conventional single-jet-trigger-based approach, and does so without adding any additional cost or requirements on the trigger system, readout, or storage. The pile-up dataset is validated through comparisons with a special trigger-unbiased dataset recorded by ATLAS, and its utility is demonstrated by means of a measurement of the jet energy resolution in dijet events, where the statistical uncertainty is significantly reduced for jet transverse momenta below 65 GeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Laboratory-Based Micro-X-ray Computed Tomography of Energy Materials at Idaho National Laboratory

Abstract The Idaho National Laboratory (INL) has implemented laboratory-based micro-X-ray computed tomography in a laboratory equipped for the examination of highly radioactive samples. This capability provides nondestructive three-dimensional volumetric information on samples to inform subsequent traditional destructive examinations as well as real-world inputs for high-fidelity scientific modeling. Samples can be imaged with spatial resolutions ranging from several hundred nm/voxel up to ~ 100 µm/voxel. The best usable spatial resolution achieved to date is 384 nm/voxel with this instrument, while the highest radiological dose rate of a sample imaged is ~ 60 R/h β/γ on contact. Advanced data analysis, including custom tomographic reconstruction and segmentation methods, have also been developed to support this capability. In addition to traditional digital X-ray radiography and tomography, this instrument is also able to visualize in situ tensile and compression testing as well as perform diffraction contrast tomography. This work describes the X-ray computed tomography post-irradiation examination capabilities at INL, as well as detailing a variety of applications this instrument has examined.

36 MATERIALS SCIENCE↗

Damage progression and failure of SiC/SiC composite tubes under hard-contact radial expansion

The response of silicon carbide (SiC) fiber-reinforced SiC matrix (SiC/SiC) composite cladding to mechanical interaction with fissile fuel is a knowledge gap that must be overcome to design and assess SiC-based cladding systems for advanced nuclear applications. This study developed the relevant mechanical testing capability and identified the failure behavior and the critical microstructural features and processing defects. Sections of SiC composite tube were subjected to a modified expansion-due-to-compression (EDC) test in an X-ray computed tomography microscope: a polyurethane plug pressed surrogate Al 2 O 3 into the inner walls of the SiC/SiC composite tubes to achieve hard contact. A pure EDC test with just a polyurethane plug was also performed as a reference. Through the use of displacement fields, digital volume correlation revealed inhomogeneous deformation fields in the tubes, even for pure EDC, which was related to the inherent defects in the structure. Deep learning–aided segmentation and systematic data analysis revealed that the presence of inhomogeneous deformation applied by the hard contact was exaggerated by the presence of inner surface imperfections left behind from the matrix densification process. In conclusion, the findings provide insights into the applications, highlighting the necessity for improvements in inner surface roughness and the incorporation of localized contacts in pellet–cladding mechanical interaction computational models.

Composites↗

Quasi-optical beam tracing module development for millimeter-wave high-wavenumber collective scattering on the NSTX-U and EAST tokamaks

A Python3-based beam tracing code utilizing Quasi-Optics has been developed to track both incident and receiving beams in high-k collective millimeter wave scattering systems within magnetic fusion plasmas. In contrast to existing ray tracing codes that solely consider refraction, this beam tracing code incorporates diffraction phenomena, providing a more comprehensive calculation. Here, this enhanced capability allows for a more accurate calculation of the scattering volume and spatial resolution in high-k collective scattering systems, crucial for evaluating system performance and facilitating data analysis. Unlike Geometrical Optics, Quasi-Optics employs the complex eikonal method, representing a Gaussian beam as a collection of coupled rays to accurately preserve diffraction characteristics. The developed code is intended for application in NSTX-Upgrade and EAST high-k beam tracing analyses, targeting frequencies of 693 GHz and 270 GHz, respectively. The high-k system's primary objective is the observation of electron-scale instabilities. Employing a symplectic integrator, the code ensures numerical accuracy, assessed through the conservation of the Hamiltonian. With its precision and efficiency, the code facilitates rapid inter-shot analyses.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗