Search NASA⌕ Search

SEARCH · Search NASA

Results for “High dimensional data,”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

A Unified Analytical Method Greenness Score ( uAMGS ) Quantifies How Microscopic Imaging Is Greener Than Conventional Liquid Chromatography

Green chemistry is a set of principles for assessing, developing, and implementing methods that are safer, more efficient, and less detrimental to the environment. The analytical method greenness score (AMGS) is one of many metrics that attempt to evaluate traditional liquid chromatography (LC) based on the energy consumption of the instrument and the safety, health risks, and environmental impact of the solvents employed. Unfortunately, in practice, the AMGS is primarily focused on traditional separation methods in the pharmaceutical industry and is not amenable to cutting-edge separation science, including miniaturization. To broaden this scope, the unified Analytical Method Greenness Score (uAMGS) is presented here, which clarifies and expands on the underlying mathematics and incorporates both dimensional and uncertainty analysis, enabling its application to a broader range of analytical techniques. The uAMGS is used to compare the greenness of two distinct methods: single-molecule microscopy (SMM) and high-performance liquid chromatography (HPLC), which were used to collect equivalent data. uAMGS determines that SMM is significantly greener than HPLC due primarily to decreased solvent consumption. Overall, the uAMGS should allow chemists ranging from undergraduates to industrial PhDs to assess the greenness of a wide range of separations.

chemical separations↗

Sparks in the dark

This study presents a novel method for the definition of signal regions in searches for new physics at collider experiments. By leveraging multi-dimensional histograms with precise arithmetic and utilizing the SparkDensityTree library, it is possible to identify high-density regions within the available phase space, potentially improving sensitivity to very small signals. Inspired by a search for dark mesons at the ATLAS experiment, CMS open data is used for this proof-of-concept intentionally targeting an already excluded signal. Signal regions are defined based on density estimates of signal and background. These preliminary regions align well with the physical properties of the signal while effectively rejecting background events.

Gudnadottir, Olga Sunneborn [Uppsala Univ. (Sweden↗

A resolution independent neural operator

The Deep operator network (DeepONet) is a powerful yet simple neural operator architecture that utilizes two deep neural networks to learn mappings between infinite-dimensional function spaces. This architecture is highly flexible, allowing the evaluation of the solution field at any location within the desired domain. However, it imposes a strict constraint on the input space, requiring all input functions to be discretized at the same locations; this limits its practical applications. Here, in this work, we introduce a general framework for operator learning from input–output data with arbitrary number and locations of sensors. This begins by introducing a resolution-independent DeepONet (RI-DeepONet), enabling it to handle input functions that are arbitrarily, but sufficiently finely, discretized. To this end, we propose two dictionary learning algorithms to adaptively learn a set of appropriate continuous basis functions, parameterized as implicit neural representations (INRs), from correlated signals defined on arbitrary point cloud data. These basis functions are then used to project arbitrary input function data as a point cloud onto an embedding space (i.e., a vector space of finite dimensions) with dimensionality equal to the dictionary size, which can be directly used by DeepONet without any architectural changes. In particular, we utilize sinusoidal representation networks (SIRENs) as trainable INR basis functions. The introduced dictionary learning algorithms are then used in a similar way to learn an appropriate dictionary of basis functions for the output function data, which defines a new neural operator architecture referred to as the R esolution I ndependent N eural O perator (RINO). In the RINO, the operator learning task simplifies to learning a mapping from the coefficients of input basis functions to the coefficients of output basis functions. We demonstrate the robustness and applicability of RINO in handling arbitrarily (but sufficiently richly) sampled input and output functions during both training and inference through several numerical examples.

Deep operator network (DeepONet)↗

Adaptively coupled phase retrieval in multi-peak Bragg coherent diffraction imaging

Recent advances in Bragg coherent diffraction imaging (BCDI) experimental techniques permit routine measurement of multiple Bragg peaks from a single crystalline grain. The resulting images contain the full lattice distortion vector field which can be differentiated to provide lattice strain and rotation. With the advent of fourth-generation synchrotron light sources, such multi-peak datasets are produced at high rates, facilitating the need for rapid phase retrieval of the multiple peaks and subsequent image analysis. Here we describe and demonstrate a new implementation of a coupled phase retrieval technique for multi-peak BCDI which simultaneously treats each Bragg peak of the dataset and produces a three-dimensional image of the crystal's morphology and lattice distortion field. In addition, this method uses the redundant information contained in the various Bragg diffraction patterns to detect and suppress spurious signal appearing on the detector in a subset of the measurements. Compared with manual data editing, adaptive coupling produces a more consistent phase profile in reciprocal space and sharper surfaces in direct space, with no significant difference in computational cost. These improvements reduce the need for manual preprocessing and enable robust high-throughput analysis of multi-peak BCDI data, supporting near-real-time strain microscopy at modern synchrotron facilities.

36 MATERIALS SCIENCE↗

Synthesizing realistic sand assemblies with denoising diffusion in latent space

Abstract The shapes and morphological features of grains in sand assemblies have far‐reaching implications in many engineering applications, such as geotechnical engineering, computer animations, petroleum engineering, and concentrated solar power. Yet, our understanding of the influence of grain geometries on macroscopic response is often only qualitative, due to the limited availability of high‐quality 3D grain geometry data. In this paper, we introduce a denoising diffusion algorithm that uses a set of point clouds collected from the surface of individual sand grains to generate grains in the latent space. By employing a point cloud autoencoder, the three‐dimensional point cloud structures of sand grains are first encoded into a lower‐dimensional latent space. A generative denoising diffusion probabilistic model is trained to produce synthetic sand that maximizes the log‐likelihood of the generated samples belonging to the original data distribution measured by a Kullback‐Leibler divergence. Numerical experiments suggest that the proposed method is capable of generating realistic grains with morphology, shapes and sizes consistent with the training data inferred from an F50 sand database. We then use a rigid contact dynamic simulator to pour the synthetic sand in a confined volume to form granular assemblies in a static equilibrium state with targeted distribution properties. To ensure third‐party validation, 50,000 synthetic sand grains and the 1542 real synchrotron microcomputed tomography (SMT) scans of the F50 sand, as well as the granular assemblies composed of synthetic sand grains are made available in an open‐source repository.

Vlassis, Nikolaos N.↗

Higher Dimensionality in the Mg–Co–B System: Synthesis and Structure of Incommensurate Composite Mg 1+ε Co 4 B 4

Guided by high-temperature in situ X-ray diffraction, the discovery and synthesis of Mg 1+ε Co 4 B 4 (ε ≈ 0.272) using a MgH 2 hydride precursor is reported, along with a detailed crystal structure description and measurement of magnetic properties. The mismatch in lattice periodicities between Mg and Co–B substructures places Mg 1+ε Co 4 B 4 in the family of incommensurate composite crystals and prompted structural refinement in a (3 + 1)-dimensional model. The structure of Mg 1+ε Co 4 B 4 (P4 2 /ncm(00γ)s00s, a = 6.75847(7) Å, c = 3.94007(8) Å, q = (0, 0, 1.2721(3))) was refined from neutron powder diffraction and high-resolution powder X-ray diffraction data and confirmed by scanning transmission electron microscopy and electron diffraction. Mg 1+ε Co 4 B 4 is isostructural to Nd 1+ε Fe 4 B 4 and several related ternary borides with 0.07 ≤ ε ≤ 0.17, with Mg occupying the rare-earth site. Satellite reflections in the electron diffraction patterns hinted at positional modulation of the transition metal–boron substructure by Mg atoms, but this could not be refined from the neutron or X-ray diffraction data. Low-temperature magnetic measurements show no indications of long-range magnetic ordering or superconductivity down to 5 K. DFT calculations confirmed the absence of a magnetically ordered ground state and the stability of a 5:4 supercell (ε = 0.25) relative to the fully commensurate structure. Neutron diffraction and synthesis from elemental Mg demonstrated that Mg 1+ε Co 4 B 4 is not a hydrogen-stabilized phase. Mg 1+ε Co 4 B 4 represents the second compound reported in the Mg–Co–B system and the first superspace symmetry model of a Nd 1+ε Fe 4 B 4 -type incommensurate composite compound refined from powder diffraction data.

chemical structure↗

Demonstration and performance of an online data selection algorithm for liquid argon time projection chambers using MicroBooNE

The MicroBooNE detector is a liquid argon time projection chamber (LArTPC) that produces three-dimensional images of particle interactions using ionization charge collected by anode wire plane arrays and scintillation light collected by a light detection system. In addition to testing long-standing experimental neutrino anomalies and performing measurements of neutrino interactions with argon nuclei using the Fermilab Booster Neutrino Beam, MicroBooNE aims to develop methodologies for rare beyond the Standard Model and off-beam physics searches. Looking ahead to the upcoming Deep Underground Neutrino Experiment (DUNE), with MicroBooNE serving as a valuable testbed, achieving high sensitivity and livetime for off-beam physics while satisfying data processing and storage constraints will require data-driven, intelligent, and online or real-time data selection techniques. These techniques are essential for reducing data rates and preserving rare signals with high accuracy. In this paper, we describe a fast data selection algorithm suitable for online execution to identify electrons from stopping cosmic ray muons in the MicroBooNE detector utilizing ionization charge information, and present its performance. This represents the first demonstration of online data selection in a LArTPC using real data and charge information exclusively and provides an important proof-of-principle for applying such techniques to other LArTPC experiments such as the Short-Baseline Near Detector and DUNE.

Abratenko, P. [Tufts U. (main)]↗

Harnessing the power of gradient-based simulations for multi-objective optimization in particle accelerators

Abstract Particle accelerator operation requires simultaneous optimization of multiple objectives. Multi-objective optimization (MOO) is particularly challenging due to trade-offs between the objectives. Evolutionary algorithms, such as genetic algorithms (GAs), have been leveraged for many optimization problems, however, they do not apply to complex control problems by design. This paper demonstrates the power of differentiability for solving MOO problems in particle accelerators using a deep differentiable reinforcement learning (DDRL) algorithm. We compare the DDRL algorithm with model-free reinforcement learning (MFRL), GA, and Bayesian optimization (BO) for simultaneous optimization of heat load and trip rates in the continuous electron beam accelerator facility. The underlying problem enforces strict constraints on both individual states and actions as well as cumulative (global) constraints on energy requirements of the beam. Using historical accelerator data, we develop a physics-based surrogate model which is differentiable and allows for back-propagation of gradients. The results are evaluated in the form of a Pareto-front with two objectives. We show that the DDRL outperforms MFRL, BO, and GA on high dimensional problems.

43 PARTICLE ACCELERATORS↗

AutoTandemML: Active Learning Enhanced Tandem Neural Networks for Inverse Design Problems

Inverse design in science and engineering involves determining optimal design parameters that achieve desired performance outcomes, a process often hindered by the complexity and high dimensionality of design spaces, leading to significant computational costs. To tackle this challenge, we propose a novel hybrid approach that combines active learning with Tandem Neural Networks to enhance the efficiency and effectiveness of solving inverse design problems. Active learning allows to selectively sample the most informative data points, reducing the required dataset size without compromising accuracy. We investigate this approach using three benchmark problems: airfoil inverse design, photonic surface inverse design, and scalar boundary condition reconstruction in diffusion partial differential equations. We demonstrate that integrating active learning with Tandem Neural Networks outperforms standard approaches across the benchmark suite, achieving better accuracy with fewer training samples.

97 MATHEMATICS AND COMPUTING↗

Bayesian inference of anisotropic 2D small-angle scattering from sparse measurement

Here, we present a Bayesian inference framework for reconstructing anisotropic two-dimensional small-angle scattering (2D SAS) patterns from sparse, noisy, or partially missing data. The method combines a symmetry-aware angular basis with radial Gaussian process priors to enable accurate, training-free interpolation and denoising. Computational benchmarks demonstrate reliable recovery of both isotropic and high-order anisotropic features under severe data reduction. Experimental validations on stretched polymers, sheared wormlike micelles, and carbon fibers show improved fidelity and resolution compared to raw measurements, achieving comparable accuracy with up to 50-fold fewer detected neutrons. This approach enables quantitative structural analysis under low-flux, time-limited, or single-shot conditions, extending the applicability of 2D SAS techniques to compact neutron sources and mechanically driven soft matter systems undergoing transient structural changes.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

A science-driven approach to optimize the design for a biological small-angle neutron scattering instrument

Biological small-angle neutron scattering (SANS) instruments facilitate critical analysis of the structure and dynamics of complex biological systems. However, with the growth of experimental demands and the advances in optical systems design, a new neutron optical concept is necessary to overcome the limitations of current instruments. This work presents an approach to include experimental objectives ( i.e. the science to be supported by a specific neutron scattering instrument) in the optimization of the neutron optical concept. The approach for a proposed SANS instrument at the Second Target Station of the Spallation Neutron Source at Oak Ridge National Laboratory, USA, is presented here. Further, the instrument is simulated with the McStas software package. The optimization process is driven by an evolutionary algorithm using McStas output data, which are processed to calculate an objective function designed to quantify the expected performance of the simulated neutron optical configuration for the intended purpose. Each McStas simulation covers the complete instrument, from source to detector, including realistic sample scattering functions. This approach effectively navigates a high-dimensional parameter space that is otherwise intractable; it allows the design of next-generation SANS instruments to address specific scientific cases and has the potential to increase instrument performance compared with traditional design approaches.

47 OTHER INSTRUMENTATION↗

Orientation microscopy–assisted grain boundary analysis for protonic ceramic cell electrolytes

Abstract Grain boundaries in protonic ceramic cell (PCC) electrolytes hinder proton transport, reducing interfacial conductivity. In multicomponent PCC electrolytes, the inclusion of sintering aids further accentuates the complexity of grain boundaries. In this study, we synthesize nanocrystalline BaCe 0.4 Zr 0.4 Y 0.1 Yb 0.1 O 3− δ thin films via pulsed laser deposition and analyze their grain boundary character distributions using orientation data collected by precession electron diffraction technique. The results reveal an anisotropic distribution of grain boundary characters, with notably high populations of 180°‐tilt and twist grain boundaries. These findings provide critical insights into identifying the predominant grain boundaries in this PCC electrolyte material, assessing the vast five‐dimensional grain boundary space.

Patel, Sooraj [School of Aerospace and Mechanical ↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Boosted decision tree reweighting of simulated neutrino interactions for O ( 1 ) GeV neutrino cross-section measurements

This paper illustrates a generic method for multidimensional reweighting of O ( 1 ) GeV neutrino interaction Monte Carlo samples. The reweighting is based on a boosted decision tree algorithm trained on high-dimensional space in detector final-state observables. This enables one generator’s events to be reweighted so that its reconstructed particle content and kinematics distributions, as well as detector efficiency, match those of a target model. The approach establishes an efficient way to reuse legacy Monte Carlo data, avoiding regeneration. As an example, we test its use in a measurement of transverse kinematic imbalance of the μ - and proton in charged-current quasielastic like ν μ events from the MINERvA experiment.

Lin, Z. [Rochester U.] (ORCID:0009000188903698)↗

Detector pixel calibration of time-of-flight neutron diffractometers accelerated by machine learning

Modern time-of-flight neutron diffractometers at spallation neutron source are equipped with two dimensional detectors with fine pixelations. The flight path of neutrons from the moderator to the sample and to the detector needs to be precisely calibrated at detector pixel level using standard powders so the diffraction data from all the detector pixels can be correctly time-focused to produce high resolution diffraction peaks. The number of pixels can reach to millions which makes a single-pixel calibration process time-consuming, or even impossible, with conventional fitting routine. Here we presented a machine learning aided calibration process via a “training and predict” process by training machine learning models with the relations between the individual pixel time-of-flight diffraction pattern and fitted diffraction constant. The training models take a portion of the available pixels to predict the diffraction constants precisely and rapidly for massive pixel diffraction patterns.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Boosting Barlow Twins Reduced Order Modeling for Machine Learning‐Based Surrogate Models in Multiphase Flow Problems

Abstract We present an innovative approach called boosting Barlow Twins reduced order modeling (BBT‐ROM) to enhance the reliability of machine learning surrogate models for multiphase flow problems. BBT‐ROM builds upon Barlow Twins reduced order modeling that leverages self‐supervised learning to effectively handle linear and nonlinear manifolds by constructing well‐structured latent spaces of input parameters and output quantities. To address the challenge of high contrast data in multiphase flow problems due to injection wells and faults, we employ a boosting algorithm within BBT‐ROM. This algorithm sequentially trains a set of weak models (i.e., inaccurate models), improving prediction accuracy through ensemble learning. To evaluate the performance of BBT‐ROM, we conduct three three‐dimensional multiphase flow problems, including waterflooding and geologic carbon storage (GCS), with varying numbers of input parameter cases and model domain features. The results demonstrate that BBT‐ROM excels at predicting non‐wetting phase saturation (e.g., oil or saturation) and fluid pressure, with average relative errors ranging from 0.5% to 3%. Importantly, BBT‐ROM showcases robustness when faced with limited input parameter space during GCS testing.

58 GEOSCIENCES↗

A Typing Discipline for High-Assurance Control Systems

This poster describes a typing discipline for high-assurance industrial systems based on three novel type systems. The first type system, information flow control (IFC), controls the flow of data through the system. The second system, dependent session types, restricts messages exchanged during the execution of a communication protocol to avoid dangerous states. The third system uses dimensional analysis to avoid subtle bugs that adversaries can exploit to cause the system to enter a dangerous state. This poster describes how a combination of these approaches can prevent sophisticated cyber attacks, such as the infamous Stuxnet incident, from occurring. In addition, we provide experimental evidence to support the claim that these approaches can be applied in control systems that are resource-constrained.

42 - ENGINEERING↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗