Search NASA⌕ Search

SEARCH · Search NASA

Results for “High dimensional modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Explainable and Differentiable Reinforcement Learning for Multi-objective Optimization in Particle Accelerators

Operating particle accelerators involves optimizing multiple goals simultaneously, which can be challenging due to trade-offs among objectives. While evolutionary algorithms like the genetic algorithm (GA) have been used for various Multi-Objective Optimization (MOO) tasks, they are not inherently suited for complex control problems. This talk highlights two variations of Reinforcement Learning (RL) for concurrently optimizing heat load and trip rates at the Continuous Electron Beam Accelerator Facility (CEBAF). The problem involves strict constraints on individual states, actions, and overall energy requirements of the beam. First, this talk highlights how differentiability can be harnessed through a Deep Differentiable Reinforcement Learning (DDRL) approach to address MOO issues within particle accelerators. We examine the DDRL method alongside Model Free Reinforcement Learning (MFRL), GA, and Bayesian Optimization (BO). The performance of these methods is assessed by generating a Pareto-front for two objectives. Our findings indicate that DDRL excels in handling high-dimensional problems more effectively than MFRL, BO, and GA. Next, we will show integration of explainable physics-based constraints into RL algorithms to enhance trans- parency and trust in decision-making processes by enabling users to verify that agents adhere to established physical principles. This surrogate function can be modeled using neural networks or sparse dictionary mod- els. By examining the mathematical form of the learned constraint function, we are able to confirm the agent has learned to use the established physics of each environment provided but the surrogate model. In addi- tion, we find that the introduction of a mathematical functional dictionary based surrogate model enables our reinforcement learning algorithms to reliably converge for difficult high-dimensional accelerator controls environments.

Rajput, Kishansingh [Thomas Jefferson National Acc↗

Enhancing dimensionality prediction in hybrid metal halides via feature engineering and class-imbalance mitigation

We present a machine learning (ML) framework for predicting the structural dimensionality of hybrid metal halides (HMHs), including organic-inorganic perovskites, using a combination of chemically-informed feature engineering and advanced class-imbalance handling techniques. This study is motivated by the small and highly imbalanced nature of experimentally available HMH datasets, which limits the applicability and reliability of conventional ML approaches. The dataset, consisting of 494 HMH structures, is highly imbalanced across dimensionality classes (0D, 1D, 2D, 3D), posing significant challenges to predictive modeling. To mitigate this limitation, the dataset was augmented to 1336 samples using the synthetic minority oversampling technique, enabling improved learning of underrepresented dimensionality classes while preserving chemically meaningful feature relationships. We developed interaction-based descriptors designed to capture coupled steric and polarity effects relevant to dimensionality prediction, which are not readily captured by standard single-parameter or composition-only descriptors. These descriptors are integrated into a multi-stage workflow combining feature selection, ensemble stacking, and performance optimization. Our approach significantly improves F1-scores for underrepresented classes, achieving robust cross-validation performance across all dimensionalities. This work demonstrates a generalizable strategy for extracting reliable and interpretable structure–dimensionality relationships from limited experimental data, enabling pre-synthesis screening of organic cations and providing a practical blueprint for small-data ML in hybrid materials systems.

36 MATERIALS SCIENCE↗

SmileyLlama: modifying large language models for directed chemical space exploration

Here we show that large language models (LLMs) can be transformed via supervised fine-tuning of engineered prompts into SmileyLlama for exploring the chemical space of drug molecules. We benchmark SmileyLlama against pretrained LLMs and chemical language models trained from scratch for generating valid and novel drug-like molecules, and use direct preference optimization to both improve SmileyLlama’s adherence to a prompt and as part of the iMiner reinforcement learning framework to predict molecules with optimized three-dimensional conformations and high binding affinity to drug targets. By training an LLM to speak directly as a chemical language model, while retaining most of its natural language capabilities, we show that SmileyLlama can reliably generate molecules with user-specified properties rather than acting only as a chatbot with knowledge of chemistry or as a virtual assistant. While SmileyLlama is geared toward drug discovery, the supervised fine-tuning/direct preference optimization/LLM framework can be extended to other chemical, biological and materials applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Simulating Pyrocumulonimbus Clouds Using a Multiscale Wildfire Simulation Framework

Pyrocumulonimbus (pyroCb) clouds, driven by extreme fires under favorable meteorological conditions, can inject smoke into the stratosphere at magnitudes comparable to those of moderate volcanic eruptions, potentially altering the global radiative balance and atmospheric composition. However, simulating pyroCb is particularly challenging in Earth system models. Using the Energy Exascale Earth System Model (E3SM), we developed a novel global multiscale framework to model pyroCb events in California, which includes a high‐resolution fire radiative power time series, a one‐dimensional plume‐rise parameterization, a fire‐induced vertical water vapor transport scheme, and a surface wildfire sensible heat flux representation. Our simulation successfully reproduces many pyroCb features, including cloud height, spatiotemporal evolution, and convective intensity in comparison with satellite and ground‐based observations. Sensitivity experiments show that realistic pyroCb simulation depends on vertical water vapor transport. These advances provide a basis for future exploration of pyroCb impacts at regional and global scales within climate models.

E3SM↗

Architecture-Aware Models of AI Engines for High-Performance Matrix Matrix Multiplication

The AI Engine (AIE) architecture, available in systems from mobile SoCs to server-class FPGAs, aims to efficiently execute AI/ML tasks through a two-dimensional array of compute tiles. Previous work on AIEs has explored different approaches to mapping computation across spatial arrays, but the compute kernel running on each tile has not been the focus. Additionally, the AIE-ML architecture introduces memory tiles and omits programmable logic, requiring new approaches to staging and moving data throughout the array. In this work we update analytical models developed for CPUs to produce the design of high performance kernels while introducing new model considerations such as memory structure, throughput, and latency as required by the AIE hardware. We evaluate our models by developing AIE-ML kernels for matrix multiplication in low-precision data types showing performance up to 95% of compute peak for the kernel when data resides in local memory and above 90% of compute peak when data resides in main memory.

Binder, Elliott D. [Carnegie Mellon University, Pi↗

Accelerating uncertainty quantification in incremental dynamic analysis using dimension reduction-based surrogate modeling

We propose a surrogate modeling framework based on dimension reduction to facilitate the quantification of seismic risk of structural systems in performance-based earthquake engineering. The framework adopts incremental dynamic analysis (IDA) for addressing hazard variability, and promotes significant computational efficiency improvement for propagating epistemic uncertainties associated with the structural models. It utilizes both linear and nonlinear dimension reduction approaches, equipped with inverse mappings, to learn a functional between the input parameter space (e.g., the epistemic uncertainties of the structure) to the high-dimensional output space created through the IDA implementation across different ground motions and seismic intensity levels. Polynomial chaos expansion is adopted as the surrogate model to learn this functional in the reduced space. A nine-story steel moment-resisting frame with uncertain structural properties is used as a testbed. Furthermore, we select the seismic fragility curves as a measure of the structure’s seismic performance, since it provides an estimate of the probability of entering specified damage states for given levels of ground shaking.

42 ENGINEERING↗

High-dimensional maximum-entropy phase space tomography

Reconstructing 4D or 6D phase space distributions from 1D or 2D measurements is a challenging inverse problem encountered in particle accelerators. Entropy maximization is an established method to incorporate prior information in the reconstruction, but it is typically infeasible in high-dimensional spaces. In this paper, I review two recent approaches to high-dimensional entropy maximization. The first approach utilizes differentiable simulations and a class of generative models known as normalizing flows, whereas the second approach employs the method of Lagrange multipliers and Markov Chain Monte Carlo (MCMC) sampling. My aim is to provide a short explanation of each method using a common notation. I conclude by mentioning several unsolved problems in phase space tomography.

Hoover, Austin [ORNL] (ORCID:0000000153136962)↗

On the validity and limitations of 1D model for heat and mass transfer performance evaluation in a multilayer binder-free desiccant dehumidifier: isothermal dehumidification with internal cooling

Efficient humidity control is essential for maintaining indoor thermal comfort, yet conventional vapor-compression-based dehumidifiers are energy-intensive. Employing separate sensible and latent cooling through desiccant-coated heat exchangers (DCHEs) combined with evaporative coolers offers energy savings of up to 80 % compared to conventional systems. However, the dehumidification performance of DCHEs remains limited due to the use of polymer binders for coating desiccant materials onto heat exchange surfaces. In our previous study, we developed a multilayer fixed-bed binder-free desiccant dehumidifier (MFBDD) that demonstrated high dehumidification capacity and low pressure drop compared to rotary desiccant wheels. Nevertheless, its potential for further enhancement through internal cooling and the use of step-shaped adsorption isotherms has not been explored. In this study, a physics-based one-dimensional (1D) transient model is developed and validated to capture the coupled heat and mass transfer processes in the MFBDD and extended to simulate internal cooling using a high-capacity composite metal–organic framework, MIL-101/GO-6 (water uptake ≈1.6 g/g within 35–47 % RH). The model enables detailed analysis of local air and bed temperature dynamics and quantifies how internal cooling affects the dehumidification performance under a wide range of operating conditions. Results show that integrating internal cooling and using MIL-101/GO-6 enhance mass adsorbed, moisture removal capacity, and dehumidification effectiveness by 50 %–99 % compared with the M.S. Gel baseline. The study further reveals that achieving near-isothermal operation requires simultaneous enhancement of the convective heat transfer coefficient and heat exchange surface area. In conclusion, this work provides the first detailed physical insight into the interplay between internal cooling and step-shaped isotherms in a binder-free desiccant device and establishes a validated modeling framework for scaling up and system-level performance evaluation of next-generation energy-efficient dehumidification systems.

Heat and mass transfer↗

Streaming Compression of Scientific Data via Weak-SINDy

Here, in this paper, a streaming weak-SINDy algorithm is developed specifically for compressing streaming scientific data. The production of scientific data, either via simulation or experiments, is undergoing a stage of exponential growth, which makes data compression important and often necessary for storing and utilizing large scientific data sets. As opposed to classical “offline” compression algorithms that perform compression on a readily available data set, streaming compression algorithms compress data “online” while the data generated from simulation or experiments is still flowing through the system. This feature makes streaming compression algorithms well suited for scientific data compression, where storing the full data set offline is often infeasible. This work proposes a new streaming compression algorithm, streaming weak-SINDy, which takes advantage of the underlying data characteristics during compression. The streaming weak-SINDy algorithm constructs feature matrices and target vectors in the online stage via a streaming integration method in a memory efficient manner. The feature matrices and target vectors are then used in the offline stage to build a model through a regression process that aims to recover equations that govern the evolution of the data. For compressing high-dimensional streaming data, we adopt a streaming proper orthogonal decomposition (POD) process to reduce the data dimension and then use the streaming weak-SINDy algorithm to compress the temporal data of the POD expansion. We propose modifications to the streaming weak-SINDy algorithm to accommodate the dynamically updated POD basis. By combining the built model from the streaming weak-SINDy algorithm and a small amount of data samples, the full data flow could be reconstructed accurately at a low memory cost, as shown in the numerical tests.

97 MATHEMATICS AND COMPUTING↗

LDRD Abbreviated report: High-Order General-Discrete-Ordinates Method Enabling Efficient Deterministic Transport in Hydrodynamic Simulations

Deterministic transport simulations for national-security and energy applications often operate in high-dimensional phase-space, where accuracy and cost both become major challenges. A common numerical artifact in such problems is the “ray-effect,” which appears as unphysical streaks. Beyond misinterpretation, these artifacts can contaminate tightly coupled physics, such as fluid dynamics, radiation-hydrodynamics, and laser-plasma interactions, eroding the predictive capability of entire multiphysics workflows. Our objective was to make high-dimension studies practical on modern hardware while mitigating the ray-effect without relying on prohibitively expensive sampling approaches such as Monte Carlo methods. We developed the Generic Discretization Library (GenDiL), a Graphics Processing Unit (GPU)-first framework that uses high-order Discontinuous Galerkin (DG) methods and matrix-free algorithms to reduce memory usage and improve computational efficiency, critical for phase-space simulations. GenDiL supports phase-space adaptivity in both mesh size and polynomial order (hp-adaptivity) to place resolution only where it is needed. A central capability is Local Dimensional Refinement (LDR), which couples lower-dimension continuum models to higher-dimension kinetic models through stable and conservative interfaces, so that high-fidelity physics is applied only in regions where it is essential. Building on the GenDiL framework, we developed the General SN (GSN) family of algorithms as a true generalization of the polar SN approach (discrete ordinates, often denoted SN). Rather than tying discrete ordinates to a specific polar change of coordinates, GSN formulates transport on an arbitrary change of coordinates chosen to reduce ray-effect. We studied two complementary variants: an analytic variant, where the coordinate map is prescribed in advance by a closed-form function; and a data-driven variant, where a quantity of interest, such as the net flux, guides the coordinate system. GenDiL provides the library infrastructure for efficient GPU execution, but the GSN concept is algorithmic and independent of any one library. Across representative high-dimension tests, including non-symmetric solutions, both variants delivered strong ray-effect mitigation at practical cost, moving four- to six-dimensional analysis toward repeatable, routine studies.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Results from a synthetic model of the ITER XRCS-Core diagnostic based on high-fidelity x-ray ray tracing

A high-fidelity synthetic diagnostic has been developed for the ITER core x-ray crystal spectrometer diagnostic based on x-ray ray tracing. This synthetic diagnostic has been used to model expected performance of the diagnostic, to aid in diagnostic design, and to develop engineering tolerances. The synthetic model is based on x-ray ray tracing using the recently developed xicsrt ray tracing code and includes a fully three-dimensional representation of the diagnostic based on the computer aided design. The modeled components are: plasma geometry and emission profiles, highly oriented pyrolytic graphite pre-reflectors, spherically bent crystals, and pixelated x-ray detectors. Plasma emission profiles have been calculated for Xe 44+ , Xe 47+ , and Xe 51+ , based on an ITER operational scenario available through the Integrated Modelling & Analysis Suite database, and modeled within the ray tracing code as a volumetric x-ray source; the shape of the plasma source is determined by equilibrium geometry and an appropriate wavelength distribution to match the expected ion temperature profile. All individual components of the x-ray optical system have been modeled with high-fidelity producing a synthetic detector image that is expected to closely match what will be seen in the final as-built system. Particular care is taken to maintain preservation of photon statistics throughout the ray tracing allowing for quantitative estimates of diagnostic performance.

47 OTHER INSTRUMENTATION↗

Generative models on phase space

Deep generative models such as diffusion and flow matching are powerful machine learning tools capable of learning and sampling from high-dimensional distributions. They are particularly useful when the training data appears to be concentrated on a submanifold of the data embedding space. For high-energy physics data, consisting of collections of relativistic energy-momentum 4-vectors, this submanifold can enforce extremely strong physically-motivated priors, such as energy and momentum conservation. If these constraints are learned only approximately, rather than exactly, this can inhibit the interpretability and reliability of such generative models. To remedy this deficiency, we introduce generative models which are, by construction, confined at every step of their sampling trajectory to the manifold of massless N-particle Lorentz-invariant phase space in the center-of-momentum frame. In the case of diffusion models, the "pure noise" forward process endpoint corresponds to the uniform distribution on phase space, which provides a clear starting point from which to identify how correlations among the particles emerge during the reverse (de-noising) process. We demonstrate that our models are able to learn both few-particle and many-particle distributions with various singularity structures, paving the way for future interpretability studies using generative models trained on simulated jet data.

Bogorad, Zachary [Fermilab]↗

A Panoramic View of MXenes via an Atomic Coordination‐Based Design Strategy

Two‐dimensional (2D) transition metal carbides and nitrides, known as MXenes, possess unique physical and chemical properties, enabling diverse applications in fields ranging from energy storage to communication, catalysis, sensing, healthcare, and beyond. Despite extensive research and notable advancements, a fundamental understanding of MXenes’ phase diversity and its connection to their hierarchical precursors, including the intermediate MAX phases and the ancestral bulk phases, remains limited. Here, in this study, it is hypothesized that the atomic coordination environments adopted by transition metal and nonmetallic atoms in their three‐dimensional (3D) bulk precursors may persist in 2D MXenes to govern their phase diversity. Using high‐throughput modeling based on first‐principles density functional theory, a wide range of MXene phases is unveiled and comprehensively evaluate their relative stabilities across a large chemical space. The key to the approach lies in considering various atomic coordination environments drawn from four types of ancestral bulk phases. Through this comprehensive structural library of MXenes, general guiding principles are uncovered, such as a close alignment between the phase stability of MXenes and that of their 3D precursors. These findings introduce a new design strategy in which the atomic coordination environments in bulk phases can serve as reliable predictors for accessing the diverse structural landscape of MXenes.

MXenes↗

Hourly PM 2.5 Estimates across California from 2018 to 2023

This study presents a new data set of hourly PM 2.5 concentrations across California from 2018 to 2023 at a three-kilometer resolution. This data set was developed by assimilating observations from PurpleAir and the U.S. EPA Air Quality System monitors into wildfire smoke forecasts from the High-Resolution Rapid Refresh Smoke (HRRR-Smoke) model using the Gridpoint Statistical Interpolation (GSI) three-dimensional variational data assimilation framework. Archived forecasts of modeled wildfire smoke PM 2.5 from HRRR-Smoke create the background field for assimilation, which is then corrected using surface observations of total PM 2.5 . The resulting reanalysis from GSI provides an estimate of total PM 2.5 that minimizes error from both the observational and the model data. Validation results indicate strong performance, with monthly R 2 values ranging from 0.73 to 0.91 across the six-year data set, comparable to other PM 2.5 data sets. Case studies are presented for three major fire events, the 2018 Camp Fire, 2019 Kincade Fire, and 2020 Lightning Complex Fires to demonstrate the data set’s fidelity in resolving plume dynamics and local exposure patterns. Root-mean-squared error averaged over each month scales with average PM 2.5 concentrations, resulting in a low error under typical conditions but higher absolute errors during extreme smoke events. This is the first long-term, hourly PM 2.5 data set of its kind for California and enables the generation of subdaily exposure metrics, such as peak hourly concentrations, exceedance durations, and time-of-day exposure peaks. The novelty and strong validation of this data set make it a compelling resource for future studies on the impact and significance of subdaily PM 2.5 exposure.

PM2.5↗

Transfer learning nonlinear plasma dynamic transitions in low dimensional embeddings via deep neural networks

Deep learning algorithms provide a new paradigm to study high-dimensional dynamical behaviors, such as those in fusion plasma systems. Development of novel, data-driven model reduction methods, coupled with detection of abnormal modes with plasma physics, opens a unique opportunity to identify plasma instabilities through automated construction of parsimonious models that can be tuned to balance accuracy and cost. Our fusion transfer learning (FTL) model demonstrates success in rapidly reconstructing nonlinear kink mode structures by learning from a limited amount of nonlinear simulation data. The knowledge transfer process leverages a pre-trained neural encoder–decoder network, initially trained on linear simulations, to effectively capture nonlinear dynamics. The low-dimensional embeddings extract the coherent structures of interest, while preserving the inherent dynamics of the complex system. Experimental results highlight FTL’s capacity to capture transitional behaviors and dynamical features in plasma dynamics—a task often challenging for conventional methods. The model developed in this study is generalizable and can be extended broadly through transfer learning to address various magnetohydrodynamics modes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Performance evaluation of the USGS velocity model for the San Francisco Bay Area

In this study, we evaluated the performance of the United States Geological Survey velocity model developed for the San Francisco Bay Area (SFBA), version 21.1. The evaluation was performed through high-resolution three-dimensional physics-based ground motion simulations of seven small-magnitude earthquakes (ranging from magnitude 3.8 to 4.4) that occurred on the eastern side of the San Francisco Bay. The simulations were performed in the frequency range from 0 to 5 Hz with a minimum shear-wave velocity of 250 m/s, which allowed the capture of wave propagation effects of the near-surface soft materials that characterize local basins. Based on the direct comparison of Fourier amplitude spectra between recorded and simulated ground motions for more than 250 stations, we found that the velocity model generally performs well in the frequency range of 0.2–5 Hz. The median value of the Fourier amplitude residuals was found to be near zero for all seven earthquakes. The slight over-prediction of 0.2 log-natural units at frequencies above 3 Hz in our simulations was attributed to the potentially inaccurate representation of the source radiation pattern by a double-couple point source model, and simple representation of shallow small-scale underground structural complexity in the velocity model. Maps of spectral amplitude differences between the simulated and recorded data were used to identify areas responsible for systematic ground motion over-predictions or under-predictions. For example, while some sub-domains over soft sediments show over-prediction patterns, the block east of the Hayward fault is prone to exhibit patterns of under-prediction. These maps can be used to guide future refinements of the SFBA velocity model. Since our simulation methodology allows for the decoupling of the source and wave propagation effects, the ground motion data generated by our simulations can also be used to quantify the epistemic uncertainty due to the velocity model, in empirically based ground motion estimates for the SFBA.

58 GEOSCIENCES↗

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING↗

Asynchronicity in opposed-piston RCMs: Does it matter?

Rapid Compression Machines (RCMs) are widely utilized to study combustion phenomena at engine-relevant conditions, and significant efforts are typically made to create a quiescent environment, particularly for investigations of autoignition chemistry. Opposed-piston configurations can be advantageous due to shorter compression times and reduced surface area to volume ratios. Each side must be actuated simultaneously, but this can be challenging in practice. These devices, like most RCMs, utilize hydraulics for actuation, speed control and arrestation of the piston at the end of the stroke; there is no mechanical control or linkage of the two piston trajectories. To quantify the magnitudes and effects of piston asynchronous behavior, this work employs both detailed experimental measurements and, for the first time, high-fidelity, Direct Numerical Simulation (DNS). The boundary conditions are carefully considered applying insight from high-resolution linear variable differential transformer (LVDT) measurements of the piston trajectory and a zero-dimensional kinematics model of the piston-shaft assembly. Sufficient resolution in the piston crevice region is used. The complicated fluid dynamical behavior that can evolve during piston compression and the ensuing delay processes due to offset timings from t offset = 0-10 ms is elucidated. It is found that near t offset = 6 ms and beyond, the boundary layer on the face of the first-seating piston can be sufficiently perturbed, due initially to reemergence of gas from the crevice of the firstseating piston, so that the adiabatic core can become degraded at long ignition delay times. Substantial mixing of colder gas into the interior of the reaction chamber can alter the measurements, similar to effects previously observed for improper piston crevice configuration. In conclusion, experimental techniques to mitigate asynchronous behavior are discussed and demonstrated.

33 ADVANCED PROPULSION SYSTEMS↗