Search NASA⌕ Search

SEARCH · Search NASA

Results for “large scale analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Software For Advanced Large-scale Analysis Of Magnetic Confinement For Numerical Design, Engineering & Research (salamander)

As magnetic confinement fusion energy gains traction internationally to enable abundant energy production, designing components for fusion systems is a pressing challenge. During the planned lifetime of a fusion device, components evolve in extreme environments and must withstand large, repeated thermal loads and bombardment by 14 MeV neutrons, plasma ions, and neutral particles (deuterium, tritium, and helium), corrosive conditions, etc. All these physical processes take place simultaneously, interact in intricate ways, and impose important constraints that can affect performance. Experimental data is rare and costly to obtain, making design particularly challenging. Predictive computational frameworks must be an integral part of an accelerated and cost-effective design process by modeling fusion system performance in simulated environments. To better understand component degradation and operational impacts on their performance, the Software for Advanced Large-scale Analysis of MAgnetic confinement for Numerical Design, Engineering & Research (SALAMANDER) is designed as an open-source, fully integrated, multiphysics, multiscale, NQA-1 compliant framework facilitating 3D, high-fidelity fusion system modeling. To that end, SALAMANDER is a MOOSE-based framework, and therefore leverages MOOSE upstream libraries such as PETSc and libMesh to deliver sophisticated finite element, finite volume, and nonlinear solver technology for fusion energy simulations. SALAMANDER couples MOOSE physics module capabilities—such as thermal hydraulics, heat conduction, Navier-Stokes, and thermomechanics—with tritium transport via TMAP8, neutronics via Cardinal, and nascent particle-in-cell capabilities. Direct simulation Monte Carlo methods will be used to address neutral transport near the walls. By coupling all these physics in an integrated application, SALAMANDER will enable high-fidelity modeling of irradiation levels and plasma exposure conditions of plasma facing components and their impact on heat and tritium distributions, as well as the resulting mechanical constraints experienced by the plasma facing components and performance of blanket systems. Furthermore, SALAMANDER will be particularly suited for engineering studies thanks to the stochastic tool module readily available in MOOSE, allowing for extended uncertainty quantification and risk analysis studies. It is also able to use computer-aided design (CAD) meshes to model complex geometries, which is indispensable for fusion systems. SALAMANDER therefore supports design, safety, engineering, and research projects for magnetic confinement fusion systems

Simon, Pierre-Clement [Idaho National Laboratory (↗

A Large-Scale Analysis to Optimize the Control and V2V Communication Protocols for CDA Agreement-Seeking Cooperation

Cooperative driving automation (CDA) Class C, agreement-seeking cooperation, is an innovative and practical solution that can promote cooperation among general passenger vehicles on the road. However, more comprehensive studies are needed before establishing the standard protocols of agreement-seeking cooperation, such as communication frequency and the duration of cooperation. Here, this article presents an initiative study on the impacts of communication capabilities on agreement-seeking cooperation. Through a large-scale analysis by regulating vehicle-to-vehicle (V2V) communication metrics, this work suggests desirable system parameters that can maximize the benefits of cooperation and ensure reliable operability while avoiding exhaustive communication loads. As the first step, an example agreement-seeking cooperation system is created for a car-following scenario, including decision-making and control algorithms for autonomous vehicles. Then, software-in-the-loop tests explore the performance of the developed system as it encounters various communication risks, such as latency and message packet drops. The system performance metrics are evaluated from various angles, including the time consumed for the agreement-seeking process, cooperation ratio, and the ratio of faulty cooperation. Energy saving from the cooperation is assessed by using simulation software that can run multiple high-fidelity vehicle models simultaneously. Based on the analyses, this article suggests the V2V communication requirements for the reliable operation of CDA agreement-seeking, which can be referred to when developing the standard protocols of agreement-seeking cooperation.

42 ENGINEERING↗

Large-scale analysis of structural brain asymmetries in schizophrenia via the ENIGMA consortium

Left–right asymmetry is an important organizing feature of the healthy brain that may be altered in schizophrenia, but most studies have used relatively small samples and heterogeneous approaches, resulting in equivocal findings. We carried out the largest case–control study of structural brain asymmetries in schizophrenia, with MRI data from 5,080 affected individuals and 6,015 controls across 46 datasets, using a single image analysis protocol. Asymmetry indexes were calculated for global and regional cortical thickness, surface area, and subcortical volume measures. Differences of asymmetry were calculated between affected individuals and controls per dataset, and effect sizes were meta-analyzed across datasets. Small average case–control differences were observed for thickness asymmetries of the rostral anterior cingulate and the middle temporal gyrus, both driven by thinner left-hemispheric cortices in schizophrenia. Analyses of these asymmetries with respect to the use of antipsychotic medication and other clinical variables did not show any significant associations. Assessment of age- and sex-specific effects revealed a stronger average leftward asymmetry of pallidum volume between older cases and controls. Case–control differences in a multivariate context were assessed in a subset of the data (N = 2,029), which revealed that 7% of the variance across all structural asymmetries was explained by case–control status. Subtle case–control differences of brain macrostructural asymmetry may reflect differences at the molecular, cytoarchitectonic, or circuit levels that have functional relevance for the disorder. Reduced left middle temporal cortical thickness is consistent with altered left-hemisphere language network organization in schizophrenia.

60 APPLIED LIFE SCIENCES↗

Design and implementation of I/O performance prediction scheme on HPC systems through large-scale log analysis

Abstract Large-scale high performance computing (HPC) systems typically consist of many thousands of CPUs and storage units used by hundreds to thousands of users simultaneously. Applications from large numbers of users have diverse characteristics, such as varying computation, communication, memory, and I/O intensity. A good understanding of the performance characteristics of each user application is important for job scheduling and resource provisioning. Among these performance characteristics, I/O performance is becoming increasingly important as data sizes rapidly increase and large-scale applications, such as simulation and model training, are widely adopted. However, predicting I/O performance is difficult because I/O systems are shared among all users and involve many layers of software and hardware stack, including the application, network interconnect, operating system, file system, and storage devices. Furthermore, updates to these layers and changes in system management policy can significantly alter the I/O behavior of applications and the entire system. To improve the prediction of the I/O performance on HPC systems, we propose integrating information from several different system logs and developing a regression-based approach to predict the I/O performance. Our proposed scheme can dynamically select the most relevant features from the log entries using various feature selection algorithms and scoring functions, and can automatically select the regression algorithm with the best accuracy for the prediction task. The evaluation results show that our proposed scheme can predict the write performance with up to 90% prediction accuracy and the read performance with up to 99% prediction accuracy using the real logs from the Cori supercomputer system at NERSC.

97 MATHEMATICS AND COMPUTING↗

The DESI-Lensing Mock Challenge: large-scale cosmological analysis of 3x2-pt statistics

The current generation of large galaxy surveys will test the cosmological model by combining multiple types of observational probes. Realising the statistical promise of these new datasets requires rigorous attention to all aspects of analysis including cosmological measurements, modelling, covariance and parameter likelihood. In this paper we present the results of an end-to-end simulation study designed to test the analysis pipeline for the combination of the Dark Energy Spectroscopic Instrument (DESI) Year 1 galaxy redshift dataset and separate weak gravitational lensing information from the Kilo-Degree Survey, Dark Energy Survey and Hyper-Suprime-Cam Survey. Our analysis employs the 3x2-pt correlation functions including cosmic shear and galaxy-galaxy lensing, together with the projected correlation function of the spectroscopic DESI lenses. We build realistic simulations of these datasets including galaxy halo occupation distributions, photometric redshift errors, weights, multiplicative shear calibration biases and magnification. We calculate the analytical covariance of these correlation functions including the Gaussian, noise and super-sample contributions, and show that our covariance determination agrees with estimates based on the ensemble of simulations. We use a Bayesian inference platform to demonstrate that we can recover the fiducial cosmological parameters of the simulation within the statistical error margin of the experiment, investigating the sensitivity to scale cuts. This study is the first in a sequence of papers in which we present and validate the large-scale 3x2-pt cosmological analysis of DESI-Y1.

79 ASTRONOMY AND ASTROPHYSICS↗

Hydropower Infrastructure - LAkes, Reservoirs, and RIvers (HILARRI), v4

HILARRI is a database of links between major datasets of operational hydropower dams and powerplants, and inland water bodies. These connections are critical for conducting large-scale analysis of hydropower infrastructure and their associated natural and engineered water systems. Features include: – Dams from the National Inventory of Dams (2025) and the Global Reservoir and Dam Database (GRanD v1.3) – Hydropower plants from the Existing Hydropower Assets dataset (EHA 2025) – Power plants that are listed in the 2025 U.S. Hydropower Development Pipeline Data or were listed in previous versions of the dataset These hydropower infrastructure features are linked to several major datasets that provide hydrologic and hydraulic information relevant for analysis of hydropower systems that includes the integral water resources. That information comes from: – Products from the National Hydrography Dataset (NHD) – NHDPlusV2 Medium Resolution river network flowlines, – NHD waterbodies (limited to lakes and reservoirs), – NHD Watershed Boundary Dataset (HUC12-level for the Conterminous United States (CONUS)) – NHD High Resolution waterbodies – HydroLAKES water bodies (lakes and reservoirs) – LAGOS-US lakes and reservoirs – EPA National Lakes Assessment (2007, 2012, 2017, and 2022) – The Reservoir Sedimentation Database (RESSED) – EPA SuRGE sampling locations Unique identifiers are used to facilitate joining to the original full datasets. For example, characteristics of NHD flowlines such as estimated average flow rate can be joined from the NHDPlusV2 dataset to a dam or power plant listed in HILARRI based on the ID field, “COMID”, that is common to both datasets. HILARRI only includes basic information about identifiers, location, and data quality or usage notes. It does not contain the attributes or time series data associated with these sites. The HILARRI dataset incorporates information from several datasets to facilitate more effective and accurate analysis of hydropower infrastructure and their associated waterbodies. For example, dams were checked against the most recent American Rivers Dam Removal Database to identify and flag facilities that may no longer exist. Additionally, dams that are listed multiple times in the NID are identified and flagged to avoid double-counting when analyzing and summarizing information. Other quality flags include certainty of operational hydropower (i.e., if one or more datasets indicates hydropower at a particular location), whether an associated water body is accurate or composed of multiple polygons, or whether there is a known issue with reported characteristics in one of the underlying datasets. These additional data flags are designed to increase confidence in data usage for individual to large-scale analyses.

Hansen, Carly [ORNL] (ORCID:0000000193280838)↗

Probabilistic Error Bounds for Low-Rank Tensor Decompositions Used in Large-Scale Data Analysis Applications (LDRD Final Report)

This report documents a research project on analyzing low-rank tensor models for data analysis that took place at Sandia National Laboratories from October 2023–September 2025. The focus of this work was to extend theoretical frameworks from statistics and probability theory for use with models for scalar, vector, and matrix data to models with tensor, or general multi-dimensional array, data. Through this work, we have provided a new set of tools for bounding errors on low-rank tensor models of both complete and sampled data. The remainder of this report is organized as follows. In Section 1, we describe the proposed work at the start of the project. Section 2 describes the research advances made as part of the project. Other research contributions in the form of conference presentations and software development is provided in Section 3. Workforce development at Sandia and Florida Atlantic University (via a subcontract on this project) is provided in Section 4.

97 MATHEMATICS AND COMPUTING↗

Large-scale spatially explicit analysis of carbon capture at cellulosic biorefineries

The large-scale production of cellulosic biofuels would involve spatially distributed systems including biomass fields, logistics networks and biorefineries. Better understanding of the interactions between landscape-related decisions and the design of biorefineries with carbon capture and storage (CCS) in a supply chain context is needed to enable efficient systems. Here we analyse the cost and greenhouse gas mitigation potential for cellulosic biofuel supply chains in the US Midwest using realistic spatially explicit land availability and crop productivity data and consider fuel conversion technologies with detailed CCS design for their associated CO 2 streams. Optimization methods identify trade-offs and design strategies leading to systems with attractive environmental and economic performance. Strategic and operational decisions depend on underlying spatial features and are sensitive to biofuel demand and CCS incentives. US CCS incentives neglect to motivate greenhouse gas mitigation from all supply chain emission sources, which leverage spatial interactions between CCS, electricity prices and the biomass landscape.

09 BIOMASS FUELS↗

Thermal Hydraulics Analysis of a Divertor Monoblock Using SALAMANDER

Divertors are critical components in magnetic confinement fusion devices. One of the Divertor's crictial roles is absorbing the highest heat flux from the plasma. The Divertor accomplishes this through the use of divertor monoblocks and cooling channels. This work presents a thermal hydraulic analysis of a divertor monoblock using the Software for Advanced Large-scale Analysis of MAgnetic confinement for Numerical Design, Engineering and Research (SALAMANDER). SALAMANDER is an open source tool developed at Idaho National Laboratory for conducting multiphysics and multscale analysis of Plasma Facing Components (PFCs). For the monoblock, solid heat conduction is modeled in the block with a constant heat flux of 1E7 W/m^2 applied to the top of the block to represent peak heat flux from the plasma. The cooling channel (which is pressurized water at a Reynolds number of ~1,000,000) is modeled using a k-epsilon turbulence model with standard wall functions. The heat transfer between the solid and fluid domain is modeled using the Dittus-Boelter correlation for the convective heat transfer coefficient. This work aims to show the importance of a multiphysics modeling approach when performing simulations on PFCs.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Optimization Formulations for Storage Devices with Disjoint Operating Modes

Optimizing State of Charge for Storage Resources Storage of generated energy has existed as a contributing resource within the power sector for more than a century in the case of pumped hydro storage. Even though most regional transmission organizations/independent system operators do not currently optimize state of charge (SOC) for storage, the SOC is an essential aspect of the operating characteristics of storage. This paper investigates the impact on computational performance of optimizing storage with explicit representation of SOC. To investigate the model, the paper considers two contexts, stand-alone and large-scale. Analysis of the stand-alone context helps to explain why the combination of features in the storage model results in a difficult problem. Numerical results for large-scale test cases verify that new tighter valid inequalities can improve the computation compared with the standard model in the literature.

Business & Economics↗

Long‐Term Large‐Scale Atmospheric Forcing Data From Three‐Dimensional Constrained Variational Analysis for the ARM SGP Site

Here, this study presents a long‐term three‐dimensional large‐scale forcing data set (VARANAL3D) derived from the three‐dimensional constrained variational analysis (3DCVA) method at the Atmospheric Radiation Measurement (ARM) program Southern Great Plains (SGP) site from 2004 to 2018. Building on the same input data sets as the conventional continuous forcing data set (VARANAL), VARANAL3D maintains overall consistency in domain‐averaged fields while introducing spatial variability, offering critical insights into the influence of mesoscale synoptic systems on cloud‐related processes. Evaluations are conducted across four cloud and precipitation regimes: Clear‐sky, Shallow‐clouds, Afternoon‐precipitation, and Nocturnal‐precipitation, presenting high consistency of the domain‐mean forcing data sets while emphasizing the role of subdomain forcing variability particularly in precipitating regimes. Single column model (SCM) simulations demonstrate that subdomain VARANAL3D forcing improves cloud and precipitation representation, with the ensemble outperforming domain‐mean forcing in three cloudy and precipitating regimes. Overall, these results highlight VARANAL3D's value for investigating the impacts of spatial variability of large‐scale forcing on atmospheric processes. The VARANAL3D data set provides new opportunities for evaluating model physics, advancing the development of scale‐aware parameterizations and deepening our understanding of cloud and precipitation dynamics.

Environmental sciences↗

Simplifying Geospatial Workflows with GeoGridFusion

Growing demands to understand PV deployment and reliability across an expanding range of climates, along with increasing computational power, are driving the need for simplified computing tools that support large-scale analysis and geospatial workflows. While existing open-source libraries and PV system modeling tools offer extensive collections of empirical and physical models, they often lack the ability to scale to large geospatial datasets. The tool presented here, GeoGridFusion, enables PV modelers to store and utilize gridded geospatial data from sources such as NSRDB, PVGIS, and others. GeoGridFusion harmonizes diverse datasets, taxonomies, and nomenclatures, and includes utilities that support intuitive geospatial area selection.

14 SOLAR ENERGY↗

Hydropower Infrastructure – LAkes, Reservoirs, and RIvers (HILARRI)

HILARRI is a database of links between major datasets of operational hydropower dams and powerplants, and inland water bodies. These connections are critical for conducting large-scale analysis of hydropower infrastructure and their associated natural and engineered water systems. Features include: – Dams from the National Inventory of Dams (2024) and the Global Reservoir and Dam Database (GRanD v1.3) – Hydropower plants from the Existing Hydropower Assets dataset (EHA 2024) These hydropower infrastructure features are linked to several major datasets that provide hydrologic and hydraulic information relevant for analysis of hydropower systems that includes the integral water resources. That information comes from: – Products from the National Hydrography Dataset (NHD) – NHDPlusV2 Medium Resolution river network flowlines, – NHD waterbodies (limited to lakes and reservoirs), – NHD Watershed Boundary Dataset (HUC12-level for the Conterminous United States (CONUS)) – NHD High Resolution waterbodies – HydroLAKES water bodies (lakes and reservoirs) – LAGOS-US lakes and reservoirs – EPA National Lakes Assessment (2007, 2012, 2017, and 2022) – The Reservoir Sedimentation Database (RESSED) Unique identifiers are used to facilitate joining to the original full datasets. For example, characteristics of NHD flowlines such as estimated average flow rate can be joined from the NHDPlusV2 dataset to a dam or power plant listed in HILARRI based on the ID field, “COMID”, that is common to both datasets. HILARRI only includes basic information about identifiers, location, and data quality or usage notes. It does not contain the attributes or time series data associated with these sites. The HILARRI dataset incorporates information from several datasets to facilitate more effective and accurate analysis of hydropower infrastructure and their associated waterbodies. For example, dams were checked against the most recent American Rivers Dam Removal Database to identify and flag facilities that may no longer exist. Additionally, dams that are listed multiple times in the NID are identified and flagged to avoid double-counting when analyzing and summarizing information. Other quality flags include certainty of operational hydropower (i.e., if one or more datasets indicates hydropower at a particular location), whether an associated water body is accurate or composed of multiple polygons, or whether there is a known issue with reported characteristics in one of the underlying datasets. These additional data flags are designed to increase confidence in data usage for individual to large-scale analyses.

13 HYDRO ENERGY↗

Screening Cu-Zeolites for Methane Activation Using Curriculum-Based Training

Machine learning (ML), when used synergistically with atomistic simulations, has recently emerged as a powerful tool for accelerated catalyst discovery. However, the application of these techniques has been limited by the lack of interpretable and transferable ML models. In this work, we propose a curriculum-based training (CBT) philosophy to systematically develop reactive machine learning potentials (rMLPs) for high-throughput screening of zeolite catalysts. Our CBT approach combines several different types of calculations to gradually teach the ML model about the relevant regions of the reactive potential energy surface. The resulting rMLPs are accurate, transferable, and interpretable. We further demonstrate the effectiveness of this approach by exhaustively screening thousands of [CuOCu] 2+ sites across hundreds of Cu-zeolites for the industrially relevant methane activation reaction. Specifically, this large-scale analysis of the entire International Zeolite Association (IZA) database identifies a set of previously unexplored zeolites (i.e., MEI, ATN, EWO, and CAS) that show the highest ensemble-averaged rates for [CuOCu] 2+ -catalyzed methane activation. We believe that this CBT philosophy can be generally applied to other zeolite-catalyzed reactions and, subsequently, to other types of heterogeneous catalysts. Thus, this represents an important step toward overcoming the long-standing barriers within the computational heterogeneous catalysis community.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Factorized visual representations in the primate visual system and deep neural networks

Object classification has been proposed as a principal objective of the primate ventral visual stream and has been used as an optimization target for deep neural network models (DNNs) of the visual system. However, visual brain areas represent many different types of information, and optimizing for classification of object identity alone does not constrain how other information may be encoded in visual representations. Information about different scene parameters may be discarded altogether (‘invariance’), represented in non-interfering subspaces of population activity (‘factorization’) or encoded in an entangled fashion. In this work, we provide evidence that factorization is a normative principle of biological visual representations. In the monkey ventral visual hierarchy, we found that factorization of object pose and background information from object identity increased in higher-level regions and strongly contributed to improving object identity decoding performance. We then conducted a large-scale analysis of factorization of individual scene parameters – lighting, background, camera viewpoint, and object pose – in a diverse library of DNN models of the visual system. Models which best matched neural, fMRI, and behavioral data from both monkeys and humans across 12 datasets tended to be those which factorized scene parameters most strongly. Notably, invariance to these parameters was not as consistently associated with matches to neural and behavioral data, suggesting that maintaining non-class information in factorized activity subspaces is often preferred to dropping it altogether. Thus, we propose that factorization of visual scene information is a widely used strategy in brains and DNN models thereof.

59 BASIC BIOLOGICAL SCIENCES↗

Kinetic Plasma Simulation Capabilities in the MOOSE Framework: Verification of Particle-Particle Collisions

High-fidelity simulations of complex plasma systems allow researchers to gain key insights into and understanding of these systems. To facilitate massively parallel high-fidelity plasma simulations, finite-element-based particle-in-cell capabilities are being developed within the open-source Multiphysics Object-Oriented Simulation Environment (MOOSE) based framework called Software for Advanced Large-scale Analysis of MAgnetic confinement for Numerical Design, Engineering & Research (SALAMANDER). While SALAMANDER’s primary objective is modeling edge plasmas and plasma-facing components in fusion devices, the particle-in-cell capabilities being developed are general and will support modeling low-temperature plasmas as well. Previously, collisionless magnetostatic simulation capabilities have been verified with the two-stream and Dorey-Guest-Harris instabilities, and single particle motion. Collisions were implemented using the direct simulation Monte Carlo method, and verification of this capability will be presented here several verification problems: relaxation of a randomly initialized gas to a Maxwellian distribution, Fourier heat flow, and comparison of reaction rates to both analytic calculations and those calculated using a multi-term Boltzmann solver.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY↗