Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data, Modeling, and Analysis Program”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Improving Dose Modeling With Dynamic Modeling Tools [Slides]

Utilities aiming for higher fuel enrichment for power uprates or extended operation times before refueling must conduct a new dose analysis. Current conservative dose estimation standards may cause utilities to exceed regulatory limits for proposed increased fuel enrichment. A more accurate modeling of doses from reactor accidents can lower these conservative assumptions. Prescott et al. (2022) demonstrated that the Event Modeling Risk Assessment using Linked Diagrams (EMRALD) software tool, developed at Idaho National Laboratory (INL), can be coupled with the Modular Accident Analysis Program (MAAP5) for dynamic accident analysis in reactor plants. EMRALD forms models of potential accident scenarios, while MAAP5 simulates the accident progression and dose consequences. By integrating these software tools with utility-specific data, a more precise estimation of dose consequences from plant accidents can be achieved. Preliminary findings indicate that EMRALD provides accurate mean core damage frequencies for generalized accident scenarios. Future work includes expanding the model to account for plant-specific data and mitigation factors.

97 - MATHEMATICS AND COMPUTING↗

Characterization and Quantification of Radiation-Induced Clusters/Precipitates in RPV Steels Using STEM-EDS and Machine Learning

Over the operational lifespan of a nuclear reactor, reactor pressure vessel (RPV) steels are subjected to significant neutron irradiation, resulting in complex microstructural changes and the consequent degradation of mechanical properties. Various physically motivated correlation models have been developed to predict neutron irradiation-induced embrittlement of RPVs under different irradiation conditions. However, the efficient and accurate characterizations and quantification of radiation-induced clusters in RPVs are still challenging, which will affect the precision of the predictive models for embrittlement of RPV components. In the DOE Visiting Faculty Program (VFP) research work at Oak Ridge National Lab (ORNL), I integrate machine learning to aid Scanning Transmission Electron Microscopy – Energy Dispersive X-ray Spectroscopy (STEM-EDS) analyses, which improve the characterization and quantification of radiation-induced clusters in RPV steels, thereby enabling more accurate predictions of material behavior under irradiation. The surveillance base- and welded- RPV steels were annealed at various temperatures of 340 °C, 450 °C and 500 °C for up to 168 hours, respectively. Afterwards, I have characterized radiation-induced clusters using advanced STEM-EDS techniques and subsequently applying machine learning algorithms to analyze and refine STEM-EDS datasets, enhancing the quantification of clusters compositions and distributions. In the end, an efficient workflow for integrating STEM-EDS data analysis with machine learning to address challenges including noise reduction has been developed. The completion of this VFP work will support bridge critical gaps in the accurate quantification of radiation-induced clusters in RPV steels using STEM-EDS and support the development of more precise models for predicting RPV embrittlement in the Light Water Reactor Sustainability program supported by Department of Energy and enhancing the collaboration between ORNL and Alred University. The outcome of the VFP project will leverage a few research papers submission to peer-reviewed journals in the relevant scientific field and a few oral presentations at national and international conferences.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

pvplr-python: Python package implementation of PVplr for Performance Loss Rate (PLR) analysis

Due to software fragmentation, PV system modeling teams can be limited to language specific packages, preventing cross-sectional analysis of different modeling techniques and workflows. To this end, PVplr, a popular PV performance modeling R software package, has been ported to the Python programming language. To verify and test the robustness of the port, NSRDB data has been used to simulated PV installations at native resolution (~2 million Sites), with a variety of degradation rates, degradation patterns, and modules. Performance Ratios were calculated using the ported functions from pvplr-python and compared against Rdtools YoY values. Due to the complicated nature of degradation, a new metric has been proposed to quantify the performance loss of a system. The cumulative production loss, is the total amount of energy lost due to the degrading performance of the system. Cumulative production loss alleviates the problems with fitting linear functions to non-linear degradation. Cumulative Production loss was shown to better estimate the total lost revenue for non-linear degradation patterns. $XbX + UTC$ was found to most accurately predict the total lost revenue in simulated systems.

Kumar, Suraj↗

Progress Towards the Validation of a new RELAP5-3D model of the High Temperature Test Facility

Validation is a key step in the development of any type of systems model. As the next generation of reactors approaches, the need for codes that have been validated for these new types of systems continues to grow. An example of a prominent option is the Reactor Excursion Leak Analysis Program (RELAP5-3D), developed by Idaho National Laboratory. This code was developed for the purpose of systems level thermal-hydraulic modeling of light water reactors (LWRS) and postulated transients that can occur in LWRS.RELAP5-3D has been substantially validated against LWR data. Due to its long history as a reactor safety analysis tool, there has been an effort to adapt RELAP5-3D for the purposes of advanced reactor concepts such as prismatic high-temperature gas-cooled reactors (HTGRs). However, RELAP5-3D has not nearly been validated and verified for HTGRs to the degree of LWRs, warranting verification and validation opportunities with computational benchmarks and existing experimental facilities. Examples of such facilities include the modular high-temperature gas-cooled reactor (MHTGR) 350 and the high temperature engineering test reactor (HTTR) from Japan. The MHTGR 350 is a benchmark design concept for code-to-code verification purposes; therefore, it does not provide any experimental data for validation opportunities The HTTR provides useful multiphysics validation data but does not have the in-core instruments to generate thermal-hydraulic experimental data to help with RELAP5-3D validation. Consequently, a facility that could provide key in-core temperatures for thermal-hydraulic validation was still needed. The High Temperature Test Facility (HTTF) is an integral effects facility for HTGR thermal hydraulics developed and operated by Oregon State University. HTTF represents ¼ length scale of the General Atomics MHTGR and is rated for a total power of 2.2 MW. Axially, the core consists of an upper and lower reflector and 10 blocks, numbered from bottom to top (Block 1 is right above lower reflector). The core is heated via graphite resistive heater rods, with respective channels distributed throughout the core. The primary coolant is helium and heat can radiate out of the core to the reactor cavity cooling system (RCCS), which is cooled by water. The primary purpose of the facility is to investigate pressurized conduction cooldown (PCC) and depressurized conduction cooldown (DCC) transients, which are also referred to as the pressurized and depressurized loss of forced cooling respectively. Two experiments were chosen to perform the validation study with a RELAP5-3D model of HTTF. These experiments are PG-27 (PCC) and PG-29 (DCC). These were chosen based off of the quality of available experimental data before and during the experiment which led to their inclusion in the HTGR Thermal Hydraulics Benchmark.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Validation of the DESI DR2 Ly$α$ forest full-shape analysis

We present the validation of the Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) Lyman-$α$ (Ly$α$) forest full-shape analysis. This analysis combines three-dimensional Ly$α$ forest auto-correlations and cross-correlations with quasars to extract information from both the baryon acoustic oscillation (BAO) feature and the broadband clustering signal, with primary emphasis on the Alcock-Paczynski (AP) measurement. Compared to the DESI DR1 analysis, the DR2 validation uses substantially larger and more realistic mock datasets, including CoLoRe 2LPT and AbacusSummit Ly$α$ forest simulations. The modeling framework is also improved through analytic marginalization over small scales ($<10$$h^{-1}$Mpc) and the impact of ultraviolet background fluctuations. The validation program was completed prior to unblinding and defines quantitative requirements for the cosmological parameters of interest, which are evaluated using hundreds of mock realizations. We further test the analysis through independent fits to the auto- and cross-correlations, multiple catalog splits, and a broad suite of analysis and modeling variations applied to both mocks and blinded observational data. We find that the BAO and AP parameters satisfy all validation requirements and remain stable across all tests. In contrast, mock studies reveal a significant bias in the inferred growth-rate parameter $fσ_8$, leading us to exclude this measurement from the final analysis. The consistency across mocks, data splits, and robustness tests demonstrates that the DR2 Ly$α$ full-shape analysis provides a reliable and substantially improved broadband AP measurement over previous Ly$α$ forest studies.

Herbold, M. [Chicago U., KICP; Ohio State U.] (ORC↗

Summary Report Of The FY25 Reactor Physics Verification And Validation Exercises In The Advanced Reactor Technologies - Gas-cooled Reactor Program

Valdiation and verification of numerical tools is critical for ensuring reasonable predictions for design scoping, licensing, and safety analsyis. In this report, two reactor physics verification and validation exercises are presented. The first of these exercises focuses on burnup analysis with data from the Advanced Gas Reactor (AGR) program. Simulations are performed with Monte Carlo N-Particle (MCNP) and are compared with the experimental measurements for the AGR 1 and 2 experiments that utilize both UCO and UO2 fuel. The second exercises utilizes data from the HTR-Proteus experiments to perform reactor physics validation. Specifications of the experimental facility are provdied, along with a demonstration of initial modeling efforts in Serpent for one of the determistic packing experiments. Both cases are part of the Generation-IV international forum (GIF) Very High-Temperature Reactor (VHTR) Computational Methods, Validation, and Benchmarking (CMVB) program, an international collaborative organization dedicated to the verification and validation of High-Temperature Gas-Cooled Reactor (HTGR) analysis. Participation in the CMVB allows the US Department of Energy (DOE) to leverage these existing validation activities to provide extra value through benchmarking activities with other CMVB members.

and Benchmarking (CMVB) program↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Data-driven modeling of dynamic occupant thermostat override behavior for demand response applications

Buildings consume nearly 40% of global energy and produce similar emissions. Whiletechnological advances address efficiency, occupant behavior causes energy use variations up to 300% between identical buildings. This gap between predicted and actual building performance impacts building design, operations, and grid demand management programs. Through analyses of smart thermostat data from 1,400 single-occupant homes, the researchdemonstrates that occupants respond to 8°F thermostat setpoint changes within a median of 15 minutes, while 2°F changes trigger responses within a median of 30 minutes. This highlights an understudied temporal relationship between thermostat setbacks and response time of occupant behaviors. Models of such behavior dynamics are required to incorporate occupant impacts into building performance simulation. A key contribution of this dissertation is the Thermal Frustration Theory (TFT), which positsthat thermal discomfort driven behaviors are caused by the time-accumulation of discomfort, not simply a temperature deviation threshold or a delay from an initiating event. Using a dataset of 634 thermostats, each with 25+ manual setpoint changes, a comparative analysis of TFT and comfort zone and a delayed response theories demonstrated that personalized TFT models better predict when manual setpoint change occur. This was measured by the area under the curve statistical measure (AUC); all three models perform similarly by a Matthews Correlation Coefficient measure. Higher AUC performance is especially important for modeling occupant behavior in demand response programs where false negatives of rare occupant interactions could adversely affect grid stability. EnergyPlus based simulations were conducted with TFT-derived occupant models, demonstrating the ability to identify parameters of known TFT models from only data observable with smart thermostats, even under the presence of noise from routine overrides. Overall, the dissertation highlights that thermostat interactions are neither static,instantaneous, nor driven solely by the environment. Instead, temporal accumulation of discomfort and routine-based behavior play important roles. The methodology and results offer a pathway towards more accurate modeling of human-building interactions for policy assessment, building design, and demand response programs.

Sharma, Kunind [Northeastern University] (ORCID:00↗

Metabolic flux and resource balance in the oleaginous yeast Rhodotorula toruloides

The yeast Rhodotorula toruloides is a promising bioproduction organism due to its high lipid yields and ability to grow on cheap and abundant substrates. Quantitative, systems-level assessment of its metabolic activity is accordingly merited. Resource-balance analysis (RBA) models capture not only reaction stoichiometry but also enzyme requirements for catalysis, providing valuable tools for understanding metabolic trade-offs and optimizing metabolic engineering strategies. Here, in this work, we present systems-level measurements of R. toruloides metabolic flux based on isotope tracing and metabolic flux analysis. In combination with new proteomic measurements, these flux data are used to parameterize a genome-scale resource balance model rtRBA. We find that S. cerevisiae and R. toruloides grow at nearly indistinguishable rates using similar biosynthetic but dramatically different central metabolic programs. R. toruloides consumes one-fifth as much glucose, which it metabolizes primarily via the pentose phosphate pathway and TCA cycle unlike primarily glycolysis in S. cerevisiae . Overall, across these two divergent yeasts, protein abundances aligned more closely than metabolic flux. Resource balance modeling of these metabolic programs predicts superior theoretical yields but lower productivities in R. toruloides than S. cerevisiae for industrial chemicals, highlighting the value of rapid glucose uptake for productivity but respiratory metabolism for yields.

60 APPLIED LIFE SCIENCES↗

Application of a Physics-Informed Convolutional Neural Network for Monitoring the Temperature Fields in High-Temperature Gas Reactors

Here, this work presents current advances in applying a physics-informed convolutional neural network (CNN) to evaluate temperature distributions in advanced reactors. Our goal is to demonstrate that the CNN can reconstruct temperature fields within the solid region of a prismatic fuel assembly in a high-temperature gas reactor (HTGR) with sensor data available in only a few cooling channels. Before that, we showcase the superior performance of the physics-informed CNN in comparison to a purely data-driven multilayer perceptron (MLP), considering a canonical heated channel setup. This analysis shows the advantages of our approach and justifies its choice. The datasets employed here are obtained upon numerical simulations performed with codes under the Nuclear Energy Advanced Modeling and Simulation program. This work is important, as industry experience indicates that the assembly material in HTGR concepts is prone to large thermal-mechanical loads nearing operational limits. This makes it crucial to characterize peak temperatures and their distributions near hot spots. Modern thermocouples are unreliable in these types of harsh environments because of the high neutron fluxes and elevated temperatures involved. The CNN-based field reconstruction represents an attractive solution, enabling sensor arrays in less aggressive locations and augmenting indirect predictions for less accessible regions. The results show that the CNN reduces prediction errors by orders of magnitude in comparison to the MLP, considering the simple yet well-representative heated channel case. In the case of the HTGR fuel assembly, the CNN can successfully reconstruct temperature fields over various cooling regimes. Furthermore, we also explore the algorithm’s ability to detect abnormalities. Interestingly, the CNN proves it has the capacity to detect blockage in one of the noninstrumented cooling channels.

Machine learning↗

Lawrence, Massachusetts, Residential Building Efficiency and Electrification Analysis [Slides]

As part of the Communities Local Energy Action Program (CLEAP), the Lawrence Stakeholders Coalition (LSC) is interested in assessing and understanding the potential for and pathways to electrification for the City of Lawrence. The LSC's main questions are: What is the impact of various electrification packages on residential electricity bills and what types of buildings should the LSC target for electrification plus weatherization packages? This technical assistance, using ResStock tool modeling, aims to assist the Coalition's electrification and energy burden reduction planning by: 1. Providing cross-cutting data on housing stock characteristics, energy burden characteristics, fuel types, energy consumption, and system efficiency; and 2. Providing information on upgrade package costs, emissions reduction, and energy reductions by prioritized housing segment. The ResStock analysis presented here focuses on opportunities to reduce energy burden, energy consumption, and energy bills for single family homes, multifamily buildings, and mobile homes.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

STRUCTURAL MODELING TO SUPPORT POST-YIELD ACCEPTANCE CRITERIA FOR SPENT NUCLEAR FUEL CLADDING

Spent nuclear fuel (SNF) is evaluated for structural failure during storage and transportation scenarios. The U.S. Department of Energy’s Spent Fuel and Waste Science and Technology (SFWST) program has sponsored significant research in quantifying mechanical loads on SNF during storage and transportation scenarios using experimental and modeling methods. The SFWST program has also performed significant research on measuring the mechanical behavior of irradiated SNF as defueled cladding segments and cladding with fuel pellets to measure composite behavior. This paper considers some of the key material data from the Sibling Pin testing and uses structural modeling and analysis methods that have been informed by testing to consider post-yield acceptance criteria for SNF cladding structural analysis. Test data published by Oak Ridge National Laboratory (ORNL) and Pacific Northwest National Laboratory (PNNL) are the foundation for informing the material behavior of the models developed in this study. In particular, four-point bend (4PB) tests of fueled and defueled cladding segments provide significant information about the bending failure mode of SNF. ORNL’s 4PB test data is on fueled cladding segments, so the composite behavior of SNF is demonstrated. This paper describes PNNL’s coincident beam model that was developed to approximate the composite behavior of SNF. This paper also presents PNNL’s structural dynamic finite element models of a cask tip-over scenario, which is predicted to cause the strongest mechanical loads on SNF of all postulated storage and transportation scenarios. SNF bending loads predicted in the cask tip-over scenario and cladding acceptance criteria beyond yield are considered, with justification based on the Sibling Pin test data. ASME Boiler and Pressure Vessel code stress intensity limits are also considered. The ultimate goal of this work is to aid in the justification of structural acceptance criteria for SNF cladding beyond the cladding’s irradiated yield strength for use in structural analysis of all storage and transportation scenarios.

Klymyshyn, Nicholas A.↗

SAGIPS: A scalable Framework for scidac quantom

As part of the Scientific Discovery through Advanced Computing (SciDAC) program, the Quantum Chromodynamics Nuclear Tomography (QuantOM) project aims to analyze data from Deep Inelastic Scattering (DIS) experiments conducted at Thomas Jefferson National Accelerator Facility and the upcoming Electron Ion Collider. The DIS data analysis is performed on an event level by taking into leveraging nuclear theory models and accounting for experimental conditions. In order to efficiently run multiple analyses under varying conditions, a composable workflow was designed where each section (theory, experiment, objective minimization, etc.) has its own dedicated module. This presentation gives an overview over of the current status of this workflow, highlights present and future challenges, and highlights possible extensions to other projects with similar requirements.

Lersch, Daniel [Thomas Jefferson National Accelera↗

Synergizing human expertise and AI efficiency with language model for microscopy operation and automated experiment design

With the advent of large language models (LLMs), in both the open source and proprietary domains, attention is turning to how to exploit such artificial intelligence (AI) systems in assisting complex scientific tasks, such as material synthesis, characterization, analysis and discovery. Here, we explore the utility of LLMs, particularly ChatGPT4, in combination with application program interfaces (APIs) in tasks of experimental design, programming workflows, and data analysis in scanning probe microscopy, using both in-house developed APIs and APIs given by a commercial vendor for instrument control. We find that the LLM can be especially useful in converting ideations of experimental workflows to executable code on microscope APIs. Beyond code generation, we find that the GPT4 is capable of analyzing microscopy images in a generic sense. At the same time, we find that GPT4 suffers from an inability to extend beyond basic analyses for more in-depth technical experimental design. We argue that an LLM specifically fine-tuned for individual scientific domains can potentially be a better language interface for converting scientific ideations from human experts to executable workflows. Such a synergy between human expertise and LLM efficiency in experimentation can open new doors for accelerating scientific research, enabling effective experimental protocols sharing in the scientific community.

97 MATHEMATICS AND COMPUTING↗

DiffLense: a conditional diffusion model for super-resolution of gravitational lensing data

Abstract Gravitational lensing data is frequently collected at low resolution due to instrumental limitations and observing conditions. Machine learning-based super-resolution techniques offer a method to enhance the resolution of these images, enabling more precise measurements of lensing effects and a better understanding of the matter distribution in the lensing system. This enhancement can significantly improve our knowledge of the distribution of mass within the lensing galaxy and its environment, as well as the properties of the background source being lensed. Traditional super-resolution techniques typically learn a mapping function from lower-resolution to higher-resolution samples. However, these methods are often constrained by their dependence on optimizing a fixed distance function, which can result in the loss of intricate details crucial for astrophysical analysis. In this work, we introduce DiffLense , a novel super-resolution pipeline based on a conditional diffusion model specifically designed to enhance the resolution of gravitational lensing images obtained from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). Our approach adopts a generative model, leveraging the detailed structural information present in Hubble space telescope (HST) counterparts. The diffusion model, trained to generate HST data, is conditioned on HSC data pre-processed with denoising techniques and thresholding to significantly reduce noise and background interference. This process leads to a more distinct and less overlapping conditional distribution during the model’s training phase. We demonstrate that DiffLense outperforms existing state-of-the-art single-image super-resolution techniques, particularly in retaining the fine details necessary for astrophysical analyses.

Computer Science↗