Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data processing methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Complexity of many-body interactions in transition metals via machine-learned force fields from the TM23 data set

Abstract This work examines challenges associated with the accuracy of machine-learned force fields (MLFFs) for bulk solid and liquid phases ofd-block elements. In exhaustive detail, we contrast the performance of force, energy, and stress predictions across the transition metals for two leading MLFF models: a kernel-based atomic cluster expansion method implemented using sparse Gaussian processes (FLARE), and an equivariant message-passing neural network (NequIP). Early transition metals present higher relative errors and are more difficult to learn relative to late platinum- and coinage-group elements, and this trend persists across model architectures. Trends in complexity of interatomic interactions for different metals are revealed via comparison of the performance of representations with different many-body order and angular resolution. Using arguments based on perturbation theory on the occupied and unoccupieddstates near the Fermi level, we determine that the large, sharpddensity of states both above and below the Fermi level in early transition metals leads to a more complex, harder-to-learn potential energy surface for these metals. Increasing the fictitious electronic temperature (smearing) modifies the angular sensitivity of forces and makes the early transition metal forces easier to learn. This work illustrates challenges in capturing intricate properties of metallic bonding with current leading MLFFs and provides a reference data set for transition metals, aimed at benchmarking the accuracy and improving the development of emerging machine-learned approximations.

Chemistry↗

Subspace-Driven Learning for Anomaly Detection in Process Transients

Nuclear power plant (NPP) monitoring and diagnostic centers are actively investigating and implementing automated anomaly detection algorithms to help plants catch anomalies sooner, thereby preventing or reducing the duration of unexpected shutdowns. Current machine learning-based anomaly detection methods are expected to be highly effective during stable, full-power operations because NPPs typically operate as baseload power generators, meaning there are extensive operating data available from plant equipment. However, it is expected that anomaly detection methods will face significant challenges during transient conditions (i.e., when power output falls below full power) because plants only occasionally operate at these lower power levels, generating sparse transient operational data, and resulting in false alarms or missed detections. Here, to address this issue, transfer learning is used, which for this problem leverages knowledge (in the form of learned features) from stable, full-power operations to improve detection accuracy during transient conditions, even with limited data. In this effort, a novel subspace approach is developed to transfer a subset of the data features from full power operation to transients. This approach is validated through experiments using synthetic data and was found to outperform two baseline transfer learning approaches in anomaly detection performance across a range of amounts of transient data used in the training process.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Extraction of Vibration Data with Imaging

To date, the primary sensing technology used to measure the vibration response has been accelerometers and strain gages mounted directly to the structure and using either wired or, more recently, wireless telemetry. Cost issues with these sensors and the associated data acquisition systems typically limit the numbers that are deployed on in situ structures. Although there are a few structures with larger sensing counts that in some cases exceed over 1000 sensors, more typical numbers range from ten to one hundred sensors resulting in low spatial resolution when they are applied to physically large systems. When one considers that nuclear power plant structures usually have complex geometries, material properties, connectivity and boundary conditions, it is clear these current approaches to vibration measurements can only provide limited information about a system’s dynamics response characteristics. As an alternative, many non-contact measurement technologies have emerged, including point wise measurement methods such as Global Positioning System (GPS), microwave interferometry, and laser Doppler vibrometry (LDV), as well as simultaneous full-field measurement methods such as electronic speckle pattern interferometry, holography interferometry, and muon tomography, some of which can provide high spatial resolution measurements. Among these methods, digital video imaging techniques have emerged as a feasible solution for full-field vibration measurements that provide significantly more detailed dynamic response information because every pixel becomes a measurement point. Furthermore, recent advances in image processing and computer vision algorithms have been successfully used to process video data for experimental and operational modal analysis. Such full-field measurements have the potential to significantly improve many current structural assessment procedures including system identification (modal parameter estimation), structural health monitoring, load reconstruction, model validation, and model updating. Furthermore, more recent full-field imaging techniques can be accomplished with relatively low-cost, commercially-available off-the-shelf cameras. However, these measurement procedures have other limitations that must be considered such as the ability to only measure visibly accessible points on a structure and a more limited dynamic range and bandwidth than can be achieved with accelerometers or strain gages.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Development of a Method for Shape Optimization for a Gas Turbine Fuel Injector Design Using Metal-Additive Manufacturing

Adjoint shape optimization has enabled physics-based optimal designs for aerodynamic surfaces. Additive manufacturing (AM) makes it possible to manufacture complex shapes. However, there has been a gap between optimal and manufacturable surfaces due to the inherent limitations of commercial computational fluid dynamics (CFD) codes to implement geometric constraints during adjoint computation. In such cases, the design sensitivities are exported and used to perform constrained shape modifications using parametric information stored in computer aided design (CAD) files to satisfy manufacturability constraints. However, modifying the design using adjoint methods in CFD solvers and performing constrained shape modification in CAD can lead to inconsistencies due to different shape parameterization schemes. This paper describes a method to enable the simultaneous optimization of the fluid domain and impose AM manufacturability constraints, resolving one of the key issues of geometry definition for isogeometric analysis. Similar to a grid convergence study, the proposed method verifies the consistencies between shape parameterization techniques present within commercial CAD and CFD software during mesh movement as a part of the adjoint shape optimization routine. By identifying the appropriate parameters essential to a shape optimization study, the error metric between the different parameterization techniques converges to demonstrate sufficient consistencies for justifiable exchange of data between CAD and CFD. For the identified shape optimization parameters, the error metric to measure the deviation between the two parameterization schemes lies within the AM laser-powder bed fusion (L-PBF) process tolerance. Additionally, comparison for subsequent objective function calculations between iterations of the optimization loop showed acceptable differences within 1% variation between the modified geometries obtained using the two parameterization schemes. This method provides justification for the use of multiphysics guided adjoint design sensitivities computed in CFD software to perform shape modifications in CAD to incorporate AM manufacturability constraints during the shape optimization loop such that optimal designs are also additively manufacturable.

33 ADVANCED PROPULSION SYSTEMS↗

Challenges and Opportunities for Electric Utility Modeling and Asset Valuation Frameworks: Case Study on Valuing New Pumped Storage Hydropower

Asset valuation by electric utilities is becoming increasingly difficult in the rapidly changing electric sector. Rapid deployment of variable generation and inverter-based storage systems along with uncertain demand growth, climate, policies, and other factors create a challenging environment for understanding the value proposition of a new potential asset. This report describes an effort between the Tennessee Valley Authority (TVA) and three U.S. Department of Energy laboratories to perform a detailed review of utility modeling and analysis practices for asset valuation and identify challenges and opportunities for advancing its methods into the future. It focuses on a case study of new potential pumped storage hydropower (PSH) because of growing interest in new PSH capacity to provide energy balancing, firm capacity, and a range of ancillary services. Staff from the DOE labs conducted systematic interviews about current practices in capacity expansion modeling, production-cost modeling, hydrological modeling, and transmission stability modeling while also discussing how scenario analysis is conducted and how models and data are integrated. The effort resulted in a set of model, integration, and scenario recommendations that could be valuable to TVA, other utilities, system operators, and other stakeholders conducting integrated grid analysis. Individual model recommendations suggest exploring computational tradeoffs with detail and resolution across spatiotemporal structure, supply- and demand-side details, transmission overlays, market interactions, and ancillary services. Automated processes to pass data between models and conduct larger scenario suites could also enhance valuation practices by enabling a more consistent study of asset value across a broader range of uncertain future grid conditions where PSH could be particularly valuable. TVA and other industry stakeholders can learn from and adapt applied research-grade methods developed by DOE laboratories and other research institutions to improve decision making and accelerate progress towards a reliable, economic, sustainable energy system.

13 HYDRO ENERGY↗

Implementing Superresolution of Nonstationary Tides with Wavelets: An Introduction to CWT_Multi

Abstract Tides are often nonstationary due to nonastronomical influences. Investigating variable tidal properties implies a trade-off between separating adjacent frequencies (using long analysis windows) and resolving their time variations (short analysis windows). Previous continuous wavelet transform (CWT) tidal methods resolved tidal species. Here, we present CWT_Multi, a MATLAB code that 1) uses CWT linearity (via the “response coefficient method”) to implement superresolution, i.e., resolving tidal constituents beyond the Rayleigh criterion; 2) provides a Munk–Hasselmann constituent selection criterion appropriate for superresolution; and 3) introduces an objective, time-variable form of inference (“dynamic inference”) based on time-varying data properties. CWT_Multi resolves tidal species on time scales of days, and multiple constituents per species with fortnightly filters. It outputs astronomical phase lags and admittances, analyzes multiple records, and provides power spectra of the signal(s), residual(s), and reconstruction(s); confidence limits; and signal-to-noise ratios. Artificial data and water levels from the Lower Columbia River Estuary (LCRE) and San Francisco Bay Delta (SFBD) are used to test CWT_Multi and compare it to harmonic analysis programs NS_Tide and UTide. CWT_Multi provides superior reconstruction, detiding, dynamic analysis utility, and time resolution of constituents (but with broader confidence limits). Dynamic inference resolves closely spaced constituents (like K 1 , S 1 , and P 1 ) on fortnightly time scales, quantifying impacts of diel power peaking (with a 24-h period, like S 1 ) on water levels in the LCRE. CWT_Multi also helps quantify the impacts of high flows and a salt barrier closing on tidal properties in the SFBD. On the other hand, CWT_Multi does not excel at prediction, and results depend on analysis details, as for any method applied to nonstationary data. Significance Statement Ocean tides, especially in coastal and estuarine systems, are often nonstationary, in the sense that the mean and standard deviation of tidal properties vary over time, usually in response to some nontidal process. We introduce here a MATLAB code, CWT_Multi, that uses wavelet transforms to resolve both tidal species and constituents on time scales from a few days to months. Our code accommodates multiple scalar time series and has typical tidal analysis features like constituent selection and inference, plus two forms of uncertainty analyses. It is flexible, allowing the user to adapt analysis properties to diverse datasets. CWT_Multi is applicable to many problems involving time-variable tides, including sea level rise, compound flooding, sediment transport, and wetland habitat analyses. Application to vector data is a straightforward extension, but further development of our uncertainty analysis is merited. Because nonstationary tidal analysis is rapidly advancing, we also define the features of a “well-formed” analysis code.

Lobo, Matthew↗

Stream Chemistry, Synoptic Surveys, East Fork Poplar Creek Watershed, TN, USA; April 2023 to February 2025

Impacts of developed land cover on stream chemistry can be difficult to discern from natural variability, particularly in carbonate watersheds where weathering of urban infrastructure and lithology generate similar signatures. We evaluated how spatial patterns of stream chemistry varied across perennial and non-perennial tributaries spanning an urban-to-forested gradient in a mid-order, carbonate-dominated watershed. This data package contains a processed and compiled summary of stream chemistry and properties obtained from 12 synoptic surveys of 54 stream sites across the East Fork Poplar Creek watershed located near Oak Ridge, TN, United States. The sites include non-perennial tributaries, perennial tributaries, and the main stem and span forested to urban (highly developed) land cover gradients. The data package includes the processed and flagged chemical data (WaDE_SynopticSummary_FinalChemistry), metadata describing data flagging and analysis (WaDE_SynopticSummary_Metadata), information about each site and its contributing subcatchment (WaDE_SynopticSummary_SiteInformation), and a comparison of instrument and field detection limits used to determine method detection limits for the study (WaDE_SynopticSummary_DetectionLimitComparison). Stream chemistry includes stream parameters measured in situ using multiparameter probes (dissolved oxygen, pH, specific conductance, temperature) and solutes including nutrients (nitrate, ammonium, soluble reactive phosphorus), dissolved organic carbon, dissolved inorganic carbon, major cations (calcium, magnesium, potassium, sodium), major anions (chloride, sulfate), and a broad suite of minor and trace elements.

EARTH SCIENCE > TERRESTRIAL HYDROSPHERE > SURFACE ↗

Extended Application of State LiDAR Datasets in Locating Orphaned Wells in Appalachian Region

Location inaccuracies in historical and state oil and gas well databases present a major challenge in locating these orphaned wells. To address this, modern scientific methods such as Light Detection and Ranging (LiDAR), aerial magnetic remote sensing, and digital GIS products have been employed. LiDAR technology uses light to detect surface area changes, providing detailed surface views. This is a workflow to process LiDAR data for use in locating orphaned wells.

Gorantla, Vijaya [NETL Site Support Contractor, Na↗

Identifying preferential flow from soil moisture time series: Review of methodologies

Abstract Identifying and quantifying preferential flow (PF) through soil—the rapid movement of water through spatially distinct pathways in the subsurface—is vital to understanding how the hydrologic cycle responds to climate, land cover, and anthropogenic changes. In recent decades, methods have been developed that use measured soil moisture time series to identify PF. Because they allow for continuous monitoring and are relatively easy to implement, these methods have become an important tool for recognizing when, where, and under what conditions PF occurs. The methods seek to identify a pattern or quantification that indicates the occurrence of PF. Most commonly, the chosen signature is either (1) a nonsequential response to infiltrated water, in which soil moisture responses do not occur in order of shallowest to deepest, or (2) a velocity criterion, in which newly infiltrated water is detected at depth earlier than is possible by nonpreferential flow processes. Alternative signatures have also been developed that have certain advantages but are less commonly utilized. Choosing among these possible signatures requires attention to their pertinent characteristics, including susceptibility to errors, possible bias toward false negatives or false positives, reliance on subjective judgments, and possible requirements for additional types of data. We review 77 studies that have applied such methods to highlight important information for readers who want to identify PF from soil moisture data and to inform those who aim to develop new methods or improve existing ones. Core Ideas Soil moisture data can be used to identify the occurrence of preferential flow (PF) and its initiating conditions. Various data‐analysis methods to identify PF differ in susceptibility to error, bias, and subjectivity. These methods can utilize vast amounts of data from soil moisture monitoring networks to develop understanding of when, where, and under what conditions PF occurs. Newly developed methods may lead to better accuracy and reliability, and reduce the need for subjective judgments. Plain Language Summary Preferential flow through soil occurs when a large amount of water is suddenly available, as during an intense storm. This type of flow moves rapidly through the soil in distinct narrow pathways rather than moving evenly throughout the body of soil, with major consequences for groundwater resources, ecosystems, spreading of contaminants, and other vital concerns. Methods of detecting preferential flow have been developed that utilize measurements of soil water content made by sensors installed at various depths. This measurement technology has been widely implemented, many locations now having datasets years in length, and various methods have been developed for using these to identify preferential flow. The various methods are based on different features in the soil moisture records and vary in their advantages and shortcomings. In this review, we explain and evaluate these methods, highlighting important information for their implementation to identify preferential flow from soil moisture data and for efforts to develop new methods or improve existing ones.

Nimmo, John R↗

Cryo2StructData: A Large Labeled Cryo-EM Density Map Dataset for AI-based Modeling of Protein Structures

The advent of single-particle cryo-electron microscopy (cryo-EM) has brought forth a new era of structural biology, enabling the routine determination of large biological molecules and their complexes at atomic resolution. The high-resolution structures of biological macromolecules and their complexes significantly expedite biomedical research and drug discovery. However, automatically and accurately building atomic models from high-resolution cryo-EM density maps is still time-consuming and challenging when template-based models are unavailable. Artificial intelligence (AI) methods such as deep learning trained on limited amount of labeled cryo-EM density maps generate inaccurate atomic models. To address this issue, we created a dataset called Cryo2StructData consisting of 7,600 preprocessed cryo-EM density maps whose voxels are labelled according to their corresponding known atomic structures for training and testing AI methods to build atomic models from cryo-EM density maps. Cryo2StructData is larger than existing, publicly available datasets for training AI methods to build atomic protein structures from cryo-EM density maps. We trained and tested deep learning models on Cryo2StructData to validate its quality showing that it is ready for being used to train and test AI methods for building atomic models.

59 BASIC BIOLOGICAL SCIENCES↗

Robust error calibration for serial crystallography

Serial crystallography is an important technique with unique abilities to resolve enzymatic transition states, minimize radiation damage to sensitive metalloenzymes and perform de novo structure determination from micrometre-sized crystals. This technique requires the merging of data from thousands of crystals, making manual identification of errant crystals unfeasible. cctbx.xfel.merge uses filtering to remove problematic data. However, this process is imperfect, and data reduction must be robust to outliers. We add robustness to cctbx.xfel.merge at the step of uncertainty determination for reflection intensities. This step is a critical point for robustness because it is the first step where the data sets are considered as a whole, as opposed to individual lattices. Robustness is conferred by reformulating the error-calibration procedure to have fewer and less stringent statistical assumptions and incorporating the ability to down-weight low-quality lattices. We then apply this method to five macromolecular XFEL data sets and observe the improvements to each. The appropriateness of the intensity uncertainties is demonstrated through internal consistency. This is performed through theoretical CC 1/2 and I /σ relationships and by weighted second moments, which use Wilson's prior to connect intensity uncertainties with their expected distribution. This work presents new mathematical tools to analyze intensity statistics and demonstrates their effectiveness through the often underappreciated process of uncertainty analysis.

Mittan-Moreau, David W.↗

Event Classifications on DNE2 Main Experiment Data using a Convolutional Neural Network Ensemble

The Dynamic Networks (DN) Experiment for FY24 (DNE2) is an experiment within DN with the goal of quantitatively evaluating the effectiveness of solutions developed so far by various researchers under the Low Yield Nuclear Monitoring (LYNM) program using a shared set of metrics and datasets. A key component of this experiment is the mimicking of a signature processing pipeline, and comparing currently accepted and standard-use processing methods to more state-of-the-art processes developed under DN. In this work, we focus specifically on the Event Characterization (EC) Focus Area (FA) of the pipeline, where a seismic event’s magnitude, yield and class are identified. We use Deep Learning (DL) to classify the type of events being processed as either earthquakes (EQs) or explosions (EXs) for three iterations of experiment datasets. The model is noticeably more confident and accurate in classifying explosions than earthquakes, reflecting a known shortcoming of the model, that being of a bias towards predicting explosions over earthquakes in the west coast due to training data biases.

97 MATHEMATICS AND COMPUTING↗

Analysis of differential scanning calorimetry data for aged plutonium

Differential scanning calorimetry data for samples of a 52 year old plutonium alloy with 3.3 at. % Ga that were heated beyond the melting point is analyzed using transition state theory to find activation energies for the δ to ε and ε to liquid phase transitions. A Bayesian statistical method involving a Gaussian process model is used to find mean values and confidence intervals for the activation energies. The activation energy for the δ to ε phase transition increases by 3.3 ± 3.8% per decade, relative to the case when all age related plutonium lattice point defects have been removed through annealing. The corresponding increase in activation energy for the ε to liquid transition is shown to be 7.1 ± 1.8% per decade. It is postulated that the change in activation energy with age for both phase transitions is caused, in part, by the accumulation of the same type of lattice point defects associated with the observed increase in elastic bulk modulus over time.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Development of an ERT‐Based Framework for Bentonite Buffers Monitoring From Laboratory Tests: 1. Characterizing Thermal–Hydrological–Mechanical Processes

Abstract Bentonite clay is widely used in engineered barrier systems for the permanent disposal of high‐level radioactive waste due to its low permeability, high swelling capacity, and thermal stability. However, the complex thermal‐hydrological‐mechanical (THM) processes induced by heating from decaying radioactive waste and hydration from surrounding rock can lead to heterogeneous changes that are difficult to measure and predict. This study develops an Electrical Resistivity Tomography (ERT)‐based framework for monitoring THM processes, progressing from sample‐scale to bench‐scale tests, to inform field‐scale applications. Sample‐scale tests analyzed small bentonite samples under controlled variations in water content, temperature, and porosity to establish fundamental resistivity relationships. Bench‐scale tests involved larger bentonite columns subjected to heating (up to 200°C) and hydration under controlled pressure, simulating repository conditions. ERT measurements, complemented by X‐ray CT imaging, temperature monitoring, and tracing sensors, revealed coupled THM processes, such as hydration‐induced compression, swelling, and thermal gradients, leading to complex resistivity patterns. The results demonstrate the potential of ERT for capturing THM‐induced resistivity changes, though challenges remain in upscaling and quantitative analysis. This study evaluates laboratory test capabilities and proposes future improvements for understanding THM‐induced resistivity responses. A conceptual framework for ERT implementation in field‐scale monitoring is presented, synthesizing findings from both scales and exploring how ERT data can inform long‐term modeling and reduce prediction uncertainties. Overall, this ERT‐based framework offers a robust method for monitoring bentonite buffers, aiding in early issue detection and supporting the safe long‐term disposal of radioactive waste in geological repositories, while highlighting the need for future development. Plain Language Summary Bentonite clay is crucial in engineered barrier systems (EBS) for containing high‐level radioactive waste due to its ability to absorb water, swell, seal and remain stable under high temperatures. When bentonite absorbs water and heats up from radioactive decay, it experiences complex changes in its physical and mechanical properties. Understanding these changes is important for ensuring the long‐term safety and effectiveness of EBS. This study used Electrical Resistivity Tomography (ERT), a non‐invasive method that measures electrical conductivity to monitor these changes during laboratory experiments. The ERT data revealed significant variations in resistivity corresponding to changes in water content, temperature, and density, providing detailed spatial and temporal insights into the behavior of bentonite. These findings enhance our ability to predict the long‐term performance of bentonite barriers, ensuring the safe containment of radioactive waste. By improving our understanding of bentonite's behavior, this research supports the development of more reliable and effective barrier systems for radioactive waste disposal, protecting the environment and public health. Key Points ERT monitoring was employed to capture resistivity changes in bentonite during controlled heating and hydration experiments, providing insights into THM processes ERT data reveal significant resistivity changes correlated with water content, temperature, and mechanical effects, enhancing the understanding of THM dynamics in bentonite This study explores the potential of the framework for application in field‐scale EBS monitoring, emphasizing the need for integrating additional geophysical methods for comprehensive subsurface imaging

Chen, Hang↗

A Mössbauer Spectroscopy Investigation of Nickel‐Zinc Ferrites Synthesized by a Self‐Combustion Method for Soft Magnetic Core Applications

Soft ferrites are materials of interest for magnetic cores, as used for wireless charging transformers. Their low permeabilities, high resistivity, and magnetic polarization make them interesting for high-power electric vehicle charging and drive systems. The nickel-zinc-doped ferrites are of particular interest; however, the compositional space is quite large with respect to dopant concentrations, stoichiometric ratios and synthesis technique. Nickel-zinc spinel ferrites with varying nickel-zinc ratios prepared by a self-combustion reaction followed by heat treatment exhibit good crystallinity, and their low-temperature Mössbauer spectra show local magnetism and site occupation in agreement with materials prepared by solid-state reaction. Thus, the combustion synthesis method offers a facile tunability of compositions, which, combined with the possibility of rapid characterization of atomic-scale magnetism by Mössbauer spectroscopy, enables advances in the compositional and processing space at a fast pace. Low-temperature Mössbauer spectroscopy data for samples with increasing nickel content reveals a systematic increase in average hyperfine field (2.8 T/Ni) and decrease in average isomer shift (−0.036 mm/s/Ni) that can determine the nickel/zinc content, even in the absence of applied magnetic field data. Furthermore, a gradual evolution of color is also observed with increasing nickel content, albeit trends in color depend on sintering conditions.

Mössbauer spectroscopy↗

Thermal conductance of interfaces between titanium nitride and group IV semiconductors at high temperatures

Measuring the temperature dependence of material properties is a standard method for better understanding the microscopic origins for that property. Surprisingly, only a few experimental studies of thermal boundary conductance at high temperatures exist. This lack of high temperature data makes it difficult to evaluate competing theories for how inelastic processes contribute to thermal conductance. To address this, we report time domain thermoreflectance measurements of the thermal boundary conductance for TiN on diamond, silicon-carbide, silicon, and germanium between 120 and 1000 K. In all systems, the interface conductance increases monotonically without stagnating at higher temperatures. For TiN/SiC interfaces, G ranges from 330 to 1000 MW/m2-K, with a room temperature conductance of 750 MW/m2-K. The interface conductance for TiN/diamond ranges from 140 to 950 MW/m2-K. Notably, for all four interfacial systems, the conductance continues to increase with temperature even after all phonon modes in the vibrationally soft material are thermally excited. This observation suggests that inelastic processes are significant contributors to the thermal conductance in all four interfacial systems, regardless of whether the materials forming the interface are vibrationally similar or dissimilar. Our study fills a notable gap in the literature for how interfacial conductance evolves at high temperatures and tests burgeoning theories for the role of inelastic processes in interfacial thermal transport.

Physics↗

SRBench++: Principled Benchmarking of Symbolic Regression With Domain-Expert Interpretation

Symbolic regression searches for analytic expressions that accurately describe studied phenomena. The main promise of this approach is that it may return an interpretable model that can be insightful to users, while maintaining high accuracy. The current standard for benchmarking these algorithms is SRBench, which evaluates methods on hundreds of datasets that are a mix of real-world and simulated processes spanning multiple domains. At present, the ability of SRBench to evaluate interpretability is limited to measuring the size of expressions on real-world data, and the exactness of model forms on synthetic data. In practice, model size is only one of many factors used by subject experts to determine how interpretable a model truly is. Furthermore, SRBench does not characterize algorithm performance on specific, challenging sub-tasks of regression such as feature selection and evasion of local minima. In this work, we propose and evaluate an approach to benchmarking SR algorithms that addresses these limitations of SRBench by 1) incorporating expert evaluations of interpretability on a domain-specific task, and 2) evaluating algorithms over distinct properties of data science tasks. We evaluate 12 modern symbolic regression algorithms on these benchmarks and present an in-depth analysis of the results, discuss current challenges of symbolic regression algorithms and highlight possible improvements for the benchmark itself.

97 MATHEMATICS AND COMPUTING↗

PaleoSTeHM v1.0: a modern, scalable spatiotemporal hierarchical modeling framework for paleo-environmental data

Abstract. Geological records of past environmental change provide crucial insights into long-term climate variability, trends, non-stationarity, and nonlinear feedback mechanisms. However, reconstructing spatiotemporal fields from these records is statistically challenging due to their sparse, indirect, and noisy nature. Here, we present PaleoSTeHM, a scalable and modern framework for spatiotemporal hierarchical modeling of paleo-environmental data. This framework enables the implementation of flexible statistical models that rigorously quantify spatial and temporal variability from geological data while clearly distinguishing measurement and inferential uncertainty from process variability. We illustrate its application by reconstructing temporal and spatiotemporal paleo-sea-level changes across multiple locations. Using various modeling and analysis choices, PaleoSTeHM demonstrates the impact of different methods on inference results and computational efficiency. Our results highlight the critical role of model selection in addressing specific paleo-environmental questions, showcasing the PaleoSTeHM framework's potential to enhance the robustness and transparency of paleo-environmental reconstructions.

58 GEOSCIENCES↗