MACHINE-LEARNING IMAGE RECONSTRUCTION SPEEDS UP NANOMETROLOGY DATA PROCESSING BY 500X WITHOUT SACRIFICING IMAGE QUALITY
Generic poster for MSRF ERB and ML workshop hosted by LANL in Santa Fe
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Generic poster for MSRF ERB and ML workshop hosted by LANL in Santa Fe
Abstract The next generation of searches for neutrinoless double beta decay ($$0 \nu \beta \beta $$ 0 ν β β ) are poised to answer deep questions on the nature of neutrinos and the source of the Universe’s matter–antimatter asymmetry. They will be looking for event rates of less than one event per ton of instrumented isotope per year. To claim discovery, accurate and efficient simulations of detector events that mimic$$0 \nu \beta \beta $$ 0 ν β β is critical. Traditional Monte Carlo (MC) simulations can be supplemented by machine-learning-based generative models. This work describes the performance of generative models that we designed for monolithic liquid scintillator detectors like KamLAND to produce accurate simulation data without a predefined physics model. We present their current ability to recover low-level features and perform interpolation. In the future, the results of these generative models can be used to improve event classification and background rejection by providing high-quality abundant generated data.
Applications of Implicit Neural Representations (INRs) have emerged as a promising deep learning approach for compactly representing large volumetric datasets. These models can act as surrogates for volume data, enabling efficient storage and on-demand reconstruction via model predictions. However, conventional deterministic INRs only provide value predictions without insights into the model’s prediction uncertainty or the impact of inherent noisiness in the data. This limitation can lead to unreliable data interpretation and visualization due to prediction inaccuracies in the reconstructed volume. Identifying erroneous results extracted from model-predicted data may be infeasible, as raw data may be unavailable due to its large size. To address this challenge, we introduce REV-INR, Regularized Evidential Implicit Neural Representation, which learns to predict data values accurately along with the associated coordinate-level data uncertainty and model uncertainty using only a single forward pass of the trained REV-INR during inference. By comprehensively comparing and contrasting REV-INR with existing well-established deep uncertainty estimation methods, we show that REV-INR achieves the best volume reconstruction quality with robust data (aleatoric) and model (epistemic) uncertainty estimates using the fastest inference time. Consequently, we demonstrate that REV-INR facilitates assessment of the reliability and trustworthiness of the extracted isosurfaces and volume visualization results, enabling analyses to be solely driven by model-predicted data.
The Sustainable Aviation Fuel (SAF) Grand Challenge (Langholtz, 2024 ) seeks to generate 35 billion gallons of SAF each year by 2050, with corn stover, an agricultural byproduct, playing a key role as a feedstock. This study develops an optimization framework to enhance the quality and quantity of corn stover while ensuring economic and environmental viability. Using the Decision Support System for Agrotechnology Transfer (DSSAT) crop model, we simulate the effects of cover crops on rotation yield, soil moisture balance, and nitrogen cycling across diverse climates and soils. The model outputs, including yield data and soil quality changes, inform a Mixed-Integer Linear Programming (MILP) optimization model. This model aims to maximize economic and environmental returns by incorporating production costs, direct and indirect income, and environmental incentives. The optimization model evaluates 280 agriculture management plans composed of various crop management strategies, including corn stover removal rates, cover crop adoption, and fertilization practices. It seeks to identify the optimal combination of crop and tillage decisions for each subfield, maximizing profits while enhancing soil carbon sequestration and reducing greenhouse gas emissions. Outputs include detailed subfield locations, optimal management plans, and profits per hectare and per acre, allowing for comparison with literature values on farm profits. This study provides a robust optimization framework supporting the SAF Grand Challenge by proposing economically viable and environmentally sustainable strategies for corn stover utilization. The findings highlight corn stover's potential as a sustainable feedstock for SAF production, offering practical solutions to enhance its quality and quantity while maintaining soil health. Idaho is used as a case study to demonstrate the framework's applicability and effectiveness in real-world scenarios. Langholtz, M. H., Davis, M., Hellwinckel, C., De La Torre Ugarte, D., Efroymson, R., Jacobson, R., Milbrandt, A., Coleman, A., Davis, R., Kline, K. L., Badgett, A., Curran, S., Schmidt, E., Theiss, T., Fried, J., English, B., Lambert, L., Cook, H., Field, J., ... Walker, L. (2024). 2023 Billion-Ton Report: An Assessment of U.S. Renewable Carbon Resources. https://doi.org/10.2172/2441098 DSSAT Foundation. (2025). Decision Support System for Agrotechnology Transfer (DSSAT). Retrieved from https://dssat.net/
This dataset contains stream discharge and temperature data for water years 2019 to 2025 from the East and Taylor Watersheds in Colorado, United States. This data was collected to understand hydrological processes occurring in the East River and Taylor River Watersheds, Colorado, which is part of the Lawrence Berkeley National Laboratory Watershed Function Scientific Focus Area. Data includes instantaneous observed discharge using salt dilution and acoustic doppler velocimeter techniques, raw pressure transducer downloaded data, sub-hourly temperature as well as corrected water level and associated stream discharge and mean daily values. Notes on water level corrections, rating curve development and metadata provided. A rating curve is the translation of depth to streamflow. The rating curve can be used as a quantitative measure of the “quality of the data.” Data within this dataset is formatted using ESS-DIVE’s Hydrological Monitoring Reporting Format. This data package contains (1) a zip file (Stream_Discharge_Data_WY19-WY25.zip) containing stream discharge and temperature data organized by location; (2) an InstallationMethods file (InstallationMethods.csv) describing metadata about the installation; (3) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; (4) a data dictionary (dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (5) a locations metadata file (locations.csv); (6) and a sensor metadata file (sensors.csv). All data files are in non-proprietary formats (csv, png, or pdf formats). Please contact Rosemary Carroll, Curtis Beutler, or Austin Shirley for any support in accessing the files. Update on 2023-05-12: Additional data from WYs 2021 and 2022 were added. Additionally, the dataset was converted using ESS-DIVE’s Hydrological Monitoring Reporting Format. Data files were reformatted to match reporting format guidance, new metadata files were added, and files were converted from excel to CSV. Update on 2025-05-16: Additional data from WYs 2022 (for locations not previously included), 2023, and 2024 were added. An additional descriptive PDF (WFSFA_Streamflow_Hydrograph_Disclaimer.pdf) was added. Metadata files were updated to reflect the addition of new data and locations. Update on 2026-05-18: Additional data from WY 2025 were added, including a new location Upper Trail Creek (TR-TCG2). Metadata files were updated to reflect the addition of new data.
During the development activities of SuperCam Calibration Target, target intended for one of the two first Raman instruments to be deployed on another planetary body, our group developed a laboratory instrument that could simulate to some extent the Raman capabilities of one of such instruments and could provide data with similar quality. The use of this kind of laboratory instruments has demonstrated its utility in the evaluation of potential calibration targets or anticipating the science outcome that an instrument could provide. The present work describes our laboratory setup to support SuperCam, evaluating similarities between both instruments, despite of differences in the hardware. Evaluation of data gathered by SuperCam on Mars and the availability of one replica of SuperCam’s Calibration Target allowed the comparison on the same set of targets, demonstrating how similar Signal-to-Noise Ratio (SNR) could be achieved from both instruments. The higher energy per pulse on SimulCam is compensated by a greater analytical footprint and the use of smaller collection optics. The results show how spectra obtained at representative distances of SuperCam are comparable. Operational principles are also comparable in terms of time resolution, and close in terms of spectral resolution. This similarity has allowed different science support works using SimulCam data, as well as the support to Mars detections using our setup. We provide examples of this support that will be shared with the community in different papers, as well as examples of possible operations activities that could benefit from experiments performed with SimulCam. We show how this setup can complement the two laboratory replicas in Los Alamos and Toulouse in providing support data to different experiments.
Climate change amplifies many threats to human health. Despite advances in understanding climate change dynamics and impacts, there remains a critical gap in translating scientific knowledge into equitable, and community-driven health interventions. The inaugural One Earth, One Health workshop sought to explore this gap through human-centered design exercises involving interdisciplinary researchers from climate and Earth sciences, engineering, epidemiology, microbiology, and environmental health. Although participants did not co-develop solutions with affected communities, they used stakeholder role-playing to guide ideation and lay groundwork for actionable plans. Through these methods, participants identified community needs and proposed prototype solutions to alleviate health threats exacerbated by global environmental change. Prototypes were organized around infectious diseases, extreme weather, and air quality, as illustrative themes rather than an exhaustive set of risks. Key solutions included strategies for anticipatory systems and early warning (e.g., integrating environmental signals with health data), inclusive communication and infrastructure needs for responding to extreme weather events, and integrated platforms visualizing air quality trends to support tailored, context-aware guidance beyond one-size-fits-all alerts. The workshop highlighted opportunities such as leveraging machine learning, Earth observation, and real-time surveillance to protect communities, but also noted barriers including data quality, technological redundancy, privacy, and governance challenges. Additionally, participants emphasized the need for interdisciplinary teams capable of collaborating across sectors, breaking down silos and addressing gaps in training and education. Overall, the workshop illustrates how process-driven, human-centered approaches can help surface user needs and generate testable prototype concepts, while underscoring the importance of direct community partnership for implementation.
Understanding the geophysical response near an underground explosion is crucial for generating insights into the source and emplacement conditions that produce distinct observations in monitoring scenarios occurring at greater distances. Recently, Shot A of the Low Yield Nuclear Monitoring (LYNM) Physics Experiment 1 (PE1) series was conducted at the Nevada National Security Site to provide ground truth for subsurface explosion signal models. This experiment resulted in measuring near-source ground motion at distances ranging from 70 to 1000 m/kt with a 99% success rate, yielding high-fidelity knowledge of the near-field response that can serve as benchmarks for future numerical modeling and experiment planning. However, technical challenges exist in observing near-source phenomena while safeguarding sensitive data acquisition components from the detrimental effects of ground motion in the subsurface. This report outlines tools and techniques to address challenges associated with observing near-source accelerations and within the tunnel drift of the PE1 test bed. Additionally, we describe key systems designed with both modern advancements and legacy guidance to maximize the collection of high-quality ground motion data, which may be applied to constitutive and computational models, leading to new or improved understanding of the near- and far-field signals produced by underground explosions.
We present Rubin Data Preview 1 (DP1), the first data from the National Science Foundation–Department of Energy Vera C. Rubin Observatory, comprising raw and calibrated single-epoch images, coadds, difference images, detection catalogs, and ancillary data products. DP1 is based on 1792 optical–near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera (LSSTComCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile in late 2024. DP1 covers ∼15 deg 2 distributed across seven roughly equal-sized noncontiguous fields, each independently observed in six broad photometric bands, ugrizy. The median FWHM of the point-spread function across all bands is approximately 1"14, with the sharpest images reaching about 0." 58. The 5σ point-source depths for coadded images in the deepest field, the Extended Chandra Deep Field South, are u = 24.55, g = 26.18, r = 25.96, i = 25.71, z = 25.07, and y = 23.1. Other fields are no more than 2.2 mag shallower in any band, where they have nonzero coverage. DP1 contains approximately 2.3 million distinct astrophysical objects, of which 1.6 million are extended in at least one band in coadds, and 431 solar system objects, of which 93 are new discoveries. DP1 is approximately 3.5 TB in size and is available to Vera C. Rubin Observatory data rights holders via the Rubin Science Platform, a cloud-based environment for the analysis of petascale astronomical data. While small compared to future LSST releases, its high quality and diversity of data support a broad range of early science investigations ahead of full operations in 2026.
Geological fault detection and characterization are crucial for understanding subsurface dynamics across scales. While methods for fault delineation based on either seismicity location analysis or seismic image reflector discontinuity are well-established, a systematic approach that integrates both data types remains absent. We develop a novel machine learning model that unifies seismic reflector images and seismicity location information to automatically identify geological faults and characterize their geometrical properties. The model encodes a seismic image and a seismicity location image separately, and fuses the encoded features with a spatial-channel attention fusion module to improve the learning of important features in both inputs. We design an automated strategy to generate high-quality synthetic training data and labels. To improve the realism of the seismicity location image, we include random seismicity noise and missing seismicity location associated with some of the faults. We validate the model’s efficacy and accuracy using synthetic data examples and two field data examples. Moreover, we show that fine-tuning the trained model with a small, domain-specific dataset enhances its fidelity for field data applications. The results demonstrate that integrating seismicity location and seismic images into a unified framework allows the end-to-end neural network to achieve higher fidelity and accuracy in delineating subsurface faults and their geometrical properties compared with image-only fault detection methods. Our approach offers an adaptive data-driven tool for geological fault characterization and seismic hazard mitigation, bridging the gap between seismicity location and image-based fault detection methods.
Four-dimensional scanning transmission electron microscopy (4D-STEM) enables mapping of diffraction information with nanometer-scale spatial resolution, offering detailed insight into local structure, orientation, and strain. However, as data dimensionality and sampling density increase, particularly for in situ scanning diffraction experiments (5D-STEM), robust segmentation of structurally consistent behavior across sequential measurements becomes essential for efficient and physically meaningful analysis. Here, we introduce a clustering framework that identifies crystallographically distinct domains from 4D-STEM datasets. By using local diffraction-pattern similarity as a metric, the method extracts closed contours delineating spatially contiguous regions. This approach produces cluster-averaged diffraction patterns that improve signal quality while reducing data volume by orders of magnitude, enabling rapid and accurate orientation, phase, and strain mapping. We demonstrate the applicability of this approach to in situ liquid-cell 4D-STEM data of gold nanoparticle growth. Our method provides a scalable and generalizable route for spatially coherent segmentation, data compression, and quantitative structure–strain mapping across diverse 4D-STEM modalities. The full analysis code and example workflows are publicly available to support reproducibility and reuse.
Lithium-ion battery (LiB) technology is playing a crucial role in transforming the predominantly fossil fuel-based transportation and stationary storage sectors to achieve a low-carbon economy. Rapid innovation in the LiB materials to electrode to cell design is happening to satisfy the performance, life, and safety metrics required by those myriads of applications. Lately, advanced analytics, such as machine-learning or artificial intelligence (ML/AI) techniques, are being used more frequently to aid in expedited LiB technology development, performance validation, and life prediction. The success of these techniques often relies on a large volume of well-defined and high-quality battery test data. On the other hand, most battery developers and research and development (R&D) communities are still following a classical approach to develop batteries, which is running calendar- and/or cycle-aging tests, performing reference performance tests (RPTs), and conducting post-mortem analyses periodically without paying attention to the wealth of data often not collected during the calendar or cycle life aging tests. This sparse data collection approach is time- and resource-intensive, requiring data capture and evaluation of months to years of RPT data to diagnose accurate battery state of performance, health, and safety. Even so, the underlying aging modes and mechanisms can be missed. If collected properly, battery test data during cycling or calendaring can be efficiently combined with ML/AI techniques to create powerful tools in the rapid diagnosis of battery state of performance, health, and safety along with insights into underlying aging modes and mechanisms. In this report, we discuss the importance of effective cycle-by-cycle (CBC) data collection with example case studies. Within a reasonable timeframe, RPT data are often inadequate in capturing many of the crucial battery aging dynamics, which often predominantly show up in CBC test data. Finally, we also show examples of ML/AI techniques that use CBC data in rapid diagnosis and projection of LiB state of health (SOH) to motivate the scientific community in collecting and using CBC data to facilitate expeditious technology development and validation.
Assessing genome annotation quality is crucial for downstream analyses, but current methods are inadequate for eukaryotes. We present Accuracy-Based Annotation Quality Score (ABAQS), a novel, minimal-data-driven method that comprehensively assesses annotation quality. ABAQS evaluates multiple factors, including genome completeness, gene model validity, and protein profile accuracy, outperforming other metrics like BUSCO and PSAURON. We applied ABAQS to over 2500 eukaryotic genomes and showed its robustness and effectiveness in evaluating genome annotation quality, making it a valuable tool for researchers working with genomic data. ABAQS reveals significant variation in annotation quality and highlights the importance of filtering in improving annotation quality and accuracy.
Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.
The data were collected as part of the BSEC project, during the period from June 2025 to May 2026. Directory "broadway" contains data collected on a multi-level flux tower (US-BWf) in the Broadway East neighborhood (1808 North Patterson Park Ave., Baltimore City, MD 21213; LAT: 39o18'40.31'' N; LONG: 76o35'12.43'' W). At each of the four measurement heights (8.5 m, 11.1 m, 13.4 m, 15.9 m), a Campbell Scientific CSAT3B sonic anemometer was operated at 50 Hz to measure virtual temperature (tc) and three velocity components (u: 270 degrees; v: 180 degrees; w: vertical), and a RM Young temperature sensor (model 41382VC) was operated at 1 Hz inside a compact aspirated radiation shield (model 43502) to measure absolute temperature (T) and relative humidity (RH). Inside directory "broadway", directory "netcdf" contains data collected each day in 5-minute chunks that have been converted to NetCDF format (before quality checking), while "4hr" contains data arranged into 4-hour chunks (also in NetCDF format) that have been through basic quality checking steps (treating data points with nonzero diagnostic codes as missing data; fixing six or fewer consecutive missing data points using linear interpolation). Users are recommended to start with data in directory "4hr", while data in directory "netcdf" can be used for reference purposes.
As widespread adoption of photovoltaic (PV) technologies continues, understanding the lifetime of modules is paramount to the viability of the industry as an environmentally conscious alternative to traditional energy generation. Although power degradation can affect the total energy production of a module over its lifetime, module safety failures necessitate the removal of a module leading to a loss of not only the particular asset, but the earning potential of the device. Therefore, it is critical to ensure that the components that provide essential safety functions for PV module operate for their entire rated lifetime. PV backsheets provide necessary electrical insulation to the completed device and failure of this component is cause for a immediate removal of the module. Degradation of the PV module backsheet has led to module safety failures in large-scale installations, costing millions of dollars in damages and lost potential revenue. The spatio-temporal degradation of fielded PV modules is important to study in order to identify which modules within installations are experiencing the greatest exposure conditions and in turn have the highest chance of failure. This paper describes a comprehensive field survey protocol developed for monitoring PV module backsheet performance using solely non-destructive methods in commercial PV fields. The protocol establishes a field naming convention, sampling method, data handling requirements, and measurement procedures. By ensuring consistent data collection practices, the field survey protocol enables research groups to obtain data of uniform quality on backsheet performance over multiple years and locations. In this study, the developed protocol was implemented at forty-one PV sites. Eight different types of airside layer backsheet materials including poly(vinylidene fluoride) (PVDF), acrylic PVDF, poly(tetrafluoroethylene-co-hexafluoropropylene-co-vinylidene fluoride) (THV), poly(vinyl fluoride) (PVF), poly(ethylene terephthalate) (PET), fluoroethylene vinyl ether (FEVE), polyethylene naphthalate (PEN), and glass were identified using attenuated total reflection Fourier transform infrared (ATR-FTIR) spectroscopy. The field survey results show that the spatial distribution of degradation indicators are non-uniform within a particular module, individual site, and across site locations. The degradation of PV modules increased in severity for modules mounted at the edge of rows (across a field) and near the junction box (within a module). This study demonstrates the sensitivity of material performance to exposure length across different materials and climates.
The voluntary carbon market within the United States has expanded rapidly in recent years and enabled private companies and other organizations to provide revenue streams to carbon dioxide removal (CDR) technologies. For a CDR technology to participate in the voluntary carbon market (VCM), the emissions associated with constructing and operating the technology must be less than the CO 2 captured from the atmosphere. Assessing the extent to which this is true for direct air capture with storage (DACS), a relatively energy-intensive CDR technology, strongly depends on the accounting method used to assess the emissions intensity of purchased energy. We simulate the hourly weather-dependent operation of sorbent- and solvent-based DACS in California, Louisiana, Texas, and Wyoming, representing a wide range of local weather and electric and natural gas grid compositions. In all cases, the single most important emissions accounting decision is the method used to estimate the emissions intensity of purchased grid electricity, which varies the calculated net removal by −1049% to +108%. All other factors influencing net removal introduce a variation of at most ±14%. No electricity emissions accounting method is universally conservative across all scenarios, and none is objectively more accurate. High-spatiotemporal-resolution, high-quality, publicly available data sets and models for electricity emissions accounting do not currently exist and are urgently needed to enable standardization of emissions accounting methods to more accurately determine the true emissions impacts of DACS and other energy-intensive facilities.
powersqueeze (psqz) is a truncated power iteration library intended for high-performance computing platforms. psqz efficiently produces low-dimensional, linear measurements of graph matrix spectra by combining classical power iteration with sparse Johnson-Lindenstrauss transforms. psqz is intended to produce high-quality, fast, data-oblivious low-dimensional representations of high-dimensional sparse data such as graphs and term-document matrices. psqz is intended to replace similar workflows that depend on directly approximating a truncated eigendecomposition (e.g., the first step of spectral clustering), which is a much more expensive operation.