Search NASA⌕ Search

SEARCH · Search NASA

Results for “data quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Standard Operating Procedure for Optimal Deployment of Meteorological Instrumentation Within the Solar Radiation Research Laboratory: 2024 Edition

The objective of the National Renewable Energy Laboratory's (NREL's) Solar Radiation Research Laboratory (SRRL) is to collect and use high-quality solar radiation data sets for research leading to the widespread adoption of solar technologies. To appropriately populate and track the diverse array of instruments at the NREL-SRRL, NREL has established a Standard Operating Procedure (SOP) for optimal instrument deployment within the SRRL for both the Baseline Measurement System (BMS) and the Research Measurement System (RMS). Using best practices methodologies, the NREL-SRRL maintains a varied and extensive array of solar monitoring equipment to test, evaluate, and characterize the solar sensors used by federal and international agencies as well as the solar industry to determine the solar resource. The SOP provides the industry with guidance for solar resource assessment and is used for procedures in the long-term continuous monitoring of legacy instruments alongside state-of-the-art instruments. Based on the SOP, instruments are annually evaluated for continued deployment. Instruments that do not meet the SOP criteria are decommissioned, and new instruments that meet the criteria are deployed. Streamlining and optimizing the use of this facility ensures that the lab continues to be a world-leading solar calibration and measurement facility. This 2024 edition includes updates to the appendices to reflect the instrument changes from one year to another.

14 SOLAR ENERGY↗

Report on branching ratios for 111 Ag

Our understanding on the distribution of fragment masses following fission, or fission yields is largely impacted by the quality of nuclear data. One of the most straightforward and reliable ways to determine the number of fissions that occurred in a chain reaction is done via detection of the characteristic γ-rays emitted during the β decay of the fission product. These γ rays are emitted in only a fraction of the decays, and this fraction (the γ-ray intensities) must be known accurately to determine the total number of fissions. Many long-lived fission products, such as 111 Ag, play an important role in science-based stockpile stewardship and nuclear forensics. The γ-ray intensities from the decay of 111 Ag are known to only 5%, leading to a 5% uncertainty in fission-chain yield. This work aims to improve the precision of these γ-ray intensities for the most intense emissions following the decay of 111 Ag, as seen in Figure 1.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Standard Operating Procedure for Optimal Deployment of Meteorological Instrumentation Within the Solar Radiation Research Laboratory: 2025 Edition

The objective of the National Renewable Energy Laboratory's (NREL's) Solar Radiation Research Laboratory (SRRL) is to collect and use high-quality solar radiation data sets for research leading to the widespread adoption of solar technologies. To appropriately populate and track the diverse array of instruments at the NREL-SRRL, NREL has established a Standard Operating Procedure (SOP) for optimal instrument deployment within the SRRL for both the Baseline Measurement System (BMS) and the Research Measurement System (RMS). Using best practices methodologies, the NREL-SRRL maintains a varied and extensive array of solar monitoring equipment to test, evaluate, and characterize the solar sensors used by federal and international agencies as well as the solar industry to determine the solar resource. The SOP provides the industry with guidance for solar resource assessment and is used for procedures in the long-term continuous monitoring of legacy instruments alongside state-of-the-art instruments. Based on the SOP, instruments are annually evaluated for continued deployment. Instruments that do not meet the SOP criteria are decommissioned, and new instruments that meet the criteria are deployed. Streamlining and optimizing the use of this facility ensures that the lab continues to be a world-leading solar calibration and measurement facility. This 2025 edition includes updates to the appendices to describe the current instrumentation of the NREL-SRRL.

14 SOLAR ENERGY↗

North Slope of Alaska XSAPR b1 Data Processing Report: April 2024-April 2025

The North Slope of Alaska (NSA) atmospheric observatory, operated by the U.S. Department of Energy (DOE)’s Atmospheric Radiation Measurement (ARM) User Facility, is a measurement site in the Arctic that has been collecting crucial atmospheric data for more than 25 years. The central facility located in Utqiaġvik, Alaska (formerly known as Barrow) hosts a suite of instruments that are used to better understand arctic processes, which are often not well represented in earth system models. The NSA site sits only a few kilometers from the Arctic Ocean, which also makes it a prime location to study complex ocean-atmosphere-ice interactions. Arctic cloud and precipitation processes are also of scientific interest, and remote-sensing instruments including radars are a key component of the NSA instrument suite. One of the radars at NSA is the X-band Scanning ARM Precipitation Radar (XSAPR). This report evaluates one year of recent XSAPR data from April 2024 through April 2025 and details the process of generating b1-level data. This analysis marks the first effort by ARM staff to quality-control NSA XSAPR data with the goal of routinely producing b1-level data in the future depending on radar operations.

54 ENVIRONMENTAL SCIENCES↗

Performance of synthetic DAS as a function of array geometry

Distributed Acoustic Sensing (DAS) can record acoustic wavefields at high sampling rates and with dense spatial resolution difficult to achieve with seismometers. Using optical scattering induced by cable deformation, DAS can record strain fields with spatial resolution of a few meters. However, many experiments utilizing DAS have relied on unused, dark telecommunication fibers. As a result, the geophysical community has not fully explored DAS survey parameters to characterize the ideal array design. This limits our understanding of guiding principles in array design to deploy DAS effectively and efficiently in the field. A better quantitative understanding of DAS array behavior can improve the quality of the data recorded by guiding the DAS array design. Here we use steered response functions, which account for DAS fiber’s directional sensitivity, as well as beamforming and back-projection results from forward modelling calculations to assess the performance of varying DAS array geometries to record regional and local sources. A regular heptagon DAS array demonstrated improved capabilities for recording regional sources over other polygonal arrays, with potential improvements in recording and locating local sources. These results help reveal DAS array performance as a function of geometry and can guide future DAS deployments.

58 GEOSCIENCES↗

Editorial: Resolving atmospheric flow in complex environments: recent experiments in terrain and forest canopies

The characterization of atmospheric flows in complex environments, which may include steep terrain slopes and heterogeneous vegetation and/or forest cover, is a long-standing challenge in boundary-layer meteorology. Atmospheric observations are complicated by the presence of transient, terrain-induced flow features, forest-canopy-atmosphere interactions, and atmospheric stability effects, not to mention the logistical hurdles involved with instrument deployment, data analysis, and quality control. Furthermore, challenges in atmospheric modeling arise due to numerical errors associated with complex terrain flows, as well as reliance on simplified parameterizations for unresolved processes such as turbulent mixing and land-surface or forest-canopy-atmosphere interactions. These modeling challenges are exacerbated in the so-called “gray zone,” wherein features of interest have length scales that are similar to the model grid spacing, or when the principal flow layer is smaller than the grid spacing (e.g., slope flows).

54 ENVIRONMENTAL SCIENCES↗

Halo Nuclei from Ab Initio Nuclear Theory

A realistic description of halo nuclei, characterized by low-lying breakup thresholds, requires a proper treatment of continuum effects. We have developed an ab initio approach, the No-Core Shell Model with Continuum (NCSMC), capable of describing both bound and unbound states in light nuclei in a unified way. With chiral two- and three-nucleon interactions as the only input, we can predict the structure and dynamics of halo and other light nuclei and, by comparing to available experimental data, test the quality of chiral nuclear forces. We review NCSMC calculations of weakly bound states and resonances of the exotic halo nuclei 6He, 8B, 11Be, and 15C. For the latter, we discuss its production in the capture reaction 14C(n,𝛾 )15C. We highlight the challenges of a description of 6He as a Borromean n-n-4He system. Finally, we present our calculations of excited states in 10Be exhibiting a one-neutron halo structure and a large scale No-Core Shell Model investigation of 11Li as a precursor of a full n-n-9Li NCSMC study.

Navrátil, Petr↗

EPCAPE-PT-LANL Measurements: Cloud Condensation Nuclei Counter

Coastal cities offer a unique environment for studying aerosol-cloud interactions and the effects of urban emissions on cloud properties. As part of the Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE), the Partitioning Thrust by Los Alamos National Laboratory (EPCAPE-PT-LANL) was conducted. Our campaign focused on measuring the optical and chemical properties of aerosols and their interactions within marine stratocumulus clouds in La Jolla, California. EPCAPE-PT-LANL enhances the primary goals of EPCAPE through innovative observations of vapor-phase transitions between aerosols and cloud droplets, the impact of black carbon on aerosol-cloud dynamics, and the effects of cloud processing on aerosol optical properties. Instrument: Cloud Condensation Nuclei Counter, single column (Droplet Measurements Technology) Files: data_10sec_CCNc.csv, data_10min_CCNc.csv Header: - SuperSaturation[unitless]: The level of supersaturation, expressed as a unitless percentage, at which cloud condensation nuclei (CCN) activity is measured. - NumberConcentration[/cm3]: Number concentration of particles acting as CCN at the baseline supersaturation level, measured in particles per cubic centimeter. - NumberConcentration_SS2[/cm3]: Number concentration of particles acting as CCN at a supersaturation level of 0.2%, measured in particles per cubic centimeter. - NumberConcentration_SS4[/cm3]: Number concentration of particles acting as CCN at a supersaturation level of 0.4%, measured in particles per cubic centimeter. - QualityControl_Flag[bool]: A boolean flag indicating whether the data point passed quality control checks. - CVI_Flag[bool]: A boolean flag indicating whether the Counterflow Virtual Impactor (CVI) was active (true) or inactive (false) during the measurement.

54 ENVIRONMENTAL SCIENCES↗

STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data

Error-bounded lossy compression is one of the most efficient solutions to reduce the volume of scientific data. For lossy compression, progressive decompression and random-access decompression are critical features that enable on-demand data access and flexible analysis workflows. However, these features can severely degrade compression quality and speed. To address these limitations, we propose a novel streaming compression framework that supports both progressive decompression and random-access decompression while maintaining high compression quality and speed. Our contributions are three-fold: (1) we design the first compression framework that simultaneously enables both progressive decompression and random-access decompression; (2) we introduce a hierarchical partitioning strategy to enable both streaming features, along with a hierarchical prediction mechanism that mitigates the impact of partitioning and achieves high compression quality—even comparable to state-of-the-art (SOTA) non-streaming compressor SZ3; and (3) our framework delivers high compression and decompression speed, up to 6.7 × faster than SZ3.

Wang, Daoce [University of Nebraska, Omaha]↗

Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"

This data package contains the associated data and scripts for Nagamoto, E., Ombadi, M., Ciulla, F. et al. Widespread drought-driven declines in streamflows and water quality in the Upper Colorado River Basin during 1998-2022. Commun Earth Environ 7, 734 (2026). https://doi.org/10.1038/s43247-026-03890-5. This purpose of this study was to investigate the impact of the 21st century drought on water quantity and quality at catchments throughout the Upper Colorado River Basin (UCRB). We used stream flow, water temperature, specific conductance, air temperature, precipitation, and catchment attribute data for over 200 sites in the UCRB, collected from the National Water Information System using Basin3D (Varadharajan, 2023), GAGESII (Falcone, 2010), and the Google Earth Engine. We identified years of severe drought between 1998 and 2022 using the Standardized Precipitation Evaporation Index (SPEI), then calculated the relative change percentage of the stream flow, water temperature, and specific conductance from drought versus non-drought years. We used the attribute information from GAGESII to investigate what physical traits of catchments are associated streamflow vulnerability (greater relative change) or resilience to drought. We used land cover data from the National Land Cover Database (USGS, 2024) to assess any changes to physical attributes that may not be represented in the static attributes information in GAGESII. To increase data availability, we modeled stream temperature using methods from Willard, 2023. While the study period is water years 1998 to 2022, the raw water quantity and quality data extends to 1950 and the meteorological data extends to 1980. The data and code can be downloaded via the UCRB_drought.zip. Within the zip, the files are organized as follows: - INPUTS: Contains all input data used in UCRB_Drought_Workflow.ipynb - OUTPUTS: Contains all intermediate data created from UCRB_Drought_Workflow.ipynb as well as final products including the calculated Standardized Evapotranspiration Index (SPEI) - climatic_variables: The code used to collect meteorologic data from Google Earth Engine - feature_importance: The code used for the catchment attributes analysis - preprocessing: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - pyeto: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - calculations: Code used in UCRB_Drought_Workflow_Impacts.ipynb - plotting: Code used in UCRB_Drought_Workflow_Impacts.ipynb - README.md - UCRB_Drought_Workflow_Preprocessing.ipynb: The code used to prep raw data for the analysis - UCRB_Drought_Workflow_Impact.ipynb: The code which uses the prepped raw data for analysis, and plots all figures - requirements_ucrb-drought_v2.yml: The requirements file to create a virtual environment and Jupyter Lab kernel to run the code The INPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_RAW" folder contains raw data for streamflow, water temperature, and specific conductance in a ".h5" file. The "NLCD_RAW" folder contains ".csv" files with annual land cover percentages for counties within the UCRB. The "MET_RAW" folder contains a ".csv" file with monthly meteorological data (air temperature and precipitation) for the sites in the UCRB which was obtained from code in the climatic_variables folder. The "GAGESII" folder contains ".csv" files with physical catchment attribute variables for catchments across the country. The "WT_LSTM_data" folder contains ".csv" files with calculated WT (Willard, 2023) and the associated RMSEs. The "Upper_Colorado_River_Basin_Boundary" folder contains geographic data including a shapefile for plotting in the UCRB_Drought_Workflow.ipynb. The "RESERVOIRS_RAW" folder contains ".csv" files for each reservoir in the UCRB with daily reservoir storage. There are also two files in the INPUTS folder that have combined reservoir storage data and reservoir metadata. The OUTPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_data" folder contains a folder "Water_year" with the associated cleaned data, metadata, and data availability information in ".csv" files, a folder "Median_Relchange" with the relative change comparing drought to non-drought years in ".csv" files, and a folder "Peak95_Min5_Relchange" that has ".csv" files for the relative change in peak (95th %) and minimum (5th %) variables. The "NLCD_data" folder contains the difference in land cover from the beginning to end of the study period and the percentage of the county that is within UCRB bounds can be found in Nagamoto et al (2025)). The "MET_data" folder contains separated monthly air temperature and precipitation data and the calculated PET in ".csv" files. The "SPEI_data" folder contains ".csv" files with calculated SPEI values (one restricted to the study period and the other with information from the entire MET data period). The "Paper_Tables" folder contains two ".csv" files containing site information and data availability and information about the GAGESII trait aggregated categories. The base directory includes the file “flmd.csv” for a list and description of all files and the file “dd.csv” for data dictionaries. Scripts for preprocessing, analysis, and figure generation are located in the associated GitHub repository found at [https://github.com/iNAIADS/drought-impacts/tree/develop/UCRB-drought]. UPDATE 1: Title and code file updated to match submitted manuscript 10-15-2025. UPDATE 2: Code and data files updated to match revised manuscript 3-4-2026. UPDATE 3: Code and data files updated to match revised manuscript 6-7-2026. ** NOTE: DD and FLMD have not been updated yet. UPDATE 4: Added associated Manuscript information and DD and FLMD have been updated. To cite this code, please use the following BibTeX: @misc{nagamoto2025drought, author = {Emily Nagamoto and Fabio Ciulla and Mohammad Ombadi and Jared Willard and Rosemary Carroll and Charuleka Varadharajan}, title = {Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"}, year = {2025}, doi = {10.15485/2551894}, publisher = {ESS-DIVE Repository}, url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2551894} }

54 ENVIRONMENTAL SCIENCES↗

L-PBF High-Throughput Data Pipeline Approach for Multi-modal Integration

Abstract Metal-based additive manufacturing requires active monitoring solutions for assessing part quality. Multiple sensors and data streams, however, generate large heterogeneous data sets that are impractical for manual assessment and characterization. In this work, an automated pipeline is developed that enables feature extraction from high-speed camera video and multi-modal data analysis. The framework removes the need for manual assessment through the utilization of deep learning techniques and training models in a weakly supervised paradigm. We demonstrate this pipeline’s capability over 700,000 high-speed camera frames. The pipeline successfully extracts melt pool and spatter geometries and links them to corresponding pyrometry, radiography, and processparameter information. 715 individual prints are examined to reveal melt pool areas that exceeds 0.07 mm 2 and pyrometry signal over a threshold (375 pyrometry units) were more likely to have defects. These automated processes enable massive throughput of characterization techniques.

36 MATERIALS SCIENCE↗

Coupling Remote Sensing With a Process Model for the Simulation of Rangeland Carbon Dynamics

Rangelands provide significant environmental benefits through many ecosystem services, which may include soil organic carbon (SOC) sequestration. However, quantifying SOC stocks and monitoring carbon (C) fluxes in rangelands are challenging due to the considerable spatial and temporal variability tied to rangeland C dynamics as well as limited data availability. We developed the Rangeland Carbon Tracking and Management (RCTM) system to track long-term changes in SOC and ecosystem C fluxes by leveraging remote sensing inputs and environmental variable data sets with algorithms representing terrestrial C-cycle processes. Bayesian calibration was conducted using quality-controlled C flux data sets obtained from 61 Ameriflux and NEON flux tower sites from Western and Midwestern US rangelands to parameterize the model according to dominant vegetation classes (perennial and/or annual grass, grass-shrub mixture, and grass-tree mixture). The resulting RCTM system produced higher model accuracy for estimating annual cumulative gross primary productivity (GPP) (R 2 > 0.6, RMSE <390 g C m -2 ) relative to net ecosystem exchange of CO 2 (NEE) (R 2 > 0.4, RMSE <180 g C m -2 ). Model performance in estimating rangeland C fluxes varied by season and vegetation type. The RCTM captured the spatial variability of SOC stocks with R 2 = 0.6 when validated against SOC measurements across 13 NEON sites. Model simulations indicated slightly enhanced SOC stocks for the flux tower sites during the past decade, which is mainly driven by an increase in precipitation. Future efforts to refine the RCTM system will benefit from long-term network-based monitoring of vegetation biomass, C fluxes, and SOC stocks.

54 ENVIRONMENTAL SCIENCES↗

Proceedings for the Workshop on Applied Nuclear Data Activities 2024

The Workshop for Applied Nuclear Data Activities (WANDA) is designed to increase communication among nuclear data (ND) users in multidisciplinary federal programs, ND producers, ND funders, and other ND experts. It also presents an opportunity to cross-pollinate ideas as well as introduce ND gaps identified by federal programs to ND experts and ND capabilities to the various federal ND users. WANDA 2024 included five technical sessions, three of which focused on Fusion Energy Sciences (FES)—FES Fusion Neutronics, FES Tritium Production, and FES Material Damage—and two stand-alone sessions—Isotopes and Targetry for Nuclear Data and Uncertainty Quantification. The FES sessions successfully brought new voices to the WANDA discussions, expanding the application space in which nuclear data are critical. FES programs need accurate nuclear data with realistic uncertainty quantification to properly estimate, for example, shielding, activation, tritium production, helium production, structural material integrity, and superconducting magnet operation. This includes a variety of projectile (neutrons, photons, charged particles) and target atoms. One of the action items common to all the FES sessions was a need to perform sensitivity studies to identify the prioritization of nuclear data needs. The Isotopes and Targetry session highlighted the many capabilities available to produce high-quality targets for nuclear data measurements, including 3D printing with spherical powders, combustion synthesis coupled with spin coating & electrospraying, inkjet printing, and isotopic doping. These new methods open doors for more accurate measurement, but it was also stressed that sample characterization following any method of fabrication is of the highest importance to accurately interpret nuclear data measurement results that used that sample. The Uncertainty Quantification (UQ) session was broken into two categories: nuclear data uncertainty quantification and the use of that uncertainty quantification. Thematic to the UQ session was the loss of information when going from nuclear data measurement, to evaluation, to evaluated file, and finally to neutron transport calculations. Current evaluated ND libraries typically only contain covariances, which assume that the probability distributions are Gaussian. Beyond being a simplified assumption for many evaluations, this can lead to negative values on many observables when attempting to sample the covariance. The covariance format, however, is very efficient in that a simple set of linear equations can transform uncertainty from parameters or cross sections to the application of interest. Focused collaboration is needed between nuclear data evaluators and nuclear data users to ensure that needs are being met.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Quality Assurance of Legacy Post-Irradiation Examination Data for Metallic Fuels

The U.S. DOE-NE’s Advanced Reactor Technologies (ART) Fast Reactor Program (FRP) and the NE-4 Advanced Fuels Campaign (AFC) have jointly undertaken the qualification of the legacy post-irradiation examination (PIE) data held in the Fuels Irradiation & Physics Database (FIPD), covering metallic fuel experiments conducted in the Experimental Breeder Reactor II (EBR-II) and the Fast Flux Test Facility (FFTF). In FY26 this effort reached a milestone: six major types of PIE data—contact profilometry, isotopic gamma scan, fission gas chemistry, fission gas release, laser profilometry, and neutron radiography—have been qualified for all available experiments in FIPD, and the U.S. nuclear industry can now draw on them with confidence in licensing activities for metallic fuel-based advanced fast reactors. This report documents the qualification process and the resulting status of the PIE data.

Mo, Kun↗

Data for Identifying the best high-biomass sorghum hybrids based on biomass yield potential and feedstock quality affected by nitrogen fertility management under various environments

Data were collected from agronomy fields in Urbana and Ewing, IL, during the 2022 and 2023 growing seasons. The dataset includes dry biomass yield, nitrogen, phosphorus, and potassium concentrations and removals, and chemical composition elements (cellulose, hemicellulose, lignin, and soluble fractions) for 13 high-biomass sorghum hybrids. data_sharing.xlsx contains 20 columns and 104 rows. Below is the explanation of all variables in the file: Year: 2022; 2023 Location: Urbana, IL; Ewing, IL N rate (kg-N/ha): 0; 112 Hybrid #: H1-H13 Pedigree: Pedigree for 13 hybrids Dry biomass yield (Mg/ha): Aboveground dry biomass yield N (g/kg): Nitrogen concentration in plant tissue P (g/kg): Phosphorus concentration in plant tissue K (g/kg): Potassium concentration in plant tissue N (kg/ha): Nitrogen removal by aboveground biomass P (kg/ha): Phosphorus removal by aboveground biomass K (kg/ha): Potassium removal by aboveground biomass Cellulose (g/kg): Cellulose concentration in plant tissue Hemicellulose (g/kg): Hemicellulose concentration in plant tissue Lignin (g/kg): Lignin concentration in plant tissue Soluble (g/kg): Soluble concentration in plant tissue Cellulose (Mg/ha): Cellulose content in aboveground biomass Hemicellulose (Mg/ha): Hemicellulose content in aboveground biomass Lignin (Mg/ha): Lignin content in aboveground biomass Soluble (Mg/ha): Soluble content in aboveground biomass

environmental adaptability↗

Next-Level Energy Management in Manufacturing: Facility-Level Energy Digital Twin Framework Based on Machine Learning and Automated Data Collection

This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.

digital twin↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗