Search NASA⌕ Search

SEARCH · Search NASA

Results for “subset selections”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Dark energy survey: Modeling strategy for multiprobe cluster cosmology and validation for the full six-year dataset

Here, we introduce an updated To&Krause2021 model for joint analyses of cluster abundances and large-scale two-point correlations of weak lensing and galaxy and cluster clustering (termed CL+3×2 pt analysis) and validate that this model meets the systematic accuracy requirements of analyses with the statistical precision of the final Dark Energy Survey (DES) Year 6 (Y6) dataset. The validation program consists of two distinct approaches, (i) identification of modeling and parametrization choices and impact studies using simulated analyses with each possible model misspecification and (ii) end-to-end validation using mock catalogs from customized Cardinal simulations that incorporate realistic galaxy populations and DES-Y6-specific galaxy and cluster selection and photometric redshift modeling, which are the key observational systematics. In combination, these validation tests indicate that the model presented here meets the accuracy requirements of DES-Y6 for CL+3×2 pt based on a large list of tests for known systematics. In addition, we also validate that the model is sufficient for several other data combinations: the CL+GC subset of this data vector (excluding galaxy–galaxy lensing and cosmic shear two-point statistics) and the CL+3×2 pt+BAO+SN (combination of CL+3×2 pt with the previously published Y6 DES baryonic acoustic oscillation and Y5 supernovae data).

79 ASTRONOMY AND ASTROPHYSICS↗

Beyond traditional diagnostics: Identifying active galactic nuclei using spectral energy distribution fitting in DESI data

Active galactic nuclei (AGN) are typically identified through their distinctive X-ray or radio emissions, mid-infrared (MIR) colors, or emission lines. However, each method captures different subsets of AGN due to signal-to-noise (S/N) limitations, redshift coverage, and extinction effects, underscoring the necessity for a multiwavelength approach for comprehensive AGN samples. This study explores the effectiveness of spectral energy distribution (SED) fitting as a robust method for AGN identification. Using CIGALE optical-MIR SED fits on DESI Early Data Release galaxies, we compare SED-based AGN selection (AGNFRAC ≥ 0.1) with traditional methods including BPT diagrams, WISE colors, X-ray, and radio diagnostics. The SED fitting identifies ∼70% of narrow- and broad-line AGN and 87% of WISE-selected AGN. Incorporating high S/N WISE photometry reduces star-forming galaxy contamination from 62% to 15%. Initially, ∼50% of SED-AGN candidates are undetected by standard methods, but additional diagnostics classify ∼85% of these sources, revealing low-ionization nuclear emission-line regions and retired galaxies potentially representing evolved systems with weak AGN activity. Further spectroscopic and multiwavelength analysis will be essential to determine the true AGN nature of these sources. SED fitting provides complementary AGN identification, unifying multiwavelength AGN selections. This approach enables more complete – albeit somewhat contaminated – AGN samples, which are essential for upcoming large-scale surveys where spectroscopic diagnostics may be limited.

Seyfert↗

Michel Electron Selection with SPINE for DUNE Far Detector Simulation

Michel electrons are a valuable input for particle detector calibration due to their consistent kinetic energy distribution. This report details the evaluation of a Michel electron identification method's application to simulated data from the DUNE (Deep Underground Neutrino Experiment) far detector. This method, which relies on the neural network-based particle classification software SPINE (Scalable Particle Imaging with Neural Embeddings), was developed and calibrated using simulated data for the SBND (Short-Baseline Neutrino Detector) experiment before being applied to simulated DUNE data from a 1x2x6 subset of far detector modules.

Wilson, Dante [Colorado State U.]↗

Hierarchical Bayesian modeling for Inverse Uncertainty Quantification of system thermal-hydraulics code using critical flow experimental data

The best estimate plus uncertainty methodology in nuclear system thermal-hydraulic studies necessitates a comprehensive understanding of uncertainties in system code predictions. The forward uncertainty quantification (UQ) process involves the propagation of input uncertainties through the computational models to obtain uncertainties in the outputs. To this end, achieving an accurate estimation of input uncertainties is important, which is the focus of inverse UQ (IUQ). Traditionally, research in Bayesian IUQ within the nuclear engineering domain has largely relied on single-level Bayesian inference. While being effective for relatively small datasets, this approach encounters limitations for cases with large datasets. The use of a single-level model may prove inefficient, as the resultant posterior distributions can significantly differ when distinct subsets of data are employed. To address this issue, we employ an hierarchical Bayesian model for IUQ. Furthermore, this approach involves organizing observations into different groups based on the test conditions, thereby accommodating varying calibration parameters across these distinct groups. In this study, we developed and implemented a hierarchical Bayesian IUQ method to consider the grouping effect of critical flow measurement data from various geometries. Comparing the outcomes of IUQ under different selections of test data using hierarchical Bayesian IUQ against those obtained from single-level Bayesian IUQ, the forward propagation of hierarchical Bayesian IUQ results demonstrates a notably improved agreement with the experimental data.

42 ENGINEERING↗

Using active learning to improve quasar identification for the DESI spectra processing pipeline

The Dark Energy Spectroscopic Instrument (DESI) survey uses an automatic spectral classification pipeline to classify spectra. QuasarNET is a convolutional neural network used as part of this pipeline originally trained using data from the Baryon Oscillation Spectroscopic Survey (BOSS). In this paper we implement an active learning algorithm to optimally select spectra to use for training a new version of the QuasarNET weights file using only DESI data, with the goal of improving classification accuracy. This active learning algorithm includes a novel outlier rejection step using a Self-Organizing Map to ensure we label spectra representative of the larger quasar sample observed in DESI. We perform two iterations of the active learning pipeline, assembling a final dataset of 5600 labeled spectra, a small subset of the approximately 1.3 million quasar targets in DESI's Data Release 1. When splitting the spectra into training and validation subsets we achieve similar performance to the previously trained weights file in completeness and purity calculated on the validation dataset but do so with less than one tenth of the amount of training data. The new weights also more consistently classify objects in the same way when used on unlabeled data compared to the old weights file. In the process of improving QuasarNET's classification accuracy we discovered a systemic error in QuasarNET's redshift estimation and used our findings to improve our understanding of QuasarNET's redshifts.

Machine learning↗

Many roads to the seam: How conformational flexibility drives nonadiabatic relaxation in a prototypical tetrapyrrolic chromophore

Large and structurally flexible chromophores pose challenges for in silico modeling of photodeactivation due to the many vibrational modes that can funnel the system toward energy degeneracy. In this work, we examine how the multiple degrees of freedom in biliverdin, a prototypical tetrapyrrolic chromophore, cooperate to drive access to the S 1 /S 0 intersection seam in vacuo. We begin by mapping the ground-state potential energy surface to identify representative biliverdin conformers relevant to photoexcitation. We then use a CASSCF-based framework to map the excited-state landscape and characterize the intersection seam, identifying distinct conical-intersection types. Finally, we employ ab initio multiple spawning to resolve the dynamical pathways by which the system accesses these regions. DFT potential-energy and free-energy mappings indicate that, although several conformers are relevant, the “locked-helix” ZsZsZs conformer predominates in the ground state. The intersection seam comprises numerous geometrically distinct regions characterized by varying degrees and combinations of dihedral torsion, pyramidalization, and bond-length alternation. Yet only select regions lie within energetic reach, and moderate barriers separate them from the S 1 minimum. Nonadiabatic dynamics combined with multivariate analyses show that, despite extensive mode coupling during deactivation that guides the system toward multiple regions of the seam, a single dihedral torsion, together with bond-length alternation, predominantly drives energy degeneracy. This work offers new insight into biliverdin’s intrinsic photochemical response and underscores a general feature of flexible chromophores: many modes may participate during photorelaxation, but only a limited subset ultimately dictates seam accessibility.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Assessing the Expansion of Ground-Motion Sensing Capability in Smart Cities via Internet Fiber-Optic Infrastructure

Monitoring ground motion in smart cities can improve the public safety by providing critical insights on natural and anthropogenic hazards, for example, earthquakes, landslides, explosions, infrastructure failures, and so forth. Although seismic activity is typically measured using dedicated point sensors (e.g., geophones and accelerometers), techniques such as distributed acoustic sensing have demonstrated the utility of using fiber-optic cable to detect seismic activity over comparable distances. In this article, we present the results of a study that quantifies the expansion in an area monitored for low-amplitude ground-motion events by augmenting existing point sensors with the internet fiber-optic cable infrastructure. Here we begin by describing our methodology, which utilizes geospatial data on point sensors and internet optical fiber deployed in metropolitan statistical areas (MSAs) in the United States. We extend these data to identify the area that can be monitored by (1) considering the observed seismic noise data in target locations, (2) applying the model from Wilson et al. (2021) to understand the potential coverage area gains using optical fiber sensing, and (3) optimizing the selection of fiber segments to maximize coverage and minimize deployment costs. We implement our methodology in ArcGIS to assess the additional area that can be monitored for low-amplitude ground-motion events (i.e., magnitude >0.5) by utilizing internet fiber-optic cables in the 100 most populous MSAs in the United States. We find that the addition of internet fiber-based sensors in MSAs would increase the area monitored on average by over an order of magnitude from 1% to 12%, if the subset of fiber cable segments that maximize coverage and minimize deployment costs is chosen even if only 20% of all fibers are used.

58 GEOSCIENCES↗

Inhibition of MALT1 and BCL2 Induces Synergistic Antitumor Activity in Models of B-Cell Lymphoma

The activated B cell (ABC) subset of diffuse large B-cell lymphoma (DLBCL) is characterized by chronic B-cell receptor signaling and associated with poor outcomes when treated with standard therapy. In ABC-DLBCL, MALT1 is a core enzyme that is constitutively activated by stimulation of the B-cell receptor or gain-of-function mutations in upstream components of the signaling pathway, making it an attractive therapeutic target. We discovered a novel small-molecule inhibitor, ABBV-MALT1, that potently shuts down B-cell signaling selectively in ABC-DLBCL preclinical models leading to potent cell growth and xenograft inhibition. We also identified a rational combination partner for ABBV-MALT1 in the BCL2 inhibitor, venetoclax, which when combined significantly synergizes to elicit deep and durable responses in preclinical models. This work highlights the potential of ABBV-MALT1 monotherapy and combination with venetoclax as effective treatment options for patients with ABC-DLBCL.

59 BASIC BIOLOGICAL SCIENCES↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic prediction of regional-scale performance in switchgrass ( Panicum virgatum ) by accounting for genotype-by-environment variation and yield surrogate traits

Switchgrass is a potential crop for bioenergy or carbon capture schemes, but further yield improvements through selective breeding are needed to encourage commercialization. To identify promising switchgrass germplasm for future breeding efforts, we conducted multisite and multitrait genomic prediction with a diversity panel of 630 genotypes from 4 switchgrass subpopulations (Gulf, Midwest, Coastal, and Texas), which were measured for spaced plant biomass yield across 10 sites. Our study focused on the use of genomic prediction to share information among traits and environments. Specifically, we evaluated the predictive ability of cross-validation (CV) schemes using only genetic data and the training set (cross-validation 1: CV1), a subset of the sites (cross-validation 2: CV2), and/or with 2 yield surrogates (flowering time and fall plant height). We found that genotype-by-environment interactions were largely due to the north–south distribution of sites. The genetic correlations between the yield surrogates and the biomass yield were generally positive (mean height r = 0.85; mean flowering time r = 0.45) and did not vary due to subpopulation or growing region (North, Middle, or South). Genomic prediction models had CV predictive abilities of –0.02 for individuals using only genetic data (CV1), but 0.55, 0.69, 0.76, 0.81, and 0.84 for individuals with biomass performance data from 1, 2, 3, 4, and 5 sites included in the training data (CV2), respectively. To simulate a resource-limited breeding program, we determined the predictive ability of models provided with the following: 1 site observation of flowering time (0.39); 1 site observation of flowering time and fall height (0.51); 1 site observation of fall height (0.52); 1 site observation of biomass (0.55); and 5 site observations of biomass yield (0.84). The ability to share information at a regional scale is very encouraging, but further research is required to accurately translate spaced plant biomass to commercial-scale sward biomass performance.

09 BIOMASS FUELS↗

Verification of RESRAD-OFFSITE Code (V.4)

This report documents the verification of RESRAD-OFFSITE Version 4.0 and describes, where necessary, the verification of the following: • The data comprising the standard dose and risk coefficient libraries in the RESRAD database files Master_dcf_ICRP07.mdb and Master_dcf_2k.mdb. • The extraction and transfer of the data from the selected database file to the computational code by the RESRAD-OFFSITE 4.0 interface, ResOWin.exe. • The different processes that are modeled by the main computational code in RESRAD OFFSITE 4.0, ResOMain.exe. • The data displayed in the graphical and text reports. Many verifications were performed as part of the quality assurance quality control program associated with the development and release of RESRAD-OFFSITE 4.0, namely: • developer testing, • internal independent testing, and • release testing. Some were also performed in response to questions from users regarding the performance of the code. The main text of the report focuses on summarizing a subset of those tests, both independent and developer tests that verified the computations performed by the code. The verifications included in this report served as the basis for the development of the release tests of the computational executables and provided the quantitative results to be compared with the code output. The input and output interfaces and the data transfers between the various executables of the code were tested while performing the verification testing. They were tested intentionally during release testing. This report also provides some basic information to help in understanding the activities that were verified. The report: • outlines the components of RESRAD-OFFSITE 4.0 and the interconnections between these components, • outlines the processes modeled by the computational code, • provides summary figures and tables to offer confirmation of the verification of the computational components of the code, • reproduces the verifiers’ reports, if available, in individual appendices, • refers to the previous verification report (Yu et al. 2011) for more details about some of the verifications, and • reproduces the test cases and the testers’ reports from the release testing in individual appendices, when possible.

54 ENVIRONMENTAL SCIENCES↗

Design and Validation of a High-Throughput Reductive Catalytic Fractionation Method

Reductive catalytic fractionation (RCF) is a promising method to extract and depolymerize lignin from biomass, and bench-scale studies have enabled considerable progress in the past decade. RCF experiments are typically conducted in pressurized batch reactors with volumes ranging between 50 and 1000 mL, limiting the throughput of these experiments to one to six reactions per day for an individual researcher. Here, we report a high-throughput RCF (HTP-RCF) method in which batch RCF reactions are conducted in 1 mL wells machined directly into Hastelloy reactor plates. The plate reactors can seal high pressures produced by organic solvents by vertically stacking multiple reactor plates, leading to a compact and modular system capable of performing 240 reactions per experiment. Using this setup, we screened solvent mixtures and catalyst loadings for hydrogen-free RCF using 50 mg poplar and 0.5 mL reaction solvent. The system of 1:1 isopropanol/methanol showed optimal monomer yields and selectivity to 4-propyl substituted monomers, and validation reactions using 75 mL batch reactors produced identical monomer yields. To accommodate the low material loadings, we then developed a workup procedure for parallel filtration, washing, and drying of samples and a 1H nuclear magnetic resonance spectroscopy method to measure the RCF oil yield without performing liquid-liquid extraction. As a demonstration of this experimental pipeline, 50 unique switchgrass samples were screened in RCF reactions in the HTP-RCF system, revealing a wide range of monomer yields (21-36%), S/G ratios (0.41-0.93), and oil yields (40-75%). These results were successfully validated by repeating RCF reactions in 75 mL batch reactors for a subset of samples. We anticipate that this approach can be used to rapidly screen substrates, catalysts, and reaction conditions in high-pressure batch reactions with higher throughput than standard batch reactors.

BIOMASS FUELS,INORGANIC, ORGANIC, PHYSICAL, AND AN↗

Post-fire time series photos from five sites across the Oak Creek watershed, Washington

This dataset supports a broader study examining wildfire impacts on hydrologic connectivity across 5 sites within the Oak Creek watershed and the resulting biogeochemical impacts. Sites were selected using the Advanced Terrestrial Simulator (ATS) hydrologic model to identify locations with varying groundwater contributions and hydrologic responses across different burn severity scenarios. The Retreat Fire burned from July 23 to August 2, 2024, affecting all five sites. This dataset provides time series game camera photos, while the broader study includes continuous water quality monitoring, biogeochemical sampling of water and soils, precipitation data, and organic matter analysis. The other data types and additional metadata (include site environmental information) can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3018020. Because this study is ongoing, this data package will be updated regularly to include newly collected photos. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and (6) folders of game camera photos. The game camera photos are organized by site with subfolders by month of collection. The field metadata contains a subset of the information collected that is most relevant to photo-processing. The full set of field metadata can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3018020. All files are .csv, .pdf, or .jpg.

Burn severity↗

Diffusion and Solvation Dynamics of Ions in Water: Beyond the Brownian Approximation

The coupled dynamics of ions and water molecules in their first hydration shell impact a variety of processes including ion diffusion, selective ion transport in water-filled nanopores, and the kinetics of ion-pairing, ion adsorption, and metal-ligand binding reactions. In this work, we study these coupled dynamics for alkali metals (Li, Na, K, Rb, Cs), alkaline Earth metals (Mg, Ca, Sr, Ba), and chloride through the lens of their dependence on ion isotopic mass. Results are validated against previous measurements of the isotopic mass-dependence of ion diffusion coefficients in water and previous ab initio calculations of ion high-frequency dynamics in water. We find that the vibrational power spectra of ions in water consistently exhibit either two or three peaks, i.e., ions have several rattling frequencies within their solvations shells as previously reported for a subset of the species examined here. These frequencies have different sensititivies to isotopic mass that may serve as signatures of ion solvation processes (such as the tendency of ions to orient their first-shell water molecules) and that also may relate to Hofmeister-like effects including the relative affinity of different metals for ribonucleic acid (RNA).

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SARS-CoV-2 variant nanobodies and constructs comprising such nanobodies

A large and highly diverse nanobody library was constructed and screened against multiple variants of SARS-COV-2 to find nanobodies with high sensitivity and specificity for the variants. Four rounds of positive selection against a panel of six diverse SARS-COV-2 variant RBDs was performed with our high-diversity. At least 59 of these nanobodies were found to work well against Alpha, Beta, Gamma, Delta, Kappa, Lambda and Mu with some overlap efficacy against other variants. These nanobodies have efficacy as stand-alone nanobodies and as a construct comprising nanobodies linked to the human IgG1 constant fragment (Fc) (nanobody-hFc constructions or nb-hFcs) to make enhanced humanized sdAbs with all the attributes of nanobodies with improved half-life and optimized effector functions. Several promising nanobodies that neutralize the original SARS-COV-2 and several of its variants have been identified, including Delta, with high efficacy. In particular, a subset of these nanobodies bind to the Omicron RBD.

Harmon, Brooke Nicole↗

Informed total-error-minimizing priors: Interpretable cosmological parameter constraints despite complex nuisance effects

While Bayesian inference techniques are standard in cosmological analyses, it is common to interpret resulting parameter constraints with a frequentist intuition. This intuition can fail, for example, when marginalizing high-dimensional parameter spaces onto subsets of parameters, because of what has come to be known as projection effects or prior volume effects. We present the method of informed total-error-minimizing (ITEM) priors to address this problem. An ITEM prior is a prior distribution on a set of nuisance parameters, such as those describing astrophysical or calibration systematics, intended to enforce the validity of a frequentist interpretation of the posterior constraints derived for a set of target parameters (e.g., cosmological parameters). Our method works as follows. For a set of plausible nuisance realizations, we generate target parameter posteriors using several different candidate priors for the nuisance parameters. We reject candidate priors that do not accomplish the minimum requirements of bias (of point estimates) and coverage (of confidence regions among a set of noisy realizations of the data) for the target parameters on one or more of the plausible nuisance realizations. Of the priors that survive this cut, we select the ITEM prior as the one that minimizes the total error of the marginalized posteriors of the target parameters. As a proof of concept, we applied our method to the density split statistics measured in Dark Energy Survey Year 1 data. We demonstrate that the ITEM priors substantially reduce prior volume effects that otherwise arise and that they allow for sharpened yet robust constraints on the parameters of interest.

79 ASTRONOMY AND ASTROPHYSICS↗

Cleaned 5-Minute Resolution Air Quality and Meteorological Data from Nine TCEQ CAMS Sites in Houston, Texas (Nov 2021 – Oct 2022)

These data encompass 5-minute air monitoring and meteorological observations collected in the greater Houston, Texas metropolitan region, at nine (9) Continuous Ambient Monitoring Stations (CAMS) operated by the Texas Commission on Environmental Quality (TCEQ) between November 1, 2021 and October 31, 2022. The CAMS sites (CAMS 1, 8, 35, 45, 148, 403, 405, 410, and 1052) were chosen because their instrumentation includes measurements of PM2.5. These sites also provide continuous multi-parameter air-quality and meteorological measurements. Particulate matter (PM2.5, PM10) was sampled along with several trace gases, including ozone (O3), nitrogen oxides (NO, NO2, NOx), sulfur dioxide (SO2), and carbon monoxide (CO). The data set also contains standard surface meteorological parameters (temperature, humidity, pressure, wind speed, and wind direction). Several sites also include AutoGC-based measurements of volatile organic compounds (VOCs). Air monitoring instruments deployed at the selected sites comprise the following systems: BAM-1020 or TEOM (PM2.5), Thermo Scientific TEI 49i (O3), TEI 42i (NOx), and AutoGCs (VOCs). This data set is similar to the data included within the houairq5mX1.00 datastream, except for a few additional quality control steps. A systematic data cleaning and verification process was performed on the data set to ensure its quality and preparation for analysis. Removal of non-numeric status flags (e.g., [LIM], [QAS], [SPZ], [CAL], [PMA], [AQI], [SPN], [MAL]) was accomplished by employing rule-based string parsing to extract valid numerical values. Missing entries were set to -9999; however, invalid or anomalous values (e.g., 99999) were retained as originally reported by the TCEQ to preserve data provenance. The time sequence was verified for completeness, removal of duplicates, and uniformity at 5-minute intervals. Column labeling was standardized, and corresponding values were assessed for physical plausibility. All timestamps in the data set were reported in Coordinated Universal Time (UTC) as provided by the TCEQ. Further, the latitude and longitude coordinates were added for each CAMS site. A subset of the data (June 1–September 30, 2022) has been used in the following publication: Subba et al. 2025. “Implications of sea breeze circulations on boundary layer aerosols in the southern coastal Texas region.” EGUsphere 2025: 1–49, https://doi.org/10.5194/egusphere-2025-2659.

latitude↗

Data for Yield from Iowa’s first commercial miscanthus fields: implications of spatial variability for productivity and sustainability beyond research plots

This dataset contains biomass yield measurements and associated vegetation index data collected from commercial Miscanthus × giganteus fields in eastern Iowa during the 2022–2023 growing seasons. The data support the analyses presented in the article: “Yield From Iowa's First Commercial Miscanthus Fields: Implications of Spatial Variability for Productivity and Sustainability Beyond Research Plots.” We collected 105 ground-truth biomass samples from four mature commercial fields (>4 years old) covering 92.81 ha. Samples were taken from 3 m² quadrats that were hand-harvested in alignment with commercial harvest timing. Stem biomass (excluding leaves) was weighed, moisture-corrected, and converted to dry-matter yield expressed in Mg DM ha⁻¹. Sampling locations were selected to capture spatial variability visible in aerial imagery and were recorded using RTK GPS. Each biomass observation was paired with vegetation indices derived from high-resolution PlanetScope satellite imagery (3 m resolution). Images were acquired throughout the growing season, and indices were calculated to evaluate their ability to predict end-of-season biomass yield. Statistical and machine learning approaches were used to identify key predictors, and a linear regression model based on end-of-July Green Normalized Difference Vegetation Index (GNDVI) was developed and evaluated. This repository includes the data used in that modeling workflow. Management practices, economic data, full imagery time series, and additional methodological details are described in the associated publication and are not included here. The dataset consists of three comma-separated value (CSV) files: 1. Combine_Groundtruth_Yield_VI_22_23.csv This file contains ground-truth biomass yield measurements and associated key vegetation index values collected during the 2022 and 2023 growing seasons. Rows: 105 observations Columns: Year — Year of observation (2022 or 2023) Field — Field location identifier Sample_number — Unique sample identifier GNDVI_End_Jul — Green Normalized Difference Vegetation Index calculated at end of July GNDVI_End_Aug — Green Normalized Difference Vegetation Index calculated at end of August NDRE_End_Aug — Normalized Difference Red Edge index calculated at end of August Biomass_Stem_Yield_MgDM/ha — Measured stem biomass yield (megagrams dry matter per hectare) 2. trainData_GNDVI.csv This file contains the subset of observations used to train the predictive relationship between July GNDVI and biomass yield. Rows: 76 observations Columns: Unnamed: 0 — Row index retained from the original data processing workflow GNDVI_End_Jul — GNDVI at end of July Stem_Yield_MgDM/ha — Observed stem biomass yield (Mg DM ha⁻¹) 3. testData_GNDVI.csv This file contains the test dataset used to evaluate model performance. Rows: 29 observations Columns: Unnamed: 0 — Row index retained from the original data processing workflow GNDVI_End_Jul — GNDVI at end of July Predicted_Yield_MgDM/ha — Model-predicted stem biomass yield (Mg DM ha⁻¹) Observed_Yield_MgDM/ha — Measured stem biomass yield (Mg DM ha⁻¹)

Potential yield, yield gap, in-field management, y↗