Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Streamlined spatial and environmental expression signatures characterize the minimalist duckweed Wolffia australiana

Single-cell genomics permits a new resolution in the examination of molecular and cellular dynamics, allowing global, parallel assessments of cell types and cellular behaviors through development and in response to environmental circumstances, such as interaction with water and the light–dark cycle of the Earth. Here, we leverage the smallest, and possibly most structurally reduced, plant, the semiaquaticWolffia australiana, to understand dynamics of cell expression in these contexts at the whole-plant level. We examined single-cell-resolution RNA-sequencing data and foundWolffiacells divide into four principal clusters representing the above- and below-water-situated parenchyma and epidermis. Although these tissues share transcriptomic similarity with model plants, they display distinct adaptations thatWolffiahas made for the aquatic environment. Within this broad classification, discrete subspecializations are evident, with select cells showing unique transcriptomic signatures associated with developmental maturation and specialized physiologies. Assessing this simplified biological system temporally at two key time-of-day (TOD) transitions, we identify additional TOD-responsive genes previously overlooked in whole-plant transcriptomic approaches and demonstrate that the core circadian clock machinery and its downstream responses can vary in cell-specific manners, even in this simplified system. Distinctions between cell types and their responses to submergence and/or TOD are driven by expression changes of unexpectedly few genes, characterizingWolffiaas a highly streamlined organism with the majority of genes dedicated to fundamental cellular processes.Wolffiaprovides a unique opportunity to apply reductionist biology to elucidate signaling functions at the organismal level, for which this work provides a powerful resource.

Biochemistry & Molecular Biology↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Early photometric and spectroscopic observations of the extraordinarily bright INTEGRAL-detected GRB 221009A

Context. GRB 221009A, initially detected as an X-ray transient by Swift, was later revealed to have triggered the Fermi satellite about an hour earlier, marking it as a post-peak observation of the event’s emission. This GRB distinguished itself as the brightest ever recorded, presenting an unparalleled opportunity to probe the complexities of GRB physics. The unprecedented brightness, however, challenged observation efforts, as it led to the saturation of several high-energy instruments.Aims. Our study seeks to investigate the nature of the INTEGRAL-detected GRB 221009A and elucidate the environmental conditions conducive to these exceptionally powerful bursts. Moreover, we aim to understand the fundamental physics illuminated by the detection of teraelectronvolt (TeV) photons emitted by GRB 221009A.Methods. We conducted detailed analyses of early photometric and spectroscopic observations that span from the Fermi trigger through to the initial days following the prompt emission phase in order to characterize GRB 221009A’s afterglow, and we complemented these analyses with a comparative study.Results. Our findings from analyzing INTEGRAL data confirm GRB 221009A as the most energetic event observed to date. Early optical observations during the prompt phase negate the presence of bright optical emissions with internal or external shock origins. Spectroscopic analyses enabled us to measure GRB 221009A’s distance and line-of-sight properties. The afterglow’s temporal and spectral analysis suggests prolonged activity of the central engine and a transition in the circumburst medium’s density. Finally, we discuss the implications for fundamental physics of detecting photons as energetic as 18 TeV from GRB 221009A.Conclusions. Early optical observations have proven invaluable for distinguishing between the potential origins of optical emissions in GRB 221009A, underscoring their utility in GRB physics studies. However, the rarity of such data underscores the need for dedicated telescopes capable of synchronous multiwavelength observations. Additionally, our analysis suggests that the host galaxies of TeV GRBs share commonalities with those of long and short GRBs. Expanding the sample of TeV GRBs could further solidify these findings.Key words: techniques: photometric / techniques: spectroscopic / gamma-ray burst: general / gamma-ray burst: individual: GRB 221009A

79 ASTRONOMY AND ASTROPHYSICS↗

Multi-omic characterization of a soil microbial consortium reveals critical role of succinate and glutamate metabolism during calcium carbonate precipitation

Microbially induced calcium carbonate precipitation (MICP) holds potential for use in soil stabilization and carbon sequestration, with the overall efficiency of the process being a major determinant for use in many environmental and civil engineering applications. While the biogeochemical pathways and enzymes driving MICP are known, the microbial metabolic networks and community dynamics underlying such precipitation remain poorly characterized. To address this gap, we developed a four-member consortium of soil bacteria (Curtobacterium flaccumfaciens, Rhodococcus qingshengii, Microbacterium sp., and Bacillus toyonensis), termed carbon storing consortium - A (CSC-A), that is capable of MICP. Prior work shows that MICP production is higher in CSC-A compared to the sum of carbonate produced by each member, suggesting carbonate production is driven by consortium dynamics. To that end we used a multi-omic integration approach of genomics, transcriptomics, and metabolomics to investigate potential inter-species interactions that may influence the MICP phenotype. Genomic life history characterizations identified evidence of niche specialization by B. toyonensis and Microbacterium, while metatranscriptomic analysis suggests R. qingshengii is a keystone species during growth in urea. By comparing individual species’ metabolomes to the metabolic profile of a shared well of precipitated metabolites, we identified over 200 metabolites predicted to be produced or consumed by CSC-A members. Integrating both data types to search the KEGG reactome highlighted a network centered around glutamine metabolism and branched chain amino acid biosynthesis under regulation during CSC-A growth in urea. Succinate metabolism was also a major node in this network and laboratory assays confirmed that increasing the amount of succinate in the growth medium leads to increased carbonate precipitation by CSC-A, a critical confirmation of our modeling approach. By isolating and identifying the interconnected metabolic components underlying MICP in CSC-A, we identified keystone taxa, metabolites, and pathways important for future optimization of the application of this consortia to carbonate precipitation.

carbon storing consortium - A (CSC-A)↗

Use and Siting of Electric Vehicle Charging Stations in Juneau, Alaska

This report details a study of electric vehicle (EV) Level 2 charging stations in Juneau, Alaska. Utilization analyses of six public over five years and 250 residential chargers over two years are included, and a composite score is introduced to identify optimal locations for future charging stations that target residents of manufactured and multifamily housing (MMFH) in Juneau. We find that public charging station usage is very location-dependent, with three chargers in use more than 60% of days during the peak hour of the day (which ranges from 10 a.m. to 7 p.m.), including a charger near residential housing, illuminating potential needs for additional public chargers in those areas. Residential charging utilization typically occurs overnight - opposite to most public charging stations analyzed - and spikes after 10 p.m. This suggests that Alaska Electric Light & Power Company's time-of-use charging program, which lowers electricity rates at 10 p.m. to incentivize overnight charging, is very effective. Residential charging data also show that households tend to charge 15 hours per week, or 9% of the time, meaning that multiple households could likely share one charger if one were provided near MMFH locations. This is supported by residential charging session analysis, which shows that the median household has around two night charging sessions per week. The EV siting analysis identifies areas of high housing density, low access to public chargers, and unconstrained feeders. A cluster of MMFH parcels in Douglas demonstrated the highest composite scores considering all factors, being the only area to have a perfect score of 2.25.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Substitution or Shared Utilization? Intrahousehold Vehicle Use in Mixed-Powertrain Households

While previous research has focused heavily on understanding the factors deriving alternative fuel vehicle adoption rates, there remains a significant gap in understanding how households distribute mileage across different powertrains. This study utilizes data from the 2022 Next Generation National Household Travel Survey to investigate vehicle miles traveled within a sample of 150 plug-in electric vehicle (PEV)-owning households (in which at least one battery electric vehicle is present), characterizing how different powertrains are integrated into daily mobility. Leveraging a Seemingly Unrelated Regression (SUR) framework the study jointly models the utilization of PEVs, hybrid electric vehicles (HEV), and internal combustion engine vehicles (ICEVs) while accounting for household-level substitution effects. The results provide evidence of an asymmetric substitution effect. In households with mixed-powertrain configurations, the ICEV captures a substantially higher share of household miles (compared with the PEV), acting as a utility sponge. Conversely, the model identifies specific socioeconomic and geographic cohorts that prioritize PEV as the primary household workhorse, indicating a systematic sorting effect. Although the sample size limits broader generalizability, these findings suggest that PEVs are used for frequent, specific routine-intensive roles, whereas the ICEV remains a specialized utility vehicle. These insights highlight distinct intrahousehold vehicle use behaviors that are often obscured by aggregate fleetwide statistics.

25 ENERGY STORAGE↗

Comparison of structurally diverse simulation models for prediction of epidemic outcomes caused by a long-distance dispersed pathogen

Long-distance dispersal (LDD) pathogens pose substantial challenges for epidemic control due to their ability to generate new infection foci at great distances. While various modeling approaches have been developed to understand and manage such outbreaks, little work has compared how models of different structures behave under shared conditions. Here, in this study, we compare four structurally distinct epidemiological models — EPIMUL, GEMF, PoPS, and Warwick — each adapted to simulate the spread of wheat stripe rust (WSR), a wind-dispersed LDD pathogen, under identical epidemiological parameters and dispersal kernel. Using data from a controlled field experiment, we evaluate the ability of each model to replicate disease prevalence under nine intervention scenarios that vary in timing and culling area. While the models differ substantially in design — ranging from spatial grid-based to network-based and raster-based frameworks — the shared dispersal kernel allowed for close alignment in their predictions. All models accurately captured general epidemic trends, particularly the strong effect of early intervention on disease suppression. We qualitatively compared their behavioral responses across scenarios and also evaluated an ensemble prediction by averaging across model outputs. Our findings highlight how integrating shared epidemiological components into distinct modeling frameworks can improve consistency and accuracy, while reinforcing the importance of early culling in managing LDD pathogen outbreaks.

Dispersal kernel↗

High-throughput and data-driven search for stable optoelectronic AMSe 3 materials

The rapid advancement in emerging optoelectronic technologies demands highly efficient, affordable, and ecofriendly materials. In this context, ternary chalcogenides, especially ternary selenides, show early promise as a material class due to their stability and remarkable electronic, optical, and transport properties. In this work, we integrate first-principles-based high-throughput computations with machine learning (ML) techniques to predict the thermodynamic stability and optoelectronic properties of 920 valency-satisfied selenide compounds. Through investigating polymorphism, our study reveals the edge-sharing orthorhombic Pnma phase (NH 4 CdCl 3 -type) as the most stable structure for most ternary selenides. High-fidelity supervised ML models are trained and tested to accelerate stability and band gap predictions. These data-driven models pin down the most influential features that dominantly control key material characteristics. The multistep high-throughput computations identify the ternary selenides with optimal direct band gaps, light carrier masses, and strong optical absorption edges. The extensive materials screening considering phase stability, toxicity, and defect tolerance, finally identifies the seven most suitable candidates for photovoltaic applications. Two of these final compounds, SrZrSe 3 and SrHfSe 3 , have already been synthesized in a single-phase form, with the latter showing an optically suitable band gap, aligning well with our findings. The non-adiabatic molecular dynamics reveal sufficiently long photoexcited charge carrier lifetimes (on the order of nanoseconds) in some of these selected selenide materials, indicating their exciting characteristics. Overall, our study suggests a robust in silico framework that can be extended to screen large datasets of various material classes for identifying promising photoactive candidates.

36 MATERIALS SCIENCE↗

Implications of increased spatial and trophic overlap between juvenile Pacific salmon and Sablefish in the northern California Current

Abstract Objective The study was designed to assess long-term variability in the distribution of juvenile Pacific salmon Oncorhynchus spp. and Sablefish Anoplopoma fimbria. The study also evaluated whether Sablefish and Pacific salmon shared food resources and looked to characterize Sablefish during an understudied period of their life cycle. Methods To meet the objectives, the study used data from 26 years of surface trawls conducted in Oregon and Washington coastal waters (1998–2023). Spatial–temporal models were used to measure changes in abundance and distribution of Pacific salmon and Sablefish along with covariates of ocean temperature. The study evaluated trophic characteristics of Pacific salmon and Sablefish from 2020 for differences. The temporal variation in size and diets of Sablefish were also analyzed, along with energy density of fish caught in 2020. Result The spatial–temporal model demonstrated that there has been a nearshore expansion of juvenile Sablefish over the past 26 years that was correlated with increased ocean temperature. The nearshore expansion of Sablefish resulted in increased spatial and trophic overlap with juvenile Pacific salmon. While feeding in nearshore waters, juvenile Sablefish demonstrated competitive feeding advantages over juvenile Pacific salmon during a critical phase of salmonid early marine life history. Juvenile Sablefish exhibited significant ontogenetic diet and energetic shifts, and even the smallest (68–80 mm fork length) were piscivorous. Conclusions If juvenile Sablefish numbers continue to increase relative to Pacific salmon, they could exert more competitive pressure, especially if food resources become limited. Pacific salmon may experience adverse effects from competition, regardless of whether or not juvenile Sablefish, which have recently expanded into nearshore waters, successfully recruit to the adult population.

Daly, Elizabeth A. (ORCID:0000000195334457)↗

Selection of Global Climate Model Data for Downscaling With Generative Machine Learning and Use in the Power Planning for Alignment of Climate and Energy Systems Project

The range of results from climate models and scenarios is important to the understanding of uncertainty in power planning analysis. A U.S. Department of Energy-funded analytic project called Power Planning for Alignment of Climate and Energy Systems is developing data and analytic methods to reflect the effects of climate change on key variables for power system planning, as part of the Grid Modernization Lab Consortium. This project will select and prepare global climate model results for use in power system planning models. A related report (Evaluation of Global Climate Models for Use in Energy Analysis) assesses the performance of various global climate models from the Coupled Model Intercomparison Project Phase 6 data archive for their historical skill with respect to energy system performance and for their future projections under multiple climate change scenarios. Building from that report, we describe the selection of a climate scenario (Shared Socioeconomic Pathway [SSP] 2-4.5) and five climate models: TaiESM1, EC-Earth3-CC, GFDL-CM4, EC-Earth3-Veg, and MPI-ESM1-2-HR. We describe the model selection criteria, which were based on the quality of the match between model results under historical conditions and on the representation of the range of future values for several variables. These results will be downscaled via an open-source generative machine learning method called Super-Resolution for Renewable Energy Resource Data with Climate Change Impacts.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

New approaches to Bayesian uncertainty quantification for Nuclear Science (Final Technical Report)

Inverse problems play a central role in experimentation and theory/data comparisons for many areas of modern Nuclear Physics (NP) and High-Energy Physics (HEP). Bayes’s Theorem is a powerful tool for solving Inverse Problems, providing conceptually transparent and unbiased constraints on theoretical parameters and their uncertainties (“Bayesian Inference”) and enabling the quantification of agreement or tension between models and data. However, analyses based on Bayesian Inference are often challenging for NP and HEP applications, either because of the large number of parameters in the problem, the high computational cost, or both. We propose a multi-institutional collaboration to develop and deploy novel Bayesian analysis tools that advance the scientific scope of a broad range of current and future NP experiments. This project brings together NP domain scientists working on several high-profile NP projects for which new, high-performance Bayesian Uncertainty Quantification (“Bayesian UQ”) methods are essential to carry out the science, and data scientists who are developing state-of-the-art methods applicable to these problems. The NP projects in this proposal comprise measurements of the mass and fundamental nature of the neutrino; study of the Quark-Gluon Plasma that filled the early universe; and mapping of natural and anthropogenic radiation environments. While these NP projects have very different scientific goals, with datasets and analysis approaches that differ significantly, they share common requirements for improving computationally intensive Bayesian analyses using advanced Machine Learning algorithms and will benefit strongly from a coherent effort to develop general solutions. This proposal brings together these projects and forefront ML-based data science algorithms to develop such general solutions. The methods developed in this project will also be more widely applicable, thereby advancing science in the larger Nuclear Physics portfolio.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Mathematical Morphological Filtering with a Self-Adaptive Reconstruction Technique and Application to Local Seismic Data

Recorded seismic data are generally contaminated by noise from different sources, which masks the signals of interest. In the seismology community, frequency filtering (FF) is the standard method for noise suppression. However, when the signal of interest and noise share the same frequency band, the latter cannot be filtered out without infringing on the former. We implemented a noise suppression approach based on the mathematical morphology theorem. The method involves compound operations of dilation and erosion using structuring elements of varying lengths and decomposes an input noisy waveform into several time functions with differing characteristics. Further, the filtered waveform is constructed from the time functions using a self-adaptive reconstruction technique. Application to a data set of >4700 local waveforms suggests that the implemented mathematical morphological filtering (MMF) approach is efficient for data with low signal-to-noise ratio (SNR) and significantly outperforms FF in that SNR range. For most of the dataset, FF, machine learning (ML) denoising, and continuous wavelet transform (CWT) thresholding result in higher SNR values compared with the MMF method. However, for ~42% of the waveforms, MMF outperforms FF, and the SNR gain achieved with MMF is as large as ~23 dB. Compared to ML denoising and CWT thresholding, this proportion drops to only ~10%–14%. Our results suggests that in an operational setting, MMF cannot replace the other noise suppression methods; however, signal detection can be improved if MMF is used to supplement them in some scenarios. MMF could help detect signals in problematic low-SNR data, which are currently being missed particularly when using FF alone.

58 GEOSCIENCES↗

Increasing Mosquito Abundance Under Global Warming

Mosquitoes are a key virus vector that poses significant health threats globally, affecting 700 million individuals and causing 1 million deaths annually. Accurately predicting mosquito abundance and dispersion remains a challenge. Complex interactions between mosquito dynamics and various environmental factors, notably hydrology, contribute to this challenge. Existing models typically focus on precipitation and temperature and often overlook further impacts of hydrological variables within mosquito modeling. In this study, we developed an artificial intelligence‐based model for mosquito dynamics, explicitly accounting for different hydrological variables, such as precipitation, soil moisture and streamflow. Using Toronto, Canada, as a case study, we identified causal relationships between changes in mosquito populations, hydrological factors, vegetation (e.g., leaf area index), and climate variables (e.g., daylight length, precipitation, and temperature). We embedded these relationships into a Long Short‐Term Memory (LSTM) Neural Network Model capable of accurately detecting mosquito dynamics across annual, seasonal, and monthly time scales. The LSTM is able to explain, on average, approximately 40% of the variance in the observed mosquito abundance data. Using the calibrated model, we predicted that the summer season mosquito abundance would increase by ∼16% and ∼19% under an intermediate greenhouse emission scenario, Shared Socioeconomic Pathway (SSP) 2–4.5, and a high greenhouse emission scenario, SSP5‐8.5, respectively. We expect that this model can serve as a valuable tool and inform science‐based decisions affecting mosquito dynamics and public health. It can also build a foundation for future risk analysis at the regional and larger scales.

54 ENVIRONMENTAL SCIENCES↗

Methane fluxes in tidal marshes of the conterminous United States

Abstract Methane (CH 4 ) is a potent greenhouse gas (GHG) with atmospheric concentrations that have nearly tripled since pre‐industrial times. Wetlands account for a large share of global CH 4 emissions, yet the magnitude and factors controlling CH 4 fluxes in tidal wetlands remain uncertain. We synthesized CH 4 flux data from 100 chamber and 9 eddy covariance (EC) sites across tidal marshes in the conterminous United States to assess controlling factors and improve predictions of CH 4 emissions. This effort included creating an open‐source database of chamber‐based GHG fluxes ( https://doi.org/10.25573/serc.14227085 ). Annual fluxes across chamber and EC sites averaged 26 ± 53 g CH 4 m −2 year −1 , with a median of 3.9 g CH 4 m −2 year −1 , and only 25% of sites exceeding 18 g CH 4 m −2 year −1 . The highest fluxes were observed at fresh‐oligohaline sites with daily maximum temperature normals (MATmax) above 25.6°C. These were followed by frequently inundated low and mid‐fresh‐oligohaline marshes with MATmax ≤25.6°C, and mesohaline sites with MATmax >19°C. Quantile regressions of paired chamber CH 4 flux and porewater biogeochemistry revealed that the 90th percentile of fluxes fell below 5 ± 3 nmol m −2 s −1 at sulfate concentrations >4.7 ± 0.6 mM, porewater salinity >21 ± 2 psu, or surface water salinity >15 ± 3 psu. Across sites, salinity was the dominant predictor of annual CH 4 fluxes, while within sites, temperature, gross primary productivity (GPP), and tidal height controlled variability at diel and seasonal scales. At the diel scale, GPP preceded temperature in importance for predicting CH 4 flux changes, while the opposite was observed at the seasonal scale. Water levels influenced the timing and pathway of diel CH 4 fluxes, with pulsed releases of stored CH 4 at low to rising tide. This study provides data and methods to improve tidal marsh CH 4 emission estimates, support blue carbon assessments, and refine national and global GHG inventories.

54 ENVIRONMENTAL SCIENCES↗

SOX2-driven enhancer landscape defines the transcriptional architecture of retinogenesis

Retinal neurogenesis is mediated by the coordinated activities of a complex gene regulatory network (GRN) of transcription factors (TFs) in multipotent retinal progenitor cells (RPCs). How this GRN mechanistically guides neural competence remains poorly understood. In this study, we present integrated transcriptional, genetic and genomic analyses to uncover the regulatory mechanisms of SOX2, a key factor in establishing neural identity in RPCs. We show that SOX2 is preferentially enriched in the RPC-specific enhancer landscape associated with essential regulators of retinogenesis. Disruption of SOX2 expression impairs retinogenesis, marked by a selective loss of enhancer activity near genes essential for RPC proliferation and lineage specification. We identified the RPC transcription factor VSX2 as a binding partner for SOX2 and, together, SOX2 and VSX2 co-target a core, retina-specific chromatin repertoire characterized by enhanced TF binding and robust chromatin accessibility. This cooperative binding establishes a shared SOX2-VSX2 transcriptional code that promotes the expression of crucial regulators of neurogenesis while repressing the acquisition of alternative lineage cell fate. Our data illuminate fundamental biological insights on how transcription factors act in concert to drive chromatin-based genetic programs underlying retinal neural identity.

Chromatin↗

OPEN-Augmented Reality GUI for Bioenergy Crop Phenotyping and Precision Agriculture (Donald Danforth Plant Science Center Final Scientific Technical Report)

The project led by the Donald Danforth Plant Science Center, in collaboration with Arizona State University, George Washington University, and Saint Louis University, has made significant strides in advancing the phenotypic analysis of bioenergy crops through the development of an innovative AI processing pipeline. This initiative was primarily funded by ARPA-E, with additional cost-sharing provided by the participating institutions. The project successfully utilized a variety of sensors—3D scanners, thermal, RGB, and hyperspectral—to refine algorithms for data-driven trait signature identification and improve the classification and visualization of plant traits. The developed AI processing pipeline is capable of handling the complex, multidimensional data characteristic of dynamic agricultural environments. 1) Contributions to understanding: The research has advanced the field of plant phenomics by showcasing the synergistic use of various sensor data to enhance the precision of trait analysis in bioenergy crops. Through the integration of 3D scanners, thermal, RGB, and hyperspectral sensors, the project has developed robust data-driven trait signature algorithms and visualization techniques. These innovations have facilitated detailed monitoring and management of plant traits, providing vital insights into plant growth dynamics and stress responses. Further, the project has broadened our understanding of how machine learning can be effectively applied in multi-sensor environments to refine trait analysis. By leveraging diverse datasets, the research has not only improved the accuracy of phenotypic assessments but also established a versatile methodological framework that can be extended beyond agriculture to other fields requiring detailed phenotypic analysis. 2) Technical effectiveness and economic feasibility: The AI processing pipeline developed in this project demonstrated significant technical effectiveness, achieving high throughput analysis of extensive phenotypic data and meeting targeted accuracies. This system exemplified the capability of advanced machine learning technologies to efficiently manage and analyze large, complex datasets. Economically, the implementation of the project-developed pipelines may offer substantial cost savings across multiple sectors. It enhances data analysis processes and significantly reduces the need for manual data interpretation, thereby decreasing both the time and resources required. 3) Public benefit: The project has significantly broadened the scope of agricultural methodologies to enhance phenotypic analysis, with potential applications in various sectors beyond agriculture. Additionally, the initiative fostered an enriching educational and collaborative environment, significantly enhancing the technical skills of participants. It also made substantial contributions to the scientific community by providing open-access data sets and tools, encouraging ongoing research and development across various disciplines. Overall, the project not only met its scientific goals but also showcased the extensive utility of integrating advanced machine learning and sensor data analysis technologies. These advancements have proven instrumental in driving forward both theoretical research and practical applications, setting a strong foundation for future explorations and innovations in data-driven science.

60 APPLIED LIFE SCIENCES↗

Analyzing and Exploring Training Recipes for Large-Scale Transformer-Based Weather Prediction

Abstract The rapid rise of deep learning (DL) in numerical weather prediction (NWP) has led to a proliferation of models which forecast atmospheric variables with comparable or superior skill than traditional physics-based NWP. However, among these leading DL models, there is a wide variance in both the training settings and architecture used. Further, the lack of thorough ablation studies makes it hard to discern which components are most critical to success. In this work, we show that it is possible to attain high forecast skill even with relatively off-the-shelf architectures, simple training procedures, and moderate compute budgets. Specifically, we train a minimally modified Swin Transformer V2 (SwinV2) on ERA5 data and find that it attains superior skill in terms of mean-square errors of deterministic forecasts when compared against the European Centre for Medium-Range Weather Forecasts’ Integrated Forecasting System (IFS). Almost all DL–NWP systems share a core set of hyperparameters and design decisions. To aid and expedite future DL–NWP research, we present an in-depth, systematic exploration of different loss functions, model sizes and depths, patch sizes, and multistep training objectives. We also examine the model performance with metrics beyond the typical accuracy (ACC) and RMSE and investigate how the performance scales with model size. Through our open-source code, scoring pipelines, and models, we share our findings on key aspects of the training pipeline. These ablations reduce the necessity for expensive hyperparameter tuning and lower the barrier to entry for future DL–NWP research. Significance Statement This study investigates the potential of using large-scale transformer-based models for weather prediction, showing that it is possible to achieve high forecast accuracy with simpler, off-the-shelf architectures. By training a minimally modified SwinV2 transformer on ERA5 data, we show that the model achieves competitive forecast skill in terms of mean-square error for key variables, outperforming the European Centre for Medium-Range Weather Forecasts’ Integrated Forecasting System (IFS) at all lead times. Our findings suggest that effective training strategies, such as multistep fine-tuning and channel-weighted losses, significantly enhance the model’s performance. However, we also highlight that these improvements come with trade-offs in other areas, such as ensemble spread and high-frequency spatial detail. This work highlights the promise of deep learning in improving weather forecasts, which could lead to better preparedness and response to weather events, ultimately benefiting society by providing more reliable weather predictions.

Willard, Jared D. [Lawrence Berkeley National Labo↗