Search NASA⌕ Search

SEARCH · Search NASA

Results for “predictor”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Objective Lightning Probability Forecasting for Kennedy Space Center and Cape Canaveral Air Force Station, Phase III

The AMU created new logistic regression equations in an effort to increase the skill of the Objective Lightning Forecast Tool developed in Phase II (Lambert 2007). One equation was created for each of five sub-seasons based on the daily lightning climatology instead of by month as was done in Phase II. The assumption was that these equations would capture the physical attributes that contribute to thunderstorm formation more so than monthly equations. However, the SS values in Section 5.3.2 showed that the Phase III equations had worse skill than the Phase II equations and, therefore, will not be transitioned into operations. The current Objective Lightning Forecast Tool developed in Phase II will continue to be used operationally in MIDDS. Three warm seasons were added to the Phase II dataset to increase the POR from 17 to 20 years (1989-2008), and data for October were included since the daily climatology showed lightning occurrence extending into that month. None of the three methods tested to determine the start of the subseason in each individual year were able to discern the start dates with consistent accuracy. Therefore, the start dates were determined by the daily climatology shown in Figure 10 and were the same in every year. The procedures used to create the predictors and develop the equations were identical to those in Phase II. The equations were made up of one to three predictors. TI and the flow regime probabilities were the top predictors followed by 1-day persistence, then VT and Ll. Each equation outperformed four other forecast methods by 7-57% using the verification dataset, but the new equations were outperformed by the Phase II equations in every sub-season. The reason for the degradation may be due to the fact that the same sub-season start dates were used in every year. It is likely there was overlap of sub-season days at the beginning and end of each defined sub-season in each individual year, which could very well affect equation performance.

Crawford, Winifred C.↗

Peak Wind Tool for General Forecasting

The expected peak wind speed of the day is an important forecast element in the 45th Weather Squadron's (45 WS) daily 24-Hour and Weekly Planning Forecasts. The forecasts are used for ground and space launch operations at the Kennedy Space Center (KSC) and Cape Canaveral Air Force Station (CCAFS). The 45 WS also issues wind advisories for KSC/CCAFS when they expect wind gusts to meet or exceed 25 kt, 35 kt and 50 kt thresholds at any level from the surface to 300 ft. The 45 WS forecasters have indicated peak wind speeds are challenging to forecast, particularly in the cool season months of October - April. In Phase I of this task, the Applied Meteorology Unit (AMU) developed a tool to help the 45 WS forecast non-convective winds at KSC/CCAFS for the 24-hour period of 0800 to 0800 local time. The tool was delivered as a Microsoft Excel graphical user interface (GUI). The GUI displayed the forecast of peak wind speed, 5-minute average wind speed at the time of the peak wind, timing of the peak wind and probability the peak speed would meet or exceed 25 kt, 35 kt and 50 kt. For the current task (Phase II ), the 45 WS requested additional observations be used for the creation of the forecast equations by expanding the period of record (POR). Additional parameters were evaluated as predictors, including wind speeds between 500 ft and 3000 ft, static stability classification, Bulk Richardson Number, mixing depth, vertical wind shear, temperature inversion strength and depth and wind direction. Using a verification data set, the AMU compared the performance of the Phase I and II prediction methods. Just as in Phase I, the tool was delivered as a Microsoft Excel GUI. The 45 WS requested the tool also be available in the Meteorological Interactive Data Display System (MIDDS). The AMU first expanded the POR by two years by adding tower observations, surface observations and CCAFS (XMR) soundings for the cool season months of March 2007 to April 2009. The POR was expanded again by six years, from October 1996 to April 2002, by interpolating 1000-ft sounding data to 100-ft increments. The Phase II developmental data set included observations for the cool season months of October 1996 to February 2007. The AMU calculated 68 candidate predictors from the XMR soundings, to include 19 stability parameters, 48 wind speed parameters and one wind shear parameter. Each day in the data set was stratified by synoptic weather pattern, low-level wind direction, precipitation and Richardson Number, for a total of 60 stratification methods. Linear regression equations, using the 68 predictors and 60 stratification methods, were created for the tool's three forecast parameters: the highest peak wind speed of the day (PWSD), 5-minute average speed at the same time (A WSD), and timing of the PWSD. For PWSD and A WSD, 30 Phase II methods were selected for evaluation in the verification data set. For timing of the PWSD, 12 Phase\I methods were selected for evaluation. The verification data set contained observations for the cool season months of March 2007 to April 2009. The data set was used to compare the Phase I and II forecast methods to climatology, model forecast winds and wind advisories issued by the 45 WS. The model forecast winds were derived from the 0000 and 1200 UTC runs of the 12-km North American Mesoscale (MesoNAM) model. The forecast methods that performed the best in the verification data set were selected for the Phase II version of the tool. For PWSD and A WSD, linear regression equations based on MesoNAM forecasts performed significantly better than the Phase I and II methods. For timing of the PWSD, none of the methods performed significantly bener than climatology. The AMU then developed the Microsoft Excel and MIDDS GUls. The GUIs display the forecasts for PWSD, AWSD and the probability the PWSD will meet or exceed 25 kt, 35 kt and 50 kt. Since none of the prediction methods for timing of the PWSD performed significantly better thanlimatology, the tool no longer displays this predictand. The Excel and MIDDS GUIs display forecasts for Day-I to Day-3 and Day-I to Day-5, respectively. The Excel GUI uses MesoNAM forecasts as input, while the MIDDS GUI uses input from the MesoNAM and Global Forecast System model. Based on feedback from the 45 WS, the AMU added the daily average wind speed from 30 ft to 60 ft to the tool, which is one of the parameters in the 24-Hour and Weekly Planning Forecasts issued by the 45 WS. In addition, the AMU expanded the MIDDS GUI to include forecasts out to Day-7.

Barrett, Joe H., III↗

Use of Minute-by-Minute Cardiovascular Measurements During Tilt Tests to Strengthen Inference on the Effect of Long-Duration Space Flight on Orthostatic Hypotension

Typical methodology for evaluating the effects of spaceflight on orthostatic hypotension (OH) has been survival analysis of tolerance times from 80 head-up tilt tests. However when scheduled test durations are short, there may not be enough failures to allow survival analysis to adequately estimate and compare the effects of flight phase (e.g. pre-flight, number of days post-flight), flight duration, and their interaction, as well as interactions with effects of interventions or countermeasures. The problem is exacerbated in the presence of a repeated measures design, in which subjects participate in tilt tests during various flight phases. Here we show how it is possible to dramatically improve the efficiency of statistical inference in this setting by making use of the additional information contained in minute-by-minute observations of cardiovascular parameters thought to be reflective of progression towards presyncope during tilt testing. Methods: We retrospectively examined operational tilt test (OTT; 10 -min 80 head-up tilt) data from 20 International Space Station (ISS) and 66 Shuttle astronauts 10 d before launch (L-10), on landing day (R+0) and during recovery (R+1, R+3, R+6-10) depending on the level of participation. Data from 5 ISS astronauts tested on R+0 or R+1 who used non-standard countermeasures were excluded. In addition to OTT survival time, 8 cardiovascular parameters (CP: heart rate, systolic, diastolic, and mean arterial blood pressure, pulse pressure, stroke volume, cardiac output, and total peripheral resistance) that might be predictive of progression towards presyncope were measured every minute of each OTT. Statistical analysis was predicated on a two ]stage model of causation. In the first stage, flight duration and time from landing affect the astronauts' degree of OH, which is manifested in the time trends and variation of the above CPs during OTTs. In the second stage, the behavior of these parameters directly affects OTT survival time. Actual analysis proceeded in the opposite direction. First we identified those CPs or linear combinations that best predicted OTT survival regardless of what spaceflight conditions led to OTT completion or presyncope. From these, we calculated a summary statistic (one per OTT) that best predicted survival. We then used mixed ]model regression analysis to relate changes in the summary statistic to flight phase and duration. Inference on the effects of phase, duration, and their interaction on OH follows directly from this second analysis. Results: A linear combination (W) of diastolic blood pressure (DBP) and stroke volume (SV) was found to be the best predictor of OTT survival using the complete data set of minute-by-minute observations of CPs for each OTT. Furthermore, the log-transformed standard deviation of W (Z = log SW) was found to be a strong predictor of survival in the reduced data set consisting of one observation per OTT. In other words, this measure of variability of W during an OTT was the best indicator of whether or not the subject could complete the 10-min test, with higher variability (i.e. higher values of Z) being associated with greater probability of failure. In the mixed-model regression analysis where Z was now treated as a outcome with flight phase and duration groups (ISS and STS) as predictors, we found that there was a significantly more variability in W (higher values of Z) for both groups at R+0, but with no evidence of an interaction until R+3, when the ISS group still had inflated variability, but not the STS group. Conclusions: Variability of the cardiovascular index W recovers more slowly after long-compared to short-duration spaceflight. Since high variability of W has also been shown to be predictive of OTT failure, a primary manifestation of OH, a logical conclusion is that recovery from OH also is slower after long-duration compared to short-duration spaceflights.

Feiveson, Alan H.↗

A Comparison of Three Algorithms for Orion Drogue Parachute Release

The Orion Multi-Purpose Crew Vehicle is susceptible to ipping apex forward between drogue parachute release and main parachute in ation. A smart drogue release algorithm is required to select a drogue release condition that will not result in an apex forward main parachute deployment. The baseline algorithm is simple and elegant, but does not perform as well as desired in drogue failure cases. A simple modi cation to the baseline algorithm can improve performance, but can also sometimes fail to identify a good release condition. A new algorithm employing simpli ed rotational dynamics and a numeric predictor to minimize a rotational energy metric is proposed. A Monte Carlo analysis of a drogue failure scenario is used to compare the performance of the algorithms. The numeric predictor prevents more of the cases from ipping apex forward, and also results in an improvement in the capsule attitude at main bag extraction. The sensitivity of the numeric predictor to aerodynamic dispersions, errors in the navigated state, and execution rate is investigated, showing little degradation in performance.

Matz, Daniel A.↗

Climate change reshapes the drivers of false spring risk across European trees

(1) Temperate forests are shaped by late spring freezes after budburst—false springs—which may shift with climate change. Research to date has generated conflicting results, potentially because few studies focus on the multiple underlying drivers of false spring risk. (2) Here, we assessed the effects of mean spring temperature, distance from the coast, elevation and the North Atlantic Oscillation (NAO) using PEP725 leafout data for six tree species across 11648 sites in Europe, to determine which were the strongest predictors of false spring risk and how these predictors shifted with climate change. (3) All predictors influenced false spring risk before recent warming, but their effects have shifted in both magnitude and direction with warming. These shifts have potentially magnified the variation in false spring risk among species with an increase in risk for early‐leafout species (i.e., Aesculus hippocastanum , Alnus glutinosa , Betula pendula ) versus a decline or no change in risk among late‐leafout species (i.e., Fagus sylvatica , Fraxinus excelsior , Quercus robur ). (4) Our results show how climate change has reshaped the drivers of false spring risk, complicating forecasts of future false springs, and potentially reshaping plant community dynamics given uneven shifts in risk across species.

false spring↗

Explicit Discontinuous Galerkin Methods for Conservation Laws

The two explicit DG methods in this study are based on a ‘predictor-corrector’ formulation, the first introduced by Lörcher, Gassner, and Munz (2007, 2008) called space–time expansion discontinuous Galerkin or STE-DG scheme, and the second, introduced independently by the author (Huynh 2006, 2013) called the upwind moment scheme. The predictor step of the two methods is essentially identical using a Cauchy-Kovalevsky (CK) procedure, which involves no interaction of the data among neighboring cells. The corrector step also shares the same space-time integration formulation and is where interaction of the data among neighboring cells takes place; the difference, however, is in how the resulting space-time volume integral is estimated. As a consequence of the different estimates, for the case of advection in one spatial dimension (1D), the moment scheme has a CFL (Courant-Friedrichs-Lewy) condition of 1 for all p and is accurate to order 2p+1, i.e., it possesses the super accuracy property, whereas the STE-DG method has a more restrictive CFL condition and is accurate to the expected order of p+1. For 1D advection, compared with the CFL conditions of 1/(2p+1) of standard RK-DG (Runge-Kutta) scheme where space and time discretization are of the same order, the moment scheme allows a significantly larger time step size. It also turns out that the scheme yields a result identical to Van Leer’s scheme III (1977), which amounts to shifting the data a distance of advection corresponding to the time step and projecting the result onto the space of polynomial solutions. Contrary to Van Leer’s approach, however, the space-time ‘predictor-corrector’ formulation facilitates extensions to the case of systems of equations. Concerning 2D extensions, in the case of advection, when the flow is along the diagonal direction, the CFL conditions for the moment schemes become restrictive as will be shown by Fourier (Von Neumann) stability and accuracy analyses. Since the moment scheme employs the right Radau points as collocation points in time, the method is closely related to the implicit Radau IIA scheme, which is stable for any time step size. The role of Radau IIA in relieving stability restriction for these explicit DG schemes remains to be explored

Discontinuous Galerkin↗

Using Satellite Soil Moisture and Rainfall in the Landslide Hazard Assessment for Situational Awareness System

The Landslide Hazard Assessment for Situational Awareness system(LHASA)gives a global view of landslide hazard in nearly real time. Currently, it is being upgraded from version 1 to version 2, which entails improvements along several dimensions. These include the incorporation of new predictors, machine learning, and new event-based landslide inventories. As a result, LHASA version 2 substantially improves on the prior performanceand introduces a probabilistic element to the global landslide nowcast. Data from the soil moisture active-passive (SMAP) satellite has been assimilated into a globally consistent data product with a latency less than 3 days, known as SMAP Level 4. In LHASA, thesedata representthe antecedent conditions prior to landslide-triggering rainfall. In some cases, soil moisture may have accumulated over aperiod of many months. The model behind SMAP Level 4 also estimates the amount of snow on the ground, which is an important factor in some landslide events. LHASA also incorporates this information as an antecedent condition that modulates the response torainfall. Slope, lithology, and active faults were also used as predictor variables. These factors can have a strong influence on where landslides initiate.LHASA relies on precipitation estimates from the Global Precipitation Measurement mission to identify the locations where landslides are most probable. The low latency and consistent global coverage of these data make them ideal for real-time applications at continental to global scales. LHASA relies primarily on rainfall from the last 24 hours to spothazardous sites, which is rescaled by the local 99thpercentile rainfall.However, the multi-day latency of SMAP requires the use of a 2-day antecedent rainfall variable to represent the accumulation of rain between the antecedent soil moisture and current rainfall. LHASA merges these predictors with XGBoost, a commonly used machine-learning tool, relying on historical landslide inventories to develop the relationship between landslide occurrence and various risk factors. The resulting model relies heavily on current daily rainfall, but other factors also play an important role. LHASA outputsthe probability oflandslide occurrence ona grid of roughly one kilometer over all continents from 60 North to 60 South latitude. Evaluation over the period 2019-2020 showsthat LHASA version 2 doubles the accuracy of the global landslide nowcast without increasing the global false alarm rate. LHASA also identifies the areas where the human exposure to landslide hazard is most intense. Landslide hazard is divided into 4 levels: minimal, low, moderate, and high. Next, the number of persons and the length of major roads (primary and secondary roads)within each of these areas is calculated for every second-level administrative district (county). These results can be viewedthrough a web portal hosted at the Goddard Space Flight Center. In addition, users can download daily hazard and exposure data.LHASAversion 2uses machine learning and satellite data to identify areas of probable landslide hazard within hours of heavy rainfall. Itsglobal maps are significantly more accurate, and it now includes rapid estimates of exposed populations and infrastructure. In addition, a forecast mode will be implemented soon.

Thomas Stanley↗

Aboveground biomass density models for NASA’s Global Ecosystem Dynamics Investigation (GEDI) lidar mission

NASA’s Global Ecosystem Dynamics Investigation (GEDI) is collecting spaceborne full waveform lidar data with a primary science goal of producing accurate estimates of forest aboveground biomass density (AGBD). This paper presents the development of the models used to create GEDI’s footprint-level (~25 m) AGBD (GEDI04_A) product, including a description of the datasets used and the procedure for final model selection. The data used to fit our models are from a compilation of globally distributed spatially and temporally coincident field and airborne lidar datasets, whereby we simulated GEDI-like waveforms from airborne lidar to build a calibration database. We used this database to expand the geographic extent of past waveform lidar studies, and divided the globe into four broad strata by Plant Functional Type (PFT) and six geographic regions. GEDI’s waveform-to-biomass models take the form of parametric Ordinary Least Squares (OLS) models with simulated Relative Height (RH) metrics as predictor variables. From an exhaustive set of candidate models, we selected the best input predictor variables, and data transformations for each geographic stratum in the GEDI domain to produce a set of comprehensive predictive footprint-level models. We found that model selection frequently favored combinations of RH metrics at the 98th, 90th, 50th, and 10th height above ground-level percentiles (RH98, RH90, RH50, and RH10, respectively), but that inclusion of lower RH metrics (e.g. RH10) did not markedly improve model performance. Second, forced inclusion of RH98 in all models was important and did not degrade model performance, and the best performing models were parsimonious, typically having only 1-3 predictors. Third, stratification by geographic domain (PFT, geographic region) improved model performance in comparison to global models without stratification. Fourth, for the vast majority of strata, the best performing models were fit using square root transformation of field AGBD and/or height metrics. There was considerable variability in model performance across geographic strata, and areas with sparse training data and/or high AGBD values had the poorest performance. These models are used to produce global predictions of AGBD, but will be improved in the future as more and better training data become available.

Laura Duncanson↗

Front Range Wildland Fires: Evaluating the Efficacy of Remote Sensing Imagery in Monitoring Forest Fuels Treatment Methods

Over the last several decades, wildfire frequency and severity in forested areas along Colorado’s Front Range have increased due to a buildup of fuels. This has led to an increase in forest treatments, as well as an increased need to evaluate the success of these treatments. Remote sensing products offer an efficient and cost-effective way to monitor forest treatments; however, not all remote sensing products and analysis techniques have been explored by Coloradan land managers. Specifically, project partners at the Colorado State Forest Service (CSFS) and the Colorado Forest Restoration Institute (CFRI) were interested in using an effective and streamlined method of mapping canopy cover to better monitor forest treatment success. To support their needs, the NASA DEVELOP Front Range Wildland Fires team explored National Agricultural Imagery Program (NAIP) imagery at different spatial resolutions and numbers of training points with NASA’s Shuttle Radar Topography Mission (SRTM) Data Elevation Model (DEM) as a predictor in addition to NAIP imagery spectral predictors. From this analysis, we created classified canopy cover rasters, and compared accuracy metrics across model iterations. We also determined that the best performing model, with an overall accuracy of 0.900 uses 2021 NAIP imagery at 2-meter resolution, 800 training points, 200 testing points, does not use topographic predictors, and reclassifies shadow pixels via a pre-selected NDVI threshold.

Remote Sensing↗

Machine Learning Emulators and Empirical Models Combining Climate and Global Crop Models for Seasonal Agricultural Production

We present results from several connected efforts to apply machine learning methods to estimates of seasonal agricultural production anomalies around the world. First, we apply the XGBoost Random Forest method to fit emulators that mimic global crop models participating in the Agricultural Model Intercomparison and Improvement Project (AgMIP) Global Gridded Crop Model Intercomparison (GGCMI). These are the same models used in the agricultural sector simulations of the Inter-Sectoral Impacts Model Intercomparison Project (ISIMIP). These emulators use 8 climate variables split across 5 sub-seasonal representations of the growing season for each ½ degree grid cell around the world for maize, wheat, rice and soybeans. Emulators are useful for estimating conditions that have not already been simulated by GGCMI (e.g., in a seasonal prediction model) and also to diagnose model differences and capabilities. For example, emulators of the pDSSAT maize model tend to be more reliant on mean temperatures than the LPJmL model, and few models have strong responses to cold extremes. Second, we use a similar XGBoost approach to fit empirical models for national production data for the top 20 producing countries according to the United Nations Food and Agricultural Organization (FAO). Models utilize both climate observations and the GGCM models as predictors, resulting in skillful models for many (but not all) top producing-countries. The patterns of climate and crop model features selected indicate regions and systems that are better or worse simulated by the GGCMs. For example, information in cold extreme predictors is often combined with GGCM output predictors to provide sensitivity that models may underrepresent.

machine learning↗

Global Variability in Sonic Boom Exposure due to Macroscopic Effects

Supersonic flight over land has been prohibited since 1973 due to the loudness of sonic booms. NASA is building the X-59 aircraft as part of its Quesst mission to demonstrate low-loudness shaped sonic booms, or “sonic thumps.” The Quesst mission will gather human perception data via a series of community noise surveys across the USA. The noise dose and perceptual response data will be provided to the International Civil Aviation Organization (ICAO) and the Federal Aviation Administration for use in determining potential future supersonic aircraft noise certification standards, effectively changing the prohibition from a speed limit to a noise limit. These noise regulations must be globally effective, as long travel distances see the largest benefit to supersonic flight. The state of the atmosphere through which a sonic boom travels affects the size of the region exposed to sound, the “carpet width” (CW), as well as the loudness. The focus of this dissertation is to understand and quantify the expected loudness and CW of sonic booms due to the macroscopic atmospheric effects around the world. A pair of large-scale propagation simulation studies were conducted using the NASA PCBoom code to compare predicted sonic boom loudness and CW statistics first across the USA and then across the world. For the USA study, near-field data of the X-59 in steady cruise was propagated at 4 cardinal headings at 138 locations through 5 years of Climate Forecast System Version 2 (CFSv2) atmospheric profiles. Results of a bootstrap forest predictor screening model indicated the importance of climate zone, latitude, ground elevation, season, and heading. It also noted the unimportance of time of day for predicting loudness and CW. The data is visualized in aggregate, and then broken out geographically, by season and heading, and by climate zone. Multiple linear regression models were fit to the data from the 138 locations so that estimates of the loudness and CW can be produced anywhere in the US. The results can aid in planning when and where to fly the X-59. For the global study, near-field data from three aircraft, the X-59 in a quiet and loud configuration, B-58, and Concorde, were propagated at four cardinal headings through data from three atmospheric models, the CFSv2, the Global Forecast System (GFS), and the ECMWF Reanalysis Version 5 (ERA5), at 100 global locations over 1 year. Results of a bootstrap forest predictor screening model indicated the importance of climate zone, ground elevation, season, and heading. Similar to the US study, the model indicated time of day was not an important predictor. The model also indicated that choice of weather model was not important, so the atmospheric model data are effectively interchangeable. The ERA5 model was chosen for use in an extension of the study to include 18 additional locations to ensure sampling of every climate zone. Loudness and CW results are shown in aggregate, and split geographically and by heading, season, and climate. Multiple linear regression models were fit to the data from the 118 locations so that estimates of loudness and CW can be produced around the world. N-waves and shaped booms did not have the same global variability. Koppen-Geiger climate zones were used as the climate zone definition for the global study. These are available as present-day and future climate projections. Making use of the multiple linear regression models, the future climate zones were input to estimate the effect of the changing climate on sonic boom loudness and CW. Results indicate that a changing climate would have little impact on the effectiveness of noise regulations.

X-59↗

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT↗

Allometric relationships and trade‐offs in 11 common M editerranean‐climate grasses

Abstract Biomass allocation in plants is the foundation for understanding dynamics in ecosystem carbon balance, species competition, and plant–environment interactions. However, existing work on plant allometry has mainly focused on trees, with fewer studies having developed allometric equations for grasses. Grasses with different life histories can vary in their carbon investment by prioritizing the growth of specific organs to survive, outcompete co‐occurring plants, and ensure population persistence. Further, because grasses are important fuels for wildfire, the lack of grass allocation data adds uncertainty to process‐based models that relate plant physiology to wildfire dynamics. To fill this gap, we conducted a greenhouse experiment with 11 common California grasses varying in photosynthetic pathway and growth form. We measured plant sizes and harvested above‐ and belowground biomass throughout the life cycle of annual species, while for the establishment stage of perennial grasses to quantify allometric relationships for leaf, stem, and root biomass, as well as plant height and canopy area. We used basal diameter as a reference measure of plant size. Overall, basal diameter is the best predictor for leaf and stem biomass, height, and canopy area. Including height as another predictor can improve model accuracy in predicting leaf and stem biomass and canopy area. Fine root biomass is a function of leaf biomass alone. Species vary in their allometric relationships, with most variation occurring for plant height, canopy area, and stem biomass. We further explored potential trade‐offs in biomass allocation across species between leaf and fine root, leaf and stem, and allocation to reproduction. Consistent with our expectation, we found that fast‐growing plants allocated a greater fraction to reproduction. Additionally, plant height and specific leaf area negatively influenced the leaf‐to‐stem ratio. However, contrary to our hypothesis, there were no differences in root‐to‐leaf ratio between perennial and annual or C 4 and C 3 plants. Our study provides species‐specific and functional‐type‐specific allometry equations for both above‐ and belowground organs of 11 common California grass species, enabling nondestructive biomass assessment in California grasslands. These allometric relationships and trade‐offs in carbon allocation across species can improve ecosystem model predictions of grassland species interactions and environmental responses through differences in morphology.

54 ENVIRONMENTAL SCIENCES↗

Utility of near‐surface phenology in estimating productivity and evapotranspiration across diverse ecosystems

Abstract Agroecosystems, which include row crops, pasture, and grass and shrub grazing lands, are sensitive to changes in management, weather, and genetics. To better understand how these systems are responding to changes, we need to improve monitoring and modeling carbon and water dynamics. Vegetation Indices (VIs) are commonly used to estimate gross primary productivity (GPP) and evapotranspiration (ET), but these empirical relationships are often location and crop specific. There is a need to evaluate if VIs can be effective and, more general, predictors of ecosystem processes through time and across different agroecosystems. Near‐surface photographic (red‐green‐blue) images from PhenoCam can be used to calculate the VI green chromatic coordinate (G CC ) and offer a pathway to improve understanding of field‐scale relationships between VIs and GPP and ET. We synthesized observations spanning 76 site‐years across 15 agroecosystem sites with PhenoCam G CC and GPP or ET estimates from eddy covariance (EC) to quantify interannual variability (IAV) in the relationship between GPP and ET and G CC across. We uncovered a high degree of variability in the strength and slopes of the G CC ∼ GPP and ET relationships (R 2 = 0.1 ‐ 0.9) within and across production systems. Overall, G CC is a better predictor of GPP than ET (R 2 = 0.64 and 0.54, respectively), performing best in croplands (R 2 = 0.91). Shrub‐dominated systems exhibit the lowest predictive power of G CC for GPP and ET but have less IAV in slope. We propose that PhenoCam estimates of G CC could provide an alternative approach for predictions of ecosystem processes.

Environmental Sciences & Ecology↗

Predictive links between microbial communities and biological oxygen utilization in the Arctic Ocean

Microbial metabolism influences rates of net community production (NCP), exerting a direct biological control on marine oxygen and carbon fluxes. In the Arctic, it is increasingly important to understand and quantify this process, as ecological and oceanographic conditions shift due to changing climate. Here, we describe potential ecological links between pelagic microbial diversity and an NCP precursor, biological oxygen utilization, using machine learning and paired observations of community structure and metabolic activity from a seasonally and spatially variable transect of the Arctic Ocean (2019–2020 MOSAiC Expedition). Community structure was determined using 16S (prokaryotic) and 18S (eukaryotic) rRNA gene amplicon sequencing, and metabolic activity was derived from ΔO 2 /Ar. Using self-organizing maps, we identified clear successional patterns in observed microbial community structure that were seasonally driven in the upper ocean and vertically stratified with depth. Metabolic activity was also stratified, with a primarily net heterotrophic water column (median −1.5% biological oxygen saturation), excepting periodic oxygen supersaturation (maximum: 13.6%) within the mixed layer. Using DNA sequences as predictor variables, we then constructed a random forest regression model that reliably reconstructed biological oxygen concentrations (root mean squared error = 4.14 μmol kg −1 ). Top predictors from this model were from heterotrophic (bacteria) or potentially mixotrophic (dinoflagellate) taxa. These analyses highlight biologically driven diagnostic tools that can be used to expand biogeochemical datasets and improve the microbial perspectives and metabolisms represented in ecological models of net productivity and carbon flux in a changing Arctic Ocean.

Chamberlain, Emelia J. [Univ. of San Diego, San Di↗

Out With the Old: Empirical Trends in U.S. Land‐Based Wind Turbine Decommissioning and Repowering

A growing number of wind turbines (WTs) across the globe are now reaching or exceeding their expected service lifetime; WT decommissioning is on the rise. Accordingly, questions pertaining to WT end-of-life have risen in importance in policy and practice. Yet, research on the various factors relating to WT decommissioning is relatively sparse. Moreover, the key assumptions underpinning that prior research (e.g., the lifespan of WTs, characteristics of WTs being decommissioned, and whether the site is repowered with new WTs) have never been empirically tested across a large set of decommissioned WTs. Leveraging a uniquely comprehensive and spatially explicit dataset of decommissioned WTs in the United States, this research analyzes spatial, technological, and temporal trends in WT decommissioning and develops a novel predictive model for WT decommissioning. Our analysis pinpoints more than 12,400 WTs that have been fully decommissioned in the United States., the majority of which have been relatively old (> 30 years) and small (< 200 kW). While a WT's age alone is a good predictor of the likelihood of decommissioning, other factors such as the size of the WT and recent performance are also important and significant predictors. Most sites where decommissioning has occurred have seen subsequent repowering, with repowered plants featuring substantially fewer WTs (−86 on average) and higher rated plant capacity (+62 MW on average). Many existing WTs in the U.S. are approaching the end of their expected life with roughly 7500 being 20 or more years old. Findings can help policymakers and stakeholders begin preparing for this potential wave of future decommissioning and repowering.

Decommissioning / End-of-life↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Time-Varying Output Delay Compensation-A Model-Free Approach and its Application on Cooperative On-Ramp Merging

This paper presents a model-free approach to compensate for time-varying output delay in networked control systems. The proposed architecture combines a model-free observer and the Smith predictor. The model-free observer estimates the current state while handling modeling errors and uncertainties of the system. The Smith predictor moves the effect of time delay outside the control closed-loop using the estimated delayed output and the actual output of the plant. The proposed method is applied to a cooperative on-ramp merging problem. First, an ultra-local model predictive control is implemented to provide a computationally efficient online speed planner agnostic to the vehicle dynamics. After that, a model-free observer is designed to estimate the current state. Finally, the proposed architecture is tested against a time-varying output delay with an upper bound of 200 milliseconds. The results demonstrate the effectiveness of the proposed method with improved tracking of intervehicle distance.

Waleed khan, Muhammad [The University of Texas at ↗