Search NASA⌕ Search

SEARCH · Search NASA

Results for “Association Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Development of a Test-Bed for Testing and Refining EarthEn’s Supercritical CO 2 Based Energy Storage System

EarthEn’s energy storage concept leverages supercritical carbon dioxide (sCO 2 ) as a working fluid and relies on compact, high-performance components operating at elevated pressures and temperatures. To accelerate component development and reduce technical risk prior to larger-scale demonstrations, Oak Ridge National Laboratory (ORNL) developed a 100 kW-scale sCO 2 test-bed under a Cooperative Research and Development Agreement with EarthEn (CRADA NO. NFE-24-10050). The objective of the work was to design and construct a flexible experimental facility capable of reproducing key thermodynamic state points and heat-transfer conditions relevant to EarthEn’s thermal energy storage (TES) cycle, with particular emphasis on enabling development and evaluation of next-generation heat exchangers and TES concepts. The test-bed consists of a closed-loop sCO 2 circulation system housed within an open-topped enclosure. In its as-installed configuration, dense-phase sCO 2 is recirculated through a printed circuit recuperator, an electrically heated section, a throttling device used to simulate turbine expansion, and a water-cooled printed circuit heat exchanger that rejects heat to the building chilled-water system before returning to the pump. The pump is driven by a variable frequency drive, enabling controlled adjustment of flow and operating point. A comprehensive instrumentation suite was integrated to support both safe operation and high-quality data collection. Installed sensors include Coriolis flow meters for sCO 2 flow rate and density, resistance temperature detectors and thermocouples distributed throughout the loop (including the heated section and key heat exchanger ports), and pressure transducers for absolute and differential pressure measurements. The facility was designed to support high-pressure (19 MPa nominal) and high-temperature (575°C nominal) operation with credited overpressure protection provided by a rupture disk. Nominal operating conditions were selected to support 100 kW-class testing while maintaining flexibility for non-heated and heated shakedown, control development, and future integration of advanced TES test sections. In parallel with facility development, a system-level thermal-hydraulic model was created using Modelica-based tools to support component sizing, anticipate performance over targeted test conditions, and establish a framework for future model calibration against experimental data. At the conclusion of the project performance period, the facility was in final assembly, and the pressure boundary was nearly completed. However, several practical challenges associated with high-pressure/high-temperature systems and specialized component procurement impacted schedule and prevented initial pump-driven operation and full commissioning within the available resources. This report documents the as-built design, operating capabilities, and instrumentation, and it summarizes key lessons learned related to heater fabrication and testing, first-of-a-kind assembly factors, specialty flange supply constraints, and fill pump corrective actions. Finally, it outlines a phased plan for future commissioning and experimental campaigns, including control and instrumentation shakedown, heater characterization, model calibration, and testing at state points representative of EarthEn’s TES cycle.

25 ENERGY STORAGE↗

Atomistic Simulations of Thermal and Chemical Expansions of PrNi x Co 1‐x O 3‐δ Accelerated by Machine Learning Potentials

The electrodes and solid-state electrolytes in protonic ceramic electrochemical cells (PCECs) experience significant lattice expansions when exposed to high steam concentrations at elevated temperatures. In this paper, phonon calculations based on a new machine learning potential (MLP) are employed to elucidate the volume expansions of the proton-conducting PrNi x Co 1-x O 3-δ (PNC) lattices, manifested under a combined influence of oxygen vacancies (V$^{\cdot\cdot}_O$ ) and proton uptake (OH$^{\cdot}_O$ ) in the bulk at varying Ni/Co occupancies. It is revealed that the Ni/Co occupancy contributes to thermal and chemical expansions differently, where thermal expansions are related to Co occupancy. In contrast, chemical expansions are more closely associated with the Ni occupancy. Both V$^{\cdot\cdot}_O$ and OH$^{\cdot}_O$ lead to higher thermal expansions when compared to the pristine PNC. The temperature increase will negatively impact the hydration-induced chemical expansions. For combined thermal and chemical expansions, it is predicted that the strategies that boost the PCEC's electrochemical performance may harm the electrode–electrolyte interfacial stability, when the Ni occupancy is high, due to severe chemical expansions. Mitigating chemical expansions of the Ni-abundant PNC will benefit the interfacial stability. Finally, the presented computational methods for phonon calculations, based on emerging machine learning interatomic potential techniques are anticipated to have a lasting impact on future PCEC development.

computational prediction↗

Improved Diagnosis of Precipitation Type with LightGBM Machine Learning

Abstract Existing precipitation-type algorithms have difficulty discerning the occurrence of freezing rain and ice pellets. These inherent biases are not only problematic in operational forecasting but also complicate the development of model-based precipitation-type climatologies. To address these issues, this paper introduces a novel light gradient-boosting machine (LightGBM)-based machine learning precipitation-type algorithm that utilizes reanalysis and surface observations. By comparing it with the Bourgouin precipitation-type algorithm as a baseline, we demonstrate that our algorithm improves the critical success index (CSI) for all examined precipitation types. Moreover, when compared with the precipitation-type diagnosis in reanalysis, our algorithm exhibits increased F1 scores for snow, freezing rain, and ice pellets. Subsequently, we utilize the algorithm to compute a freezing-rain climatology over the eastern United States. The resulting climatology pattern aligns well with observations; however, a significant mean bias is observed. We interpret this bias to be influenced by both the algorithm itself and assumptions regarding precipitation processes, which include biases associated with freezing drizzle, precipitation occurrence, and regional synoptic weather patterns. To mitigate the overall bias, we propose increasing the precipitation cutoff from 0.04 to 0.25 mm h −1 , as it better reflects the precision of precipitation observations. This adjustment yields a substantial reduction in the overall bias. Finally, given the strong performance of LightGBM in predicting mixed precipitation episodes, we anticipate that the algorithm can be effectively utilized in operational settings and for diagnosing precipitation types in climate model outputs. Significance Statement Freezing rain can have significant impacts on transportation and infrastructure, making accurate prediction of precipitation types crucial. In this study, we use a machine learning method known as LightGBM to predict precipitation types. We show that the new algorithm performs better than the existing methods for all precipitation types examined. Additionally, we compute a freezing-rain climatology over the eastern United States. Although the resulting climatology pattern corresponds well to observations, the algorithm overpredicts freezing-rain occurrence. We argue that this bias can be substantially reduced by increasing the precipitation cutoff from 0.04 to 0.25 mm h −1 . Overall, this work highlights the potential of the LightGBM algorithm for both weather forecasting and diagnosing precipitation types in climate models.

Meteorology & Atmospheric Sciences↗

Model-independent measurement of the Higgs boson associated production with two jets and decaying to a pair of W bosons in proton-proton collisions at $\sqrt{s}=13$ TeV

A model-independent measurement of the differential production cross section of the Higgs boson decaying into a pair of W bosons, with a final state including two jets produced in association, is presented. In the analysis, events are selected in which the decay products of the two W bosons consist of an electron, a muon, and missing transverse momentum. The model independence of the measurement is maximized by employing a discriminating variable, developed through machine learning, that is agnostic to the signal hypothesis. The analysis is based on proton-proton collision data at $\sqrt{s}=13$ TeV collected with the CMS detector from 2016–2018, corresponding to an integrated luminosity of 138 fb −1 . The production cross section is measured as a function of the difference in azimuthal angle between the two jets. The differential cross section measurements are used to constrain Higgs boson couplings within the standard model effective field theory framework.

Hadron-Hadron Scattering↗

Design and optimization of a modular hydrogen-based integrated energy system to maximize revenue via nuclear-renewable sources

Here, this paper demonstrates a novel modular distributed framework that uses optimal energy-dispatching strategies to enable greater flexibility and profitability in nuclear-renewable integrated energy systems (NR-IES). Hydrogen is used as a commodity in this framework since its production can improve grid stability and system operational flexibility, decarbonize heavy industry, and create an additional revenue stream for electricity generators, particularly nuclear power plants with high operational expenses. The proposed solution addresses the challenges associated with merging multiple software and services from various domains by using functional mock-up units (FMU) to co-simulate diverse subsystems designed in various platforms. The tightly coupled integrated energy system (IES) is optimized to maximize revenue by utilizing the deep reinforcement learning (DRL) technique to make smart dispatching decisions based on variable electricity prices and the availability of renewable energy. Proximal policy optimization (PPO) algorithm is used in training and testing the DRL agent. Over a period of 120 days, the proposed hydrogen-based IES framework showed about 10% revenue boost compared to a non-hydrogen generating baseline IES while also providing an easily-adoptable framework which can help to improve the flexibility of future generation nuclear power plants.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

JAX-CanVeg: A Differentiable Land Surface Model

Land surface models consider the exchange of water, energy, and carbon along the soil-canopy-atmosphere continuum, which is challenging to model due to their complex interdependency and associated challenges in representing and parameterizing them. Differentiable modeling provides a new opportunity to capture these complex interactions by seamlessly hybridizing process-based models with deep neural networks (DNNs), benefiting both worlds, that is, the physical interpretation of process-based models and the learning power of DNNs. Here, we developed a differentiable land model, JAX-CanVeg. The new model builds on the legacy CanVeg by incorporating advanced functionalities through JAX in the graphic processing unit support, automatic differentiation, and integration with DNNs. We demonstrated JAX-CanVeg's hybrid modeling capability by applying the model at four flux tower sites with varying aridity. To this end, we developed a hybrid version of the Ball-Berry equation that emulates the water stress impact on stomatal closure to explore the capability of the hybrid model in (a) improving the simulations of latent heat fluxes (LE) and net ecosystem exchange (NEE), (b) improving the optimization trade-off when learning observations of both LE and NEE, and (c) benefiting a multi-layer canopy model setup. Our results show that the proposed hybrid model improved the simulations of LE and NEE at all sites, with an improved optimization trade-off over the process-based model. Additionally, the multi-layer canopy set benefited hybrid modeling at some sites. Anchored in differentiable modeling, our study provides a new avenue for modeling land-atmosphere interactions by leveraging the benefits of both data-driven learning and process-based modeling.

54 ENVIRONMENTAL SCIENCES↗

Post-2026 Environmental Impact Statement Rate Analysis for the Colorado River Storage Project

The Glen Canyon Dam (GCD) is a principal power-generating asset within the Colorado River Storage Project (CRSP), accounting for approximately 70–80% of CRSP power production over the past two decades. In June 2023, the U.S. Bureau of Reclamation (Reclamation) issued a Notice of Intent to prepare an Environmental Impact Statement (EIS) outlining operational guidelines and strategies for Colorado River Basin reservoirs after 2026 (Reclamation, 2023). Power generation is among CRSP’s statutory purposes under the Colorado River Storage Project Act of 1956 (U.S. Congress, 1956). Assessing how alternative policy frameworks affect CRSP power production and the resulting electricity rates for U.S. customers is therefore essential to inform decision-making. This report evaluates projected electricity rates and the market value of electricity from the Western Area Power Administration (WAPA) CRSP GCD under multiple post-2026 policy scenarios to support Reclamation’s EIS development. Results from advanced econometric and machine learning models indicate that the Enhanced Coordination alternative (EnhanCoor), Maximum Operational Flexibility alternative (CCA), and Supply Driven - 55 alternative (SD55) alternatives yield more favorable hydropower generation and capacity outcomes, which are objectives outlined in Reclamation’s documentations (Reclamation, 2007; Reclamation, 2016). Specifically, these scenarios are associated with higher electricity production, lower projected rate trajectories, and greater economic value to the U.S. power system from CRSP generation. The remaining five scenarios generally produce lower generation, higher rate trajectories, and reduced long-term market values.

13 HYDRO ENERGY↗

SSR APPLIED – Automated Power Plants: Intelligent, Efficient and Digitised (V.2)

This report describes work undertaken during the SSR APPLIED project. The focus of the project has been on the development of digital twins to de-risk licensing of improved operating and maintenance practices. The operation of a bespoke flowing separate effects molten salt loop at ANL, with realistic temperature gradients, will provide invaluable data for computer codes validation. Three digital twins of aspects of the SSR-W have been successfully developed using ANL expertise and software. These digital twins have demonstrated optimization of the fuel cycle, the ability to model transients using an integrated coupled neutronic – thermal-hydraulic model with a model of the fuel expansion feedback so important to the inherent safety of the SSR-W. Advanced machine learning techniques have been developed and demonstrated for optimization of heat exchanger operation and maintenance.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Search for CP violation in events with top quarks and Z bosons at $\sqrt{s}$ = 13 and 13.6 TeV

A search for the violation of the charge-parity (CP) symmetry in the production of top quarks in association with Z bosons is presented, using events with at least three charged leptons and additional jets. The search is performed in a sample of proton-proton collision data collected by the CMS experiment at the CERN LHC in 2016–2018 at a center-of-mass energy of 13 TeV and in 2022 at 13.6 TeV, corresponding to a total integrated luminosity of 173 fb –1 . For the first time in this final state, observables that are odd under the CP transformation are employed. Also for the first time, physics-informed machine-learning techniques are used to construct these observables. While for standard model (SM) processes the distributions of these observables are predicted to be symmetric around zero, CP-violating modifications of the SM would introduce asymmetries. Two CP-odd operators $\mathcal{O}$$^{I}_{tW}$ and $\mathcal{O}$$^{I}_{tZ}$ in the SM effective field theory are considered that may modify the interactions between top quarks and electroweak bosons. The obtained results are consistent with the SM prediction within two standard deviations, and exclusion limits on the associated Wilson coefficients of –2.7 < $c$$^{I}_{tW}$ < 2.5 and –0.2 < $c$$^{I}_{tZ}$ < 2.0 and are set at 95 % confidence level. The largest discrepancy is observed in $c$$^{I}_{tZ}$ where data is consistent with positive values, with an observed local significance with respect to the SM hypothesis of 2.5 standard deviations, when only linear terms are considered.

CMS↗

Evaluating multistation phase picking algorithm phase neural operator (PhaseNO) on local seismic networks

Reliable automatic phase picking is important for many seismic applications. With the development of machine learning approaches, many algorithms are proposed, evaluated and applied to different areas. Many of these algorithms are single station based, while recent proposed methods start to combine surrounding stations into consideration in the problem of phase picking. Among these algorithms, the phase neural operator (PhaseNO) shows promising results on regional data sets comparing to existing algorithms. But there are many use cases for the local seismic networks in our community, therefore in this paper we evaluate the performance of PhaseNO on four different local data sets and compare the results to PhaseNet and EQTransformer. We used both individual phase picking metrics as well as association metrics to illustrate the performance of PhaseNO. By manually reviewing the newly detected events, we find that the PhaseNO model outperforms the single station-based approaches in the local-scale use cases due to its consideration of coherent signals from multiple stations. We also explored PhaseNO’s behaviours when only using one station, as well as gradually increasing the number of stations in the seismic network to better understand its behaviour. Overall, using the off-the-shelf machine learning based phase pickers, PhaseNO demonstrated its good performance on local-scale seismic networks.

58 GEOSCIENCES↗

Risk-Aware Measurement Synchronization and Recovery for DSSE With Heterogeneous Data Sources

Power distribution systems are increasingly integrating heterogeneous sensors with varying data reporting rates and types, which pose challenges to achieving observability at the desired temporal resolution of distribution system state estimation (DSSE). Multisensor failures caused by extreme events exacerbate these issues, introducing substantial uncertainties into DSSE. This article proposes a novel solution to these challenges by ensuring high-resolution system observability despite heterogeneous data sources and multisensor failures. First, a deep learning architecture combining long short-term memory (LSTM) and graph convolutional network (GCN) is employed to synchronize meters with different reporting rates, aiming to achieve system observability. A random-walk-model-based approach is introduced to generate pseudo-measurements while properly characterizing their uncertainties under multisensor failures. Finally, a disaster-risk-informed observability metric (RiOM) is defined to quantify the uncertainty associated with state estimation results. The proposed framework offers deeper insights into the system observability on the fly compared with conventional analysis. The effectiveness of the framework is demonstrated on an IEEE standard test case and a large-scale real-world distribution feeder in mid-Minnesota in the U.S.

97 MATHEMATICS AND COMPUTING↗

Forecasting Solar Photovoltaic Power Production: A Comprehensive Review and Innovative Data-Driven Modeling Framework

The intermittent and stochastic nature of Renewable Energy Sources (RESs) necessitates accurate power production prediction for effective scheduling and grid management. This paper presents a comprehensive review conducted with reference to a pioneering, comprehensive, and data-driven framework proposed for solar Photovoltaic (PV) power generation prediction. The systematic and integrating framework comprises three main phases carried out by seven main comprehensive modules for addressing numerous practical difficulties of the prediction task: phase I handles the aspects related to data acquisition (module 1) and manipulation (module 2) in preparation for the development of the prediction scheme; phase II tackles the aspects associated with the development of the prediction model (module 3) and the assessment of its accuracy (module 4), including the quantification of the uncertainty (module 5); and phase III evolves towards enhancing the prediction accuracy by incorporating aspects of context change detection (module 6) and incremental learning when new data become available (module 7). This framework adeptly addresses all facets of solar PV power production prediction, bridging existing gaps and offering a comprehensive solution to inherent challenges. By seamlessly integrating these elements, our approach stands as a robust and versatile tool for enhancing the precision of solar PV power prediction in real-world applications.

14 SOLAR ENERGY↗

Huge ensembles – Part 1: Design of ensemble weather forecasts using spherical Fourier neural operators

Abstract. Simulating low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computational constraints, it is infeasible to generate huge ensembles (comprised of 1000–10 000 members) with traditional, physics-based numerical models. In this two-part paper, we replace traditional numerical simulations with machine learning (ML) to generate hindcasts of huge ensembles. In Part 1, we construct an ensemble weather forecasting system based on spherical Fourier neural operators (SFNOs), and we discuss important design decisions for constructing such an ensemble. The ensemble represents model uncertainty through perturbed-parameter techniques, and it represents initial condition uncertainty through bred vectors, which sample the fastest-growing modes of the forecast. Using the European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (IFS) as a baseline, we develop an evaluation pipeline composed of mean, spectral, and extreme diagnostics. With large-scale, distributed SFNOs with 1.1 billion learned parameters, we achieve calibrated probabilistic forecasts. As the trajectories of the individual members diverge, the ML ensemble mean spectra degrade with lead time, consistent with physical expectations. However, the individual ensemble members' spectra stay constant with lead time. Therefore, these members simulate realistic weather states during the rollout, and the ML ensemble passes a crucial spectral test in the literature. The IFS and ML ensembles have similar extreme forecast indices, and we show that the ML extreme weather forecasts are reliable and discriminating. These diagnostics ensure that the ensemble can reliably simulate the time evolution of the atmosphere, including low-likelihood high-impact extremes. In Part 2, we generate a huge ensemble initialized each day in summer 2023, and we characterize the simulations of extremes.

Mahesh, Ankur↗

On Road vs. Off Road Low Load Cycle Comparison

Reducing criteria pollutants while reducing greenhouse gases is an active area of research for commercial on-road vehicles as well as for off-road machines. The heavy duty on-road sector has moved to reducing NOx by 82.5% compared to 2010 regulations while increasing the engine useful life from 435,000 to 650,000 miles by 2027 in the United States (US). An additional certification cycle, the Low Load Cycle (LLC), has been added focusing on part load operation having tight NOx emissions levels. In addition to NOx, the total CO2 emissions from the vehicle will also be reduced for various model years. The off-road market is following with a 90% NOx reduction target compared to Tier 4 Final for 130-560 kW engines along with greenhouse gas targets that are still being established. The off-road market will also need to certify with a Low Load Application Cycle (LLAC), a version of which was proposed for evaluation in 2021. Since the LLAC has not been finalized, this study is being conducted to compare and contrast the LLC for on-road with the LLAC for off-road as there might be some shared learnings. A US off-road production 2023 Fiat Powertrain 13L engine and aftertreatment system was chosen for this work. This engine is used in production for both off-road and on-road products, so it is a good choice for this study. The associated off-road aftertreatment system was aged for more relevant comparisons. The engine calibration was not altered for either of the low load cycles. This study shows that the cycles are quite different in nature as the market needs are different. The LLC includes a large fraction of operation at idle and lower speeds, representing products that use the engine primarily for motive power, where lower vehicle speed means a lower engine speed and load. The LLAC has more time and load spent at high speeds and slightly higher loads. The off-road products represented by this cycle often use the engine to drive auxiliary equipment which means higher parasitic loads and hand/fixed throttle. The comparison will include the use profiles, tailpipe NOx and greenhouse gas emissions (CO2, N2O).

McCarthy, James↗

Predictive models of the genetic bases underlying budding yeast fitness in multiple environments

Abstract The ability of organisms to adapt and survive depends on the effects of genes and the environment on fitness. However, the multigenic nature of fitness and genotype-by-environment interactions hinder our understanding of the genetic basis of fitness. Here, we established fitness prediction models for 35 environments using machine learning and existing fitness data and different genetic variant types for a Saccharomyces cerevisiae population. Models revealed that the predictive ability of genetic variants varied across environments, with copy number variants explaining the majority of fitness variation in most cases. Model interpretation showed that different variant types identified distinct gene sets associated with predictive variants. These gene sets were significantly enriched in experimentally validated genes affecting fitness in only a subset of environments, indicating that many genes influencing fitness remain unexplored. Notably, non-experimentally validated genes were more important than validated ones for fitness predictions. Gene contributions to predictions were both isolate- and environment-dependent, pointing to gene-by-gene and gene-by-environment interactions. Furthermore, models uncovered experimentally validated and novel candidate genetic interactions for a well-characterized stress, the fungicide benomyl. These findings highlight the feasibility of identifying the genetic basis of fitness by using different genetic variant types and offer novel targets for future functional analysis.

DNA copy number variations↗

Predicting Pulsed-Laser Deposition SrTiO 3 Homoepitaxy Growth Dynamics Using High-Speed Reflection High-Energy Electron Diffraction

Pulsed-laser deposition (PLD) is a powerful technique for growing complex oxides with controlled stoichiometry. To understand growth dynamics therein, it is common to leverage in situ spectroscopies, such as reflection high-energy electron diffraction (RHEED), to monitor surface crystallinity. Most commercial systems rely on video-rate cameras operating at 60-120 Hz that lack sufficient temporal resolution to capture growth dynamics at practical deposition frequencies. Here, a high-speed platform to record in situ dynamics via RHEED at >500 Hz is implemented. An open-source analysis package is designed to fit diffraction spots to 2D Gaussians, allowing single-pulse surface reconstruction kinetics extraction. Using homoepitaxially deposited (001)-oriented SrTiO 3 as a model system, we demonstrate how high-speed RHEED can provide real-time insight into growth processes obscured by slower acquisition systems. By fitting the single-pulse intensity to a set of exponential functions, we observe changes in the characteristic decay time and mechanism correlated to the substrate step width and surface termination. We observe distinct surface effects, with diffraction intensity decaying on lower-energy TiO 2 -terminated surfaces and stabilizing on SrO- or mixed-terminated surfaces. Similarly, using an exponential model, the extracted characteristic time of adatom deposition decreases with increased density of bonding sites associated with mixed termination and narrower step widths. Ultimately, this work shows how increasing RHEED temporal resolution can uncover new insights into growth processes, with practical implications for the design and control of PLD processes. This experimental platform provides new capabilities to enable data-driven machine learning analysis and autonomous control systems to enhance the complexity and fecundity of PLD.

(SrO)↗

Validation of new and existing methods for time-domain simulations of turbulence and loads

We seek to obtain a second-by-second match between the simulated and measured structural loads of a utility-scale wind turbine. To obtain the one-to-one load simulations, we start with the furthest upstream component of the modeling chain: the turbulent inflow. We consider new and existing methods to generate constrained-turbulence flow fields. The new method is based on large-eddy simulations (LES) and machine learning (ML). The existing methods include Kaimal-based TurbSim and the superstatistical wind field model. The inflow measurements used to constrain these simulations are obtained with a nacelle-mounted scanning lidar. We compare the flow fields for the different inflow simulation approaches and validate their associated load predictions against measurements collected in the Rotor Aero-dynamics, Aeroelastics, and Wake (RAAW) field campaign. We find that the rotor-position control developed for this study is key in enabling the time match between measurements and simulations. When this control approach is used, the load simulation performance tracks with the inflow simulation fidelity, with LES+ML yielding errors ≤ 4% for the damage-equivalent loads of flapwise bending moment, and tower fore-aft bending moments.

17 WIND ENERGY↗

W2VPCA: A Machine Learning Method for Measuring Attitudes With Natural Language

Company strategy influences many decisions in freight transportation. Behavioral models of company decision-making therefore could benefit from including strategy variables. However, strategy is difficult to observe and quantify. Attitudinal surveys of company executives can be used to collect measurements of latent strategy to use in quantitative models. However, surveys are costly and burdensome. Text mining methods to collect measurements overcome these issues somewhat, but typically require manual intervention and ignore the context of words, which can be problematic. This study introduces a new machine learning method to generate strategy measurement data from existing big text data. The new method, called W2VPCA, combines Natural Language Processing and Principal Components Analysis. W2VPCA produces measurement data that serve as quantitative indicators of latent strategy in behavioral models. W2VPCA is unsupervised, data-driven, and uses information on word context. We apply W2VPCA to generate measurements of latent strategies using readily available, large-scale text data: annual company reports. The empirical measurements are used successfully to associate two latent strategies, one focusing on distribution and the other on products, with truck fleet and distribution center outsourcing decisions. The main empirical outcome is that the W2VPCA measurements outperform Bag-of-Words measurements in a psychometric analysis of latent firm strategies. While this study focuses on freight behavioral models, W2VPCA may also have applications in behavioral modeling in other domains.

97 MATHEMATICS AND COMPUTING↗