Search NASA⌕ Search

SEARCH · Search NASA

Results for “data model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

A Framework for Identifying Building Energy Models of Localized Utility Service Areas Using Smart Meter Data

Bottom-up load modeling of buildings offers a versatile approach to simulating baseline demand and scenarios of future technology evolution and adoption at the individual building level. This capability is essential to understanding how future load shapes may change with the adoption of electric equipment and vehicles, particularly as it relates to grid planning and infrastructure investments. Traditionally, grid planning techniques have used historical load data to predict future load and infrastructure needs. However, with the anticipated rise in adoption of electrification technologies such as heat pumps and electric vehicles, historical data become less reliable predictors of the future. By employing ResStock, a high-fidelity building stock modeling tool, we can fine-tune electrification scenarios and aggregate models to represent varying geographic resolutions of the grid system, while considering the underlying features of homes. This may enable a more accurate and responsive approach to anticipate and plan for the evolving landscape of energy demands. We present a new framework that leverages building stock energy modeling to identify building models that align with the load shapes and housing attributes of buildings with AMI data. This approach applies two model layers: (1) a classification step that identifies the presence of air conditioning, electric heating, and electric water heating, and (2) an optimization routine that identifies building energy models aligning with load profile data from advanced metering infrastructure meters. This report demonstrates one approach to deploying this framework, and presents results for three test cases that use both modeled and AMI data to assess performance. For a test case using AMI data in Fort Collins, Colorado, we observed a median monthly electricity load CV-RMSE of 16.6%, and a top ten daily heating and cooling median absolute percent error of 7.7% and 8.3%, respectively. For each AMI meter, we identify a set of potential energy models so that downstream use-cases can account for uncertainty driven by variability of baseline technologies and occupant behavior, which impact the response to electrification and energy efficiency scenarios. Our results indicate that ResStock has potential as a scalable solution for modeling residential energy demand at local grid resolutions. Its performance depends on location-specific factors, underlying building characteristics, and the level of aggregation, offering a path towards more precise and adaptive distribution grid planning for the evolving energy landscape.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Assessing Impacts of Waves on Hub-Height Winds off the U.S. West Coast Using Lidar Buoys and Coupled Modeling Approaches

Given the importance of offshore wind energy development to the U.S. clean energy targets, it is vital to be able to characterize the wind resource in that environment accurately. Toward that end, two Bureau of Ocean Energy Management buoys equipped with Doppler lidar are being maintained by Pacific Northwest National Laboratory on behalf of the Department of Energy and deployed to regions of potential offshore wind development. In addition to standard meteorological and oceanographic measurements, the buoys document the wind profile between about 40 m and 250 m above the sea surface through Doppler lidar retrievals. After a multiyear deployment of two buoys along the U.S. East Coast, the buoys were redeployed to the U.S. West coast from 2020 – 2022 to locations near the Humboldt and Morro Bay lease areas. The buoys provide nearly continuous, multiyear datasets that can be used to evaluate predictions of hub-height (~100 m) wind speed for standard atmospheric models in the region. In the absence of measurements at the study site, offshore wind developers rely on model-based data to assess site conditions. Potential sources of model error in this environment include under-resolution or misrepresentation of coastal topographically forced flows, marine boundary layer dynamics and the evolution of their associated cloud and turbulence fields, the role of upwelling and other currents on surface heat fluxes into the boundary layer, and the impact of wave fields on surface momentum fluxes and thus the wind speed profile. In particular, most predictive models of wind speed do not predict wave fields at all, relying on parameterizations to represent their effects. In thus study, we focus on evaluating the role of wind / wave interactions on modeled hub-height wind speed and error by using the Coupled Ocean–Atmosphere–Wave–Sediment–Transport Modeling System to capture two-way interactions between an atmospheric model (Weather Research and Forecasting (WRF)) and a wave model (WAVEWATCHIII (WW3)) and compare to both stand-alone WRF and one-way coupled WRF / WW3 configurations. Our approach is similar to that used in Gaudet et al. (2022) to evaluate wind / wave coupling over the U.S. East Coast, but applied to the very different environment of the U.S. West Coast. We show examples for two cases, a cold-season frontal case and a warm-season low-level jet case. We find that wind / wave coupling makes little impact on model error for these cases at the location of the lidar buoys, for which other misrepresentations of model physics seems to be responsible for model-observation discrepancies. However, domain-wide evaluations, which also make use of the National Buoy Data Center network, show that a two-way coupling approach is less prone to systematic errors in the hub-height wind field than the one-way coupled approach. WRF resolution of kilometer-scale or less is needed to properly capture the sharp wind speed gradients that can be found along the coastline, and WW3 simulations driven by the downscaled WRF produce better bulk and spectral wave fields when compared to observations. Implications of the results for wind resource characterization are then discussed.

17 WIND ENERGY↗

Neural Network‐Based Methods for Ocean Surface Wave Measurement Using Submarine Distributed Acoustic Sensing (DAS)

Two new data-driven models for estimating ocean surface waves from distributed acoustic sensing (DAS) submarine cable strain rate are developed using supervised machine learning on a 10-day data set collected offshore of Oliktok Point, Alaska. The new models were trained on target data from seafloor pressure moorings at three sites spaced evenly along 27.1 km of cable and were benchmarked against an empirical transfer function method previously used to estimate waves from DAS. A model which uses convolutional neural networks to transform 2-km frequency-wavenumber strain spectra to seafloor pressure spectra outperforms the benchmark in wave height prediction (RMSE of 0.15 vs. 0.41 m) and period prediction (0.29 vs. 0.37 s) when evaluated on a held-out test data set. When applied to a DAS data set collected on the same cable 2 years prior, the CNN-based model maintained similar significant wave height performance (RMSE = 0.23 m) relative to available satellite altimetry data. A two-hidden-layer, fully connected neural network which transforms 1-D strain spectra to seafloor pressure spectra also outperforms the benchmark in wave height prediction (RMSE of 0.19 vs. 0.41 m), but does not generalize as well to the prior data. Regression-based machine learning is useful for estimating waves from DAS data when the pressure-strain relationship varies temporally and spatially across different wave conditions. Models can be applied to DAS data to measure waves with higher spatial resolution and longer temporal coverage than traditional methods, which often measure waves only at a single point.

Davis, Jacob R. [Univ. of Washington, Seattle, WA ↗

In Situ Data Analysis Through Physics-informed Tensor Decompositions (LDRD Final Report)

We introduce a new low-dimensional model of high-dimensional numerical simulation data based on low-rank tensor decompositions. Our new model aims to minimize differences between the model data and simulation data as well as functions of the model data and functions of the simulation data. This novel approach to dimensionality reduction of simulation data provides a means of directly incorporating quantities of interests and invariants associated with conservation principles associated with the simulation data into the low-dimensional model, thus enabling more accurate analysis of the simulation without requiring access to the full set of high-dimensional data. Computational results of applying this approach to two standard low-rank tensor decompositions of data arising from simulation of combustion and plasma physics are presented.

97 MATHEMATICS AND COMPUTING↗

Quasi-Classical Trajectory Calculation of Rate Constants Using an Ab Initio Trained Machine Learning Model (aML-MD) with Multifidelity Data

Machine learning (ML) provides a great opportunity for the construction of models with improved accuracy in classical molecular dynamics (MD). However, the accuracy of a ML trained model is limited by the quality and quantity of the training data. Generating large sets of accurate ab initio training data can require significant computational resources. Furthermore, inconsistent or incompatible data with different accuracies obtained using different methods may lead to biased or unreliable ML models that do not accurately represent the underlying physics. Recently, transfer learning showed its potential for avoiding these problems as well as for improving the accuracy, efficiency, and generalization of ML models using multifidelity data. In this work, ab initio trained ML-based MD (aML-MD) models are developed through transfer learning using DFT and multireference data from multiple sources with varying accuracy within the Deep Potential MD framework. Further, the accuracy of the force field is demonstrated by calculating rate constants for the H + HO 2 → H 2 + 3 O 2 reaction using quasi-classical trajectories. We show that the aML-MD model with transfer learning can accurately predict the rate constants while reducing the computational cost by more than five times compared to the use of more expensive quantum chemistry training data sets. Hence, the aML-MD model with transfer learning shows great potential in using multifidelity data to reduce the computational cost involved in generating the training set for these potentials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Streamlining Ocean Dynamics Modeling with Fourier Neural Operators: A Multiobjective Hyperparameter and Architecture Optimization Approach

Training an effective deep learning model to learn ocean processes involves careful choices of various hyperparameters. We leverage DeepHyper’s advanced search algorithms for multiobjective optimization, streamlining the development of neural networks tailored for ocean modeling. The focus is on optimizing Fourier neural operators (FNOs), a data-driven model capable of simulating complex ocean behaviors. Selecting the correct model and tuning the hyperparameters are challenging tasks, requiring much effort to ensure model accuracy. DeepHyper allows efficient exploration of hyperparameters associated with data preprocessing, FNO architecture-related hyperparameters, and various model training strategies. We aim to obtain an optimal set of hyperparameters leading to the most performant model. Moreover, on top of the commonly used mean squared error for model training, we propose adopting the negative anomaly correlation coefficient as the additional loss term to improve model performance and investigate the potential trade-off between the two terms. The numerical experiments show that the optimal set of hyperparameters enhanced model performance in single timestepping forecasting and greatly exceeded the baseline configuration in the autoregressive rollout for long-horizon forecasting up to 30 days. Utilizing DeepHyper, we demonstrate an approach to enhance the use of FNO in ocean dynamics forecasting, offering a scalable solution with improved precision.

97 MATHEMATICS AND COMPUTING↗

Machine learning-powered data cleaning for LEGEND: a semi-supervised approach using affinity propagation and support vector machines

Neutrinoless double-beta decay ($0\nu\beta\beta$) is a rare nuclear process that, if observed, will provide insight into the nature of neutrinos and help explain the matter-antimatter asymmetry in the Universe. The large enriched germanium experiment for neutrinoless double-beta decay (LEGEND) will operate in two phases to search for $0\nu\beta\beta$. The first (second) stage will employ 200 (1000) kg of High-Purity Germanium (HPGe) enriched in 76 Ge to achieve a half-life sensitivity of 10 27 (10 28 ) years. In this study, we present a semi-supervised data-driven approach to remove non-physical events captured by HPGe detectors powered by a novel artificial intelligence model. We utilize affinity propagation to cluster waveform signals based on their shape and a support vector machine to classify them into different categories. We train, optimize, and test our model on data taken from a natural abundance HPGe detector installed in the Full Chain Test experimental stand at the University of North Carolina at Chapel Hill. We demonstrate that our model yields a maximum sacrifice of physics events of $0.024 ^{+0.004}_{-0.003} \%$ after data cleaning. Our model is being used to accelerate data cleaning development for LEGEND-200 and will serve to improve data cleaning procedures for LEGEND-1000.

artificial intelligence↗

A ModEx Framework for Watershed Subsurface Investigation With Limited Geophysical Data Using Machine Learning and Hydrologic Modeling

Abstract Subsurface heterogeneity influences watershed hydrology strongly but remains difficult to characterize at catchment scales with sparse and costly field data. Geophysical surveys such as electromagnetic induction (EMI) provide local spatial subsurface images yet scaling them to watershed scales and converting EMI‐derived resistivity into hydraulic properties remains a challenge. We present a Model–Experiment (ModEx) framework that integrates limited EMI data with machine learning (ML) and hydrologic modeling to improve process representation and guide field investigations. Sparse EMI surveys were scaled to the catchment scale using a Random Forest model, and the resulting resistivity fields were combined with nearby borehole constraints to parameterize a hydrologic model. The EMI‐informed hydrological simulations improved predictions of streamflow sustained by subsurface flow and shallow saturation patterns. By combining EMI data and ML with hydrologic modeling, the ModEx framework guides future subsurface surveys, providing a transferable and efficient strategy for data–model integration across diverse watersheds. Plain Language Summary Mapping the underground network of soil and rock that controls water is essential for predicting floods and droughts, but seeing underground is difficult and expensive. We cannot drill everywhere, so scientists use geophysical tools to scan broad areas. There are two key challenges: these geophysical scans are often sparse across the whole watershed, and the geophysical data is hard to translate into water‐related properties. We used artificial intelligence to solve these problems. We taught a computer to find patterns linking the limited geophysical data to the land surface properties. This allowed it to fill in the gaps and create a complete, useful subsurface map for the entire watershed. This new map improves hydrologic simulations, leading to more accurate predictions of water movement in the watershed. It also helps scientists build better models with less data and generates a priority map showing where to measure next, making future investigations more efficient. Key Points Limited EMI scaled with ML improves catchment‐scale subsurface parameterization for hydrologic models The framework integrates hydrologic modeling with limited geophysical data to support subsurface investigation design ModEx framework offers a transferable data–model integration strategy that quantifies and reduces uncertainty guiding watershed studies

Chen, Hang↗

Evaluation of the EarthSHAB Stratospheric Solar Hot Air Balloon Flight Prediction Model Using Balloon Trajectory Data

Abstract The heliotrope is a solar balloon design which is constructed out of painter’s plastic, and the exterior is coated in charcoal powder. Darkening the plastic gives the balloon a high solar absorptance, which allows it to ascend into the lower stratosphere and float for hours at a time. The balloons have previously been used to lift scientific instruments into the stratosphere to study chemical explosions, earthquakes, and stratospheric aerosols. They have also been proposed as a platform for planetary exploration. Flight predictions are crucial to preflight planning to reduce safety risks and meet flight objectives. However, there exists a wide range of possible flight paths due to varying environmental conditions and solar balloon configurations. EarthSHAB is one such software that was designed to support flight planning using the weather forecasts and balloon properties to predict the flight path of a solar balloon. We compare EarthSHAB-simulated flight paths to a set of observed flight paths for the 3.5-m diameter heliotrope design called the “Cloudskimmer.” Using the criteria that the modeled paths must fall within 5% of the observations to be considered successful, we found that EarthSHAB successfully predicted the Cloudskimmer ascent rate and average float altitude 10% and 90% of the time, respectively. We also found that the average difference in the observed and predicted landing locations was 97 km and landing times were 54 ± 38 min. Significant deviations between the observed and predicted ascent rates and excursions at float were found to be associated with heavy payloads and convective cloud development, respectively. Significance Statement Solar balloons are used to lift scientific instruments into the lower stratosphere for hours at a time to study chemical explosions, earthquakes, stratospheric aerosols, and more. The flight paths of solar balloons can be difficult to predict due to variability in their design and surrounding environment. We evaluate the accuracy of EarthSHAB, a software that predicts the altitude profile and horizontal trajectory of a solar balloon using inputs such as the weather, balloon size, and balloon mass. Our results suggest that for the balloon design used in this study, EarthSHAB is best suited for modeling the behavior of balloons with lightweight payloads that do not fly within or directly above clouds.

Lien, Jessica M. [Sandia National Laboratories, Al↗

An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry

Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy.

Wang, Tianle [Brookhaven National Laboratory (BNL)↗

A database and meta-analysis on the performance of exploding pusher implosions conducted at OMEGA

A database of 222 exploding pusher implosions conducted at the OMEGA Laser Facility is presented. The dataset consists of glass-shell capsules filled with varying pressures of D 2 , T 2 , and 3 He, which were imploded using square laser pulses with intensities ranging from 1 to 1 × 10 15 W/cm 2 . The database includes measurements of bang times, ion temperatures, and yields from the DD, D 3 He, and DT fusion reactions. A semi-analytic exploding pusher model is introduced, which effectively captures the observed trends in the data. This model predicts that the measurements scale according to a power-law relation based on the initial capsule and laser conditions. A generalized power-law scaling relation is directly fit to each dataset, providing a useful interpolation of the entire database. Overall, the database provides a valuable resource to estimating bang times, temperatures, and yields for the design of future experiments. Additionally, it provides a diverse set of data for validating more advanced implosion physics models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Weak-form latent space dynamics identification

Recent work in data-driven modeling has demonstrated that a weak formulation of model equations enhances the noise robustness of a wide range of computational methods. In this paper, we demonstrate the power of the weak form to enhance the LaSDI (Latent Space Dynamics Identification) algorithm, a recently developed data-driven reduced order modeling technique. We introduce a weak form-based version WLaSDI (Weak-form Latent Space Dynamics Identification). WLaSDI first compresses data, then projects onto the test functions and learns the local latent space models. Notably, WLaSDI demonstrates significantly enhanced robustness to noise. With WLaSDI, the local latent space is obtained using weak-form equation learning techniques. Compared to the standard sparse identification of nonlinear dynamics (SINDy) used in LaSDI, the variance reduction of the weak form guarantees a robust and precise latent space recovery, hence allowing for a fast, robust, and accurate simulation. We demonstrate the efficacy of WLaSDI vs. LaSDI on several common benchmark examples including viscid and inviscid Burgers', radial advection, and heat conduction. For instance, in the case of 1D inviscid Burgers' simulations with the addition of up to 100% Gaussian white noise, the relative error remains consistently below 6% for WLaSDI, while it can exceed 10,000% for LaSDI. Similarly, for radial advection simulations, the relative errors stay below 15% for WLaSDI, in stark contrast to the potential errors of up to 10,000% with LaSDI. Moreover, speedups of several orders of magnitude can be obtained with WLaSDI. For example applying WLaSDI to 1D Burgers' yields a 140X speedup compared to the corresponding full order model.

97 MATHEMATICS AND COMPUTING↗

Distribution Substation Planning Toolkit (dsp-toolkit) v1.0

The Distribution Substation Planning Toolkit (DSP Toolkit) is a software suite designed to streamline the planning and optimization of distribution substations. This toolkit offers a comprehensive set of tools and APIs for data curation, short-term electric load forecasting, and weather-sensitive load adjustment, making it an essential resource for utility companies, engineers, and researchers. Features • Data Preprocessing and Curation: Efficiently manage and preprocess large datasets to ensure high-quality input for analysis. • Short-Term Load Forecasting: Utilize data-driven models to predict short-term electric loads accurately. • Weather-Sensitive Modeling: Automatically adjust load forecasts based on weather data to predict future peak demands more precisely. Uses The DSP Toolkit is ideal for planning and optimizing distribution substations, providing a user-friendly interface and comprehensive documentation. It is suitable for both novice and experienced users, facilitating efficient and accurate planning processes. Advantages • Efficiency: Automates complex planning tasks, reducing manual effort and minimizing errors. • Scalability: Handles large datasets and complex models, making it suitable for large-scale projects. • Community and Support: Open-source with active community contributions, ensuring continuous improvement and support. • Extensibility: Easily extendable with custom modules and plugins, allowing users to tailor the toolkit to their specific needs. The DSP Toolkit stands out by offering a robust, flexible, and user-friendly solution for distribution substation planning. Public Abstract

Li, Han [Lawrence Berkeley National Laboratory (LB↗

A Scalable Multi-Modal Framework for High-Fidelity Distributed Human Mobility Simulations

The development of data-driven models for human mobility in urban settings requires access to substantial and diverse real-world data. However, existing historical data often presents challenges such as limited volume, variety, and veracity, as well as missing data and privacy preservation concerns. Also, urban mobility modeling is inherently time-variant, complex, and multi-modal, encompassing everything from individual walking and running to private road travel and large-scale public transportation. These challenges call for innovative solutions to overcome data limitations and compute needs to model mobility behaviors accurately. To address these challenges, we propose a distributed, co-simulation-based architecture DURMOSim that integrates real-world data with scalable, high-fidelity simulations, demonstrating distributed co-simulation feasibility with existing mobility models. DURMOSim underpins a modular integration that would enable using any available mobility simulators for greater extensibility and scalability in performing various urban scenarios. In this paper, we present the design, implementation, and performance evaluation of DURMOSim, highlighting its capability to model population-scale mobility patterns. Our initial results show its ability to dynamically synchronize multiple simulation models at runtime with negligible computational overhead. We believe DURMOSim could be a robust tool for advancing urban mobility research and intelligent transportation systems.

Yoginath, Srikanth [ORNL] (ORCID:0000000184236050)↗

Toward Understanding the Differences between Mesoscale and Large-Eddy Simulations of Tropical Cyclones

In this work, we investigate the ability of mesoscale and large-eddy simulation (LES) model configurations to predict the mean wind speed profile within the boundary layer of tropical cyclones (TCs). To this end, we perform idealized simulations of five hypothetical intense storms ranging from categories 1 to 5 on the Saffir–Simpson scale and extract time-averaged quantities near the eyewall region. We compare the model-generated data against mean wind speed profiles compiled from dropsondes launched from reconnaissance aircraft operating in the North Atlantic basin. Our analysis shows that mesoscale- and LES-generated mean wind fields display important differences in the boundary layer, including the magnitude of shear as well as the height where their low-level wind speed maxima are located. In addition, a comparison between the two model configurations with the dropsonde data shows that both modeling approaches are unable to capture the typical structure of mean winds in the lower part of the TC boundary layer (10–500 m), calling into question the use of simulations of near-axisymmetric storms for investigating the wind structure of past events. To better understand these differences, we conduct a momentum-budget analysis and show that modeled turbulent fluxes are underestimated in the mesoscale boundary layer parameterization compared to the LES model. Based on the analysis of the horizontal turbulent fluxes and their potential impact on mean flow quantities, a TC-specific boundary layer parameterization may be needed.

17 WIND ENERGY↗

A GPU‐Accelerated Generative Adversarial Model for Causal Inference

We develop a GPU-accelerated machine learning generative adversarial model designed to facilitate causal inferences from observational data. Our model's theoretical framework is conceptualized in a manner that is amenable to being operable and scalable for high-performance computing platforms. We leverage GPU acceleration to develop a parallel evolutionary algorithm to achieve large-scale parallel computation of the model within a now widely accessible computing platform. This capability both enhances computational speedup and efficiency and also extends the use of the model to a broader range of substantive research domains while maintaining the underlying theoretical properties of the model.

GPU↗