Search NASA⌕ Search

SEARCH · Search NASA

Results for “model data volumes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ESS-DIVE guidelines for archiving terrestrial model data

This dataset contains supporting documents and images for ESS-DIVE terrestrial model data archiving guidelines.Terrestrial models are broadly defined as numerical models that couple both land dynamics and energy, water, carbon, or nutrient fluxes. We created these guidelines based on input from the U.S. Department of Energy’s Biological and Environmental Research land modeling community. The guidelines are intended to help modelers determine which components of their terrestrial model data associated with publication should be archived. Based on input from the land modeling community, the guidelines recommend archiving both model input and testing data, as well as code, script, and metadata. The guidelines also recommend archiving model data output, depending on the limitations set by data repositories. Lastly, we provide recommendations for bundling data files for publication as well as a discussion about tools that can facilitate model data archiving and reuse.This dataset is an archive of the associated GitHub repository for our model archiving guidelines (https://github.com/ess-dive-community/essdive-model-data-archiving-guidelines). The ‘README.pdf’ file gives a general introduction to the guidelines, and the ‘instructions.pdf’ file provides more detailed steps for following the guidelines. We also provide 2 figures in this data package: 1) a decision tree (model_data_guidelines_decision_tree.png) that can help users determine which components of their model data to archive. and 2) the ‘model_data_guidelines_flmd.png’ file depicts the different files that can be archived in addition to the model data itself. Lastly, we include 3 digitized tables from our associated manuscript and 3 CSV files with anonymized input from DOE scientists about the importance of different aspects of model data archiving from which we developed the guidelines.Dataset updates for v1.1.0: We updated this data package on 2021-11-22 in response to review comments on our related manuscript. In this update we removed one figure so that the model archiving guidelines are conveyed in text rather than an image. We updated the file-level metadata (FLMD) figure to be in accord with the most recent FLMD recommendations. We made minor edits to the README file to update the recommended citation and added two co-authors. We also added 6 new data files (3 are anonymized input from DOE scientists that helped to inform guidelines, and 3 are digitized tables from our manuscript.

54 ENVIRONMENTAL SCIENCES↗

Spatial Transferability of Machine Learning Based Volume Estimation Models

High-quality traffic volume data is essential for efficient transportation planning and operations. However, such high-quality data is expensive to collect, owing primarily to the high capital cost of installing and maintaining continuous counting stations (CCSs). Recent availability of probe-based vehicle data offers a cost-effective solution for increasing the observability of traffic volumes. However, having ample ground truth traffic data is a prerequisite for developing robust volume estimation models. Though this might not be a big issue in many states, states with scarce CCS data might be able to benefit from robust volume estimation models developed in (adjacent) data-rich states. While there is a reasonable amount of spatial transferability research in the transportation domain, there is a dearth of knowledge on the spatial transferability of probe-based volume estimation models. To address this gap, this paper explores spatial transferability of volume estimation models developed from data in three states (Colorado, North Carolina, and Pennsylvania). Results indicate that it is extremely important to maintain temporal consistency when attempting spatial transferability of volume estimation models. It was also found that models trained on regions with lower peak traffic volumes will limit the performance of models transferred to states with higher peak hourly traffic volumes. Corroborating findings from existing spatial transferability research on other topics, it was found that a meta-model (developed using data from multiple states) performs better than volume estimation models developed within any one of the states.

ADVANCED PROPULSION SYSTEMS↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 4: Utah FORGE Well 78B-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 3: Utah FORGE Well 56-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 5: Utah FORGE Well 16B(78)-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

Analysis of Rig Parameter Data Using Drilling Process Modeling Constraints, Volume 1: Summary of Utah FORGE Wells 16A(78)-32, 56-32, 78B-32 and 16B(78)-32

Drill rig parameter measurements are routinely used during deep well construction to monitor and guide drilling conditions for improved performance and reduced costs. While insightful into the drilling process, these measurements are of reduced value without a standard to aid in data evaluation and decision making. In the main body of this work (Volume 1), a method is demonstrated whereby rock reduction model constraints are used to interpret drilling response parameters; the method could be applied in real-time to improve decision-making in the field and to further discern technology performance during post-drilling evaluations. Drilling parameters are evaluated using laboratory-validated rock reduction models for predicting the phenomenological response of drag bits (Detournay and Defourny, 1992) in computational algorithms. The method presented has applicability to development of advanced analytics on future geothermal wells using real-time electronic data recording for improved performance and reduced drilling costs. A drilling cost model is also used to show the tradeoff between rate of penetration and bit life and the influence on interval drilling costs. Details of the bit specifications and performance are cataloged in an independent volume, documented under separate cover, for each of the four wells, and include Volume 2: Utah FORGE 16A(78)-32; Volume 3: Utah FORGE 56-32; Volume 4: Utah FORGE 78B-32 and Volume 5: Utah FORGE 16B(78)-32.

15 GEOTHERMAL ENERGY↗

REV-INR: Regularized Evidential Implicit Neural Representation for Uncertainty-Aware Volume Visualization

Applications of Implicit Neural Representations (INRs) have emerged as a promising deep learning approach for compactly representing large volumetric datasets. These models can act as surrogates for volume data, enabling efficient storage and on-demand reconstruction via model predictions. However, conventional deterministic INRs only provide value predictions without insights into the model’s prediction uncertainty or the impact of inherent noisiness in the data. This limitation can lead to unreliable data interpretation and visualization due to prediction inaccuracies in the reconstructed volume. Identifying erroneous results extracted from model-predicted data may be infeasible, as raw data may be unavailable due to its large size. To address this challenge, we introduce REV-INR, Regularized Evidential Implicit Neural Representation, which learns to predict data values accurately along with the associated coordinate-level data uncertainty and model uncertainty using only a single forward pass of the trained REV-INR during inference. By comprehensively comparing and contrasting REV-INR with existing well-established deep uncertainty estimation methods, we show that REV-INR achieves the best volume reconstruction quality with robust data (aleatoric) and model (epistemic) uncertainty estimates using the fastest inference time. Consequently, we demonstrate that REV-INR facilitates assessment of the reliability and trustworthiness of the extracted isosurfaces and volume visualization results, enabling analyses to be solely driven by model-predicted data.

Saklani, Shanu [Indian Institute of Technology, Ka↗

RockFlow: Fast Generation of Synthetic Source Rock Images Using Generative Flow Models

Image-based evaluation methods are a valuable tool for source rock characterization. The time and resources needed to obtain images has spurred development of machine-learning generative models to create synthetic images of pore structure and rock fabric from limited image data. While generative models have shown success, existing methods for generating 3D volumes from 2D training images are restricted to binary images and grayscale volume generation requires 3D training data. Shale characterization relies on 2D imaging techniques such as scanning electron microscopy (SEM), and grayscale values carry important information about porosity, kerogen content, and mineral composition of the shale. Here, we introduce RockFlow, a method based on generative flow models that creates grayscale volumes from 2D training data. We apply RockFlow to baseline binary micro-CT image volumes and compare performance to a previously proposed model. We also show the extension of our model to 2D grayscale data by generating grayscale image volumes from 2D SEM and dual modality nanoscale shale images. The results show that our method underestimates the porosity and surface area on the binary baseline datasets but is able to generate realistic grayscale image volumes for shales. With improved binary data preprocessing, we believe that our model is capable of generating synthetic porous media volumes for a very broad class of rocks from shale to carbonates to sandstone.

58 GEOSCIENCES↗

Scalable Volume Visualization for Big Scientific Data Modeled by Functional Approximation

Considering the challenges posed by the space and time complexities in handling extensive scientific volumetric data, various data representations have been developed for the analysis of large-scale scientific data. Multivariate functional approximation (MFA) is an innovative data model designed to tackle substantial challenges in scientific data analysis. It computes values and derivatives with high-order accuracy throughout the spatial domain, mitigating artifacts associated with zero- or first-order interpolation. However, the slow query time through MFA makes it less suitable for interactively visualizing a large MFA model. In this work, we develop the first scalable interactive volume visualization pipeline, MFA-DVV, for the MFA model encoded from large-scale datasets. Our method achieves low input latency through distributed architecture, and its performance can be further enhanced by utilizing a compressed MFA model while still maintaining a high-quality rendering result for scientific datasets. We conduct comprehensive experiments to show that MFA-DVV can decrease the input latency and achieve superior visualization results for big scientific data compared with existing approaches.

big scientific dataset↗

Automatic recognition system for document digitization in nuclear power plants

With the increasing number of data-driven models in nuclear applications, large volumes of numerical data are required to accurately model and predict the health status of a plant component. However, many historical operation logs that contain useful information are not fully utilized due to the lack of a systematic approach of digitization. To overcome this issue, this study proposes an automatic pipeline for extracting information from handwritten tabular documents collected from nuclear power plants. In our pipeline, we first denoise scanned documents with morphological operations, and then extract relevant parts from individual pages using both traditional computer vision and neural network methods. Handwriting recognition is applied to obtain text and numbers. As the most challenging step is how to crop only relevant information, the main focus of our paper is to detect tables and cells from scanned handwritten documents. Here we evaluate the efficiency and accuracy of our proposed method on handwritten operational reports obtained from a real-world case study. The results demonstrate the high accuracy and practicality of our proposed method.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

With the end of Moore’s law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intelligent subsampling? To explore this, we develop SICKLE, a sparse intelligent curation framework for efficient learning, featuring a novel maximum entropy (MaxEnt) sampling approach, scalable training, and energy benchmarking. We compare MaxEnt with random and phase-space sampling on large direct numerical simulation (DNS) datasets of turbulence. Evaluating SICKLE at scale on Frontier, we show that subsampling as a preprocessing step can, in many cases, improve model accuracy and substantially lower energy consumption, with observed reductions of up to 38×.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

Differentiable modelling to unify machine learning and physical models for geosciences

Process-based modelling offers interpretability and physical consistency in many domains of geosciences but struggles to leverage large datasets efficiently. Machine-learning methods, especially deep networks, have strong predictive skills yet are unable to answer specific scientific questions. Here, in this Perspective, we explore differentiable modelling as a pathway to dissolve the perceived barrier between process-based modelling and machine learning in the geosciences and demonstrate its potential with examples from hydrological modelling. ‘Differentiable’ refers to accurately and efficiently calculating gradients with respect to model variables or parameters, enabling the discovery of high-dimensional unknown relationships. Differentiable modelling involves connecting (flexible amounts of) prior physical knowledge to neural networks, pushing the boundary of physics-informed machine learning. It offers better interpretability, generalizability, and extrapolation capabilities than purely data-driven machine learning, achieving a similar level of accuracy while requiring less training data. Additionally, the performance and efficiency of differentiable models scale well with increasing data volumes. Under data-scarce scenarios, differentiable models have outperformed machine-learning models in producing short-term dynamics and decadal-scale trends owing to the imposed physical constraints. Differentiable modelling approaches are primed to enable geoscientists to ask questions, test hypotheses, and discover unrecognized physical relationships. Future work should address computational challenges, reduce uncertainty, and verify the physical significance of outputs.

58 GEOSCIENCES↗

OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model

With the sharp increasing volume of user data, Deep Learning Recommendation Model (DLRM) becomes an indispensable infrastructure in large technology companies. However, large-scale DLRM on the multi-GPU platform is still inefficient due to unbalanced workload partitioning and intensive inter-GPU communication. To this end, we propose OPER, an OPtimality guided Embedding table placement for large-scale Recommendation model training and inference. OPER explores the potential of mitigating remote memory access latency in DLRM through fine-grained embedding table placement. Specifically, OPER proposes a theoretical modeling that builds up the relationship between EMT placement and the embedding communication latency in both training and inference. OPER proves the NP hardness of finding the optimal embedding table placement and proposes a heuristic algorithm that yields near optimal placement. OPER implements a SHMEM-based embedding table training system and a unified embedding index mapping to support fine-grained embedding table sharding and placement. Comprehensive experiments reveal that OPER achieves on average 3.4× and 5.1× speedup on training and inference respectively over state-of-the-art DLRM frameworks.

Wang, Zheng↗

Combining Observations and Models: A Review of the CARDAMOM Framework for Data‐Constrained Terrestrial Ecosystem Modeling

The rapid increase in the volume and variety of terrestrial biosphere observations (i.e., remote sensing data and in situ measurements) offers a unique opportunity to derive ecological insights, refine process‐based models, and improve forecasting for decision support. However, despite their potential, ecological observations have primarily been used to benchmark process‐based models, as many past and current models lack the capability to directly integrate observations and their associated uncertainties for parameterization. In contrast, data assimilation frameworks such as the CARbon DAta MOdel fraMework (CARDAMOM) and its suite of process‐based models, known as the Data Assimilation Linked Ecosystem Carbon Model (DALEC), are specifically designed for model‐data fusion. This review, motivated by a recent CARDAMOM community workshop, examines the development and applications of CARDAMOM, with an emphasis on its role in advancing ecosystem process understanding. CARDAMOM employs a Bayesian approach, using a Markov Chain Monte Carlo algorithm to enable data‐driven calibration of DALEC parameters and initial states (i.e., carbon pool sizes) through observation operators. CARDAMOM's unique ability to retrieve localized model process parameters from diverse datasets—ranging from in situ measurements to global satellite observations—makes it a highly flexible tool for analyzing spatially variable ecosystem responses to environmental change. However, assimilating these data also presents challenges, including data quality issues that propagate into model skill, as well as trade‐offs between model complexity, parameter equifinality, and predictive performance. We discuss potential solutions to these challenges, such as reducing parameter equifinality by incorporating new observations. This review also offers community recommendations for incorporating emerging datasets, integrating machine learning techniques, strengthening collaboration with remote sensing, field, and modeling communities, and expanding CARDAMOM's relevance for localized ecosystem monitoring and decision‐making. CARDAMOM enables a deep, mechanistic understanding of terrestrial ecosystem dynamics that cannot be achieved through empirical analyses of observational datasets or weakly constrained models alone.

Bayesian inference↗

Latent Stochastic Differential Equations for Modeling Quasar Variability and Inferring Black Hole Properties

Quasars are bright and unobscured active galactic nuclei (AGN) thought to be powered by the accretion of matter around supermassive black holes at the centers of galaxies. The temporal variability of a quasar’s brightness contains valuable information about its physical properties. The UV/optical variability is thought to be a stochastic process, often represented as a damped random walk described by a stochastic differential equation (SDE). Upcoming wide-field telescopes such as the Rubin Observatory Legacy Survey of Space and Time (LSST) are expected to observe tens of millions of AGN in multiple filters over a ten year period, so there is a need for efficient and automated modeling techniques that can handle the large volume of data. Latent SDEs are machine learning models well suited for modeling quasar variability, as they can explicitly capture the underlying stochastic dynamics. In this work, we adapt latent SDEs to jointly reconstruct multivariate quasar light curves and infer their physical properties such as the black hole mass, inclination angle, and temperature slope. Our model is trained on realistic simulations of LSST ten year quasar light curves, and we demonstrate its ability to reconstruct quasar light curves even in the presence of long seasonal gaps and irregular sampling across different bands, outperforming a multioutput Gaussian process regression baseline. Our method has the potential to provide a deeper understanding of the physical properties of quasars and is applicable to a wide range of other multivariate time series with missing data and irregular sampling.

79 ASTRONOMY AND ASTROPHYSICS↗

Uncertainty-Informed Volume Visualization using Implicit Neural Representation

The increasing adoption of Deep Neural Networks (DNNs) has led to their application in many challenging scientific visualization tasks. While advanced DNNs offer impressive generalization capabilities, understanding factors such as model prediction quality, robustness, and uncertainty is crucial. These insights can enable domain scientists to make informed decisions about their data. However, DNNs inherently lack ability to estimate prediction uncertainty, necessitating new research to construct robust uncertainty-aware visualization techniques tailored for various visualization tasks. In this work, we propose uncertainty-aware implicit neural representations to model scalar field data sets effectively and comprehensively study the efficacy and benefits of estimated uncertainty information for volume visualization tasks. We evaluate the effectiveness of two principled deep uncertainty estimation techniques: (1) Deep Ensemble and (2) Monte Carlo Dropout (MC-Dropout). These techniques enable uncertainty-informed volume visualization in scalar field data sets. Our extensive exploration across multiple data sets demonstrates that uncertainty-aware models produce informative volume visualization results. Moreover, integrating prediction uncertainty enhances the trustworthiness of our DNN model, making it suitable for robustly analyzing and visualizing real-world scientific volumetric data sets.

Saklani, Shanu↗

Tidal Currents in San Juan Archipelago, Washington

Re-analyzed acoustic Doppler current profiler (ADCP) data originally collected by NOAA CO-OPS (Center for Operational Oceanographic Products and Services) and equivalent point data from Pacific Northwest National Laboratory's FVCOM (Finite Volume Community Ocean Model) model of the region. Data are processed to products describing characteristics of tidal currents relevant to tidal turbines, as well as power output estimates for a notional turbine deployed from a surface platform or from the seabed at each location. These data underpin the results presented in their associated paper - see below.

16 TIDAL AND WAVE POWER↗

BISON High Burnup Structure Modeling Capabilities Validated with a Selection of the Halden IFA-650 Rods

The U.S. Department of Energy (DOE)’s Nuclear Energy Advanced Modeling and Simulation (NEAMS)program aims to develop predictive capabilities using computational methods for the analysis and design of advanced reactor and fuel cycle systems. This program has been supporting the development of BISON, a high-fidelity and high-resolution fuel performance tool at the engineering scale. This document continues analysis and refinement of capabilities added to BISON early this calendar year in regards to the incorporation of capabilities applicable to extended burnups in response to industry interest. Details are provided on high burnup thermal conductivity models, a refitting of the high-burnup structure (HBS) porosity formation model to include additional data, the coupling of the HBS volume fraction model to thermal conductivity and fine fragmentation models, and validation activities. The IFA-650.4 and IFA-650.9 loss of coolant accident (LOCA) analyses are revisited with the latest developments in this report. A new validation case, IFA-650.14 has also been added to the BISON test suite.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗