Search NASA⌕ Search

SEARCH · Search NASA

Results for “Big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Efficient First-Order Algorithms for Large-Scale, Non-Smooth Maximum Entropy Models with Application to Wildfire Science

Maximum entropy (MaxEnt) models are a class of statistical models that use the maximum entropy principle to estimate probability distributions from data. Due to the size of modern data sets, MaxEnt models need efficient optimization algorithms to scale well for big data applications. State-of-the-art algorithms for MaxEnt models, however, were not originally designed to handle big data sets; these algorithms either rely on technical devices that may yield unreliable numerical results, scale poorly, or require smoothness assumptions that many practical MaxEnt models lack. In this paper, we present novel optimization algorithms that overcome the shortcomings of state-of-the-art algorithms for training large-scale, non-smooth MaxEnt models. Our proposed first-order algorithms leverage the Kullback–Leibler divergence to train large-scale and non-smooth MaxEnt models efficiently. For MaxEnt models with discrete probability distribution of n elements built from samples, each containing m features, the stepsize parameter estimation and iterations in our algorithms scale on the order of O(mn) operations and can be trivially parallelized. Moreover, the strong ℓ1 convexity of the Kullback–Leibler divergence allows for larger stepsize parameters, thereby speeding up the convergence rate of our algorithms. To illustrate the efficiency of our novel algorithms, we consider the problem of estimating probabilities of fire occurrences as a function of ecological features in the Western US MTBS-Interagency wildfire data set. Our numerical results show that our algorithms outperform the state of the art by one order of magnitude and yield results that agree with physical models of wildfire occurrence and previous statistical analyses of wildfire drivers.

Physics↗

An overview of visualization and visual analytics applications in water resources management

Recent advances in information, communication, and environmental monitoring technologies have increased the availability, spatiotemporal resolution, and quality of water-related data, thereby leading to the emergence of many innovative big data applications. Among these applications, visualization and visual analytics, also known as the visual computing techniques, empower the synergy of computational methods (e.g., machine learning and statistical models) with human reasoning to improve the understanding and solution toward complex science and engineering problems. These approaches are frequently integrated with geographic information systems and cyberinfrastructure to provide new opportunities and methods for enhancing water resources management. Here, we present a comprehensive review of recent hydroinformatics applications that employ visual computing techniques to (1) support complex data-driven research problems, and (2) support the communication and decision-makings in the water resources management sector. Then, we conduct a technical review of the state-of-the-art web-based visualization technologies and libraries to share our experiences on developing shareable, adaptive, and interactive visualizations and visual interfaces for water resources management applications. We close with a vision that applies the emerging visual computing technologies and paradigms to develop the next generation of hydroinformatics applications.

54 ENVIRONMENTAL SCIENCES↗

CMOS-Based Single-Cycle in-Memory XOR/XNOR

Big data applications are on the rise, and so is the number of data centers. The ever-increasing massive data pool needs to be periodically backed up in a secure environment. Moreover, a massive amount of securely backed-up data is required for training binary convolutional neural networks for image classification. XOR and XNOR operations are essential for large-scale data copy verification, encryption, and classification algorithms. The disproportionate speed of existing compute and memory units makes the von Neumann architecture inefficient to perform these Boolean operations. Compute-in-memory (CiM) has proved to be an optimum approach for such bulk computations. The existing CiM-based XOR/XNOR techniques either require multiple cycles for computing or add to the complexity of the fabrication process. Here, we propose a CMOS-based hardware topology for single-cycle in-memory XOR/XNOR operations. Our design provides at least 2× improvement in the latency compared with other existing CMOS-compatible solutions. We verify the proposed system through circuit/system-level simulations and evaluate its robustness using a 5000-point Monte Carlo variation analysis. This all-CMOS design paves the way for practical implementation of CiM XOR/XNOR at scaled technology nodes.

97 MATHEMATICS AND COMPUTING↗

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information

Detecting and anticipating global proliferation expertise and capability evolution from unstructured, noisy, and incomplete public data streams is a highly desired, but extremely challenging task. Here, in this article, we present our pioneering data-driven approach to support the non-proliferation mission to detect and explain the evolution of proliferation expertise and capability development globally from terabytes of publicly available information (PAI), focusing on our knowledge extraction pipeline and descriptive analytics. We first discuss how we fuse nine open-source data streams, including multilingual data, to convert 4 TB of unstructured data to structured knowledge and encode dynamically evolving proliferation expertise representations—content and context graphs. For this, we rely on natural language processing (NLP) and deep learning (DL) models to perform information extraction, topic modeling, and distributed text representation (aka embedding) learning. We then present interactive, usable, and explainable descriptive analytics to refine domain knowledge and present it in a human-understandable form. Finally, we introduce future work avenues that will leverage our dynamic knowledge representations and descriptive analytics to enable predictive and prescriptive inferences to achieve real-time domain understanding and contextual reasoning about global proliferation expertise and capability evolution.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Fault Detection Utilizing Convolution Neural Network on Timeseries Synchrophasor Data From Phasor Measurement Units

An end-to-end supervised learning method is proposed for fault detection in the electric grid using Big Data from multiple Phasor Measurement Units (PMUs). The approach consists of preprocessing steps aimed at reducing data noise and dimensionality, followed by utilization of six classification models considered for detecting faults. Three of the models were variants of Convolutional Neural Network (CNN) architectures that consider a single type of measurement (voltage, current or frequency) at all PMUs or all types together also at all PMUs. CNN based models were compared to traditional methods of Logistic Regression (LR), Multi-layer Perceptron (MLP) and Support Vector Machine (SVM). Evaluation was conducted on two-year data measured by PMUs at 37 locations in a large electric grid. Here, the response variable for classification were extracted from the grid-wide outage event log. Experiments show that CNN-based models outperformed traditional methods on one year out-of-sample outage detection over the entire grid.

42 ENGINEERING↗

Roadmap on artificial intelligence and big data techniques for superconductivity

This paper presents a roadmap to the application of AI techniques and big data (BD) for different modelling, design, monitoring, manufacturing and operation purposes of different superconducting applications. To help superconductivity researchers, engineers, and manufacturers understand the viability of using AI and BD techniques as future solutions for challenges in superconductivity, a series of short articles are presented to outline some of the potential applications and solutions. These potential futuristic routes and their materials/technologies are considered for a 10–20 yr time-frame.

machine learning, neural network↗

Data-Driven Control: Theory and Applications

The ushering in of the big-data era, ably supported by exponential advances in computation, has provided new impetus to data-driven control in several engineering sectors. This topic's rapid and deep expansion has precipitated the need to showcase the highlights of data-driven approaches. There has been a rich history of contributions from the control systems community in data-driven control. At the same time, several new concepts and research directions have also been introduced in recent years. Many of these contributions and concepts have started to transition from theory to practical applications. This paper will overview the historical contributions and highlight recent concepts and research directions.

Soudbakhsh, Damoon↗

DeepPatent2: A Large-Scale Benchmarking Corpus for Technical Drawing Understanding

Abstract Recent advances in computer vision (CV) and natural language processing have been driven by exploiting big data on practical applications. However, these research fields are still limited by the sheer volume, versatility, and diversity of the available datasets. CV tasks, such as image captioning, which has primarily been carried out on natural images, still struggle to produce accurate and meaningful captions on sketched images often included in scientific and technical documents. The advancement of other tasks such as 3D reconstruction from 2D images requires larger datasets with multiple viewpoints. We introduce DeepPatent2, a large-scale dataset, providing more than 2.7 million technical drawings with 132,890 object names and 22,394 viewpoints extracted from 14 years of US design patent documents. We demonstrate the usefulness of DeepPatent2 with conceptual captioning. We further provide the potential usefulness of our dataset to facilitate other research areas such as 3D image reconstruction and image retrieval.

97 MATHEMATICS AND COMPUTING↗

Search for extended Lyman-α emission around 9k quasars at z = 2–3

ABSTRACT Enormous Lyα nebulae (ELANe) around quasars have provided unique insights into the formation of massive galaxies and their associations with super-massive black holes since their discovery. However, their detection remains highly limited. This paper introduces a systematic search for extended Lyα emission around 8683 quasars at z = 2.34–3.00 using a simple but very effective broad-band gri selection based on the Third Public Data Release of the Hyper Suprime-Cam Subaru Strategic Program. Although the broad-band selection detects only bright Lyα emission (≳ 1 × 10−17 erg s−1cm−2 arcsec−2) compared with narrow-band imaging and integral field spectroscopy, we can apply this method to far more sources than such common approaches. We first generated continuum g-band images without contributions from Lyα emission for host and satellite galaxies using r- and i-bands. Then, we established Lyα maps by subtracting them from observed g-band images with Lyα emissions. Consequently, we discovered extended Lyα emission (with masked area >40 arcsec2) for 7 and 32 out of 366 and 8317 quasars in the Deep and Ultra-deep (35 deg2) and Wide (890 deg2) layers, parts of which may be potential candidates of ELANe. However, none of them seem to be equivalent to the largest ELANe ever found. We detected higher fractions of quasars with large nebulae around more luminous or radio-loud quasars, supporting previous results. Future applications to the forthcoming big data from the Vera C. Rubin Observatory will help us detect more promising candidates. The source catalogue and obtained Lyα properties for all the quasar targets are accessible as online material.

79 ASTRONOMY AND ASTROPHYSICS↗

Region-adaptive, Error-controlled Scientific Data Compression using Multilevel Decomposition

The increase of computer processing speed is significantly outpacing improvements in network and storage bandwidth, leading to the big data challenge in modern science, where scientific applications can quickly generate much more data than that can be transferred and stored. As a result, big scientific data must be reduced by a few orders of magnitude while the accuracy of the reduced data needs to be guaranteed for further scientific explorations. Moreover, scientists are often interested in some specific spatial/temporal regions in their data, where higher accuracy is required. The locations of the regions requiring high accuracy can sometimes be prescribed based on application knowledge, while other times they must be estimated based on general spatial/temporal variation. In this paper, we develop a novel multilevel approach which allows users to impose region-wise compression error bounds. Our method utilizes the byproduct of a multilevel compressor to detect regions where details are rich and we provide the theoretical underpinning for region-wise error control. With spatially varying precision preservation, our approach can achieve significantly higher compression ratios than single-error bounded compression approaches and control errors in the regions of interest.We conduct the evaluations on two climate use cases – one targeting small-scale, node features and the other focusing on long, areal features. For both use cases, the locations of the features were unknown ahead of the compression. By selecting approximately 16% of the data based on multi-scale spatial variations and compressing those regions with smaller error tolerances than the rest, our approach improves the accuracy of post-analysis by approximately 2 × compared to single-error-bounded compression at the same compression ratio. Using the same error bound for the region of interest, our approach can achieve an increase of more than 50% in overall compression ratio.

Gong, Qian↗

Genome-Scale Metabolic Modeling Enables In-Depth Understanding of Big Data

Genome-scale metabolic models (GEMs) enable the mathematical simulation of the metabolism of archaea, bacteria, and eukaryotic organisms. GEMs quantitatively define a relationship between genotype and phenotype by contextualizing different types of Big Data (e.g., genomics, metabolomics, and transcriptomics). In this review, we analyze the available Big Data useful for metabolic modeling and compile the available GEM reconstruction tools that integrate Big Data. We also discuss recent applications in industry and research that include predicting phenotypes, elucidating metabolic pathways, producing industry-relevant chemicals, identifying drug targets, and generating knowledge to better understand host-associated diseases. In addition to the up-to-date review of GEMs currently available, we assessed a plethora of tools for developing new GEMs that include macromolecular expression and dynamic resolution. Finally, we provide a perspective in emerging areas, such as annotation, data managing, and machine learning, in which GEMs will play a key role in the further utilization of Big Data.

59 BASIC BIOLOGICAL SCIENCES↗

4th Big Data for Nuclear Power Plants Workshop 2023

The Ohio State University and Idaho National Laboratory organized the 4 th Big Data for Nuclear Power Plants Workshop in November, 2023 in Columbus, Ohio. Workshop topics were chosen to understand the challenges and gaps that need to be addressed to maximize the impact of data on the nuclear industry, as well as the associated applications and risks. Discussions were focused around six specific application areas: Operation and Maintenance; Machine Learning in Nuclear Materials and Advanced Manufacturing; Cybersecurity; High-Performance Computing and Massive Computation; Big Data and Digital Twins; and Nuclear Non-Proliferation. The opportunities, challenges, and risks identified in the six focus areas explored in this workshop are diverse, but some common themes emerge, such as the importance of data integrity, quality, coverage, privacy, and traceability. Big data and AI/ML tools can be leveraged to reduce costs, optimize human tasking, and reduce human error across various application areas. In order for the nuclear industry to benefit from big data and advanced analytic capabilities, it is essential to address challenges and risks, such as data privacy, model reliability, and computational resource availability. Learning from other industries that have successfully implemented big data and AI/ML technologies, like the aerospace industry, can help the nuclear industry successfully integrate these technologies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Development of a Unified Taxonomy for HVAC System Faults

Detecting and diagnosing HVAC faults is critical for maintaining building operation performance, reducing energy waste, and ensuring indoor comfort. An increasing deployment of commercial fault detection and diagnostics (FDD) software tools in commercial buildings in the past decade has significantly increased buildings’ operational reliability and reduced energy consumption. A massive amount of data has been generated by the FDD software tools. However, efficiently utilizing FDD data for ‘big data’ analytics, algorithm improvement, and other data-driven applications is challenging because the format and naming conventions of those data are very customized, unstructured, and hard to interpret. This paper presents the development of a unified taxonomy for HVAC faults. A taxonomy is an orderly classification of HVAC faults according to their characteristics and causal relations. The taxonomy includes fault categorization, physical hierarchy, fault library, relation model, and naming/tagging scheme. The taxonomy employs both a physical hierarchy of HVAC equipment and a cause-effect relationship model to reveal the root causes of faults in HVAC systems. A structured and standardized vocabulary library is developed to increase data representability and interpretability. The developed fault taxonomy can be used for HVAC system ‘big data’ analytics such as HVAC system fault prevalence analysis or the development of an HVAC FDD software standard. A common type of HVAC equipment-packaged rooftop unit (RTU) is used as an example to demonstrate the application of the developed fault taxonomy. Two RTU FDD software tools are used to show that after mapping FDD data according to the taxonomy, the meta-analysis of the multiple FDD reports is possible and efficient.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

RAPIDS: Reconciling Availability, Accuracy, and Performance in Managing Geo-Distributed Scientific Data

In modern science, big data plays an increasingly important role. Many scientific applications, such as running simulations on supercomputers or conducting experiments on advanced instruments, produce huge amount of data at unprecedented speed. Analyzing and understanding such big data is the key for scientists to make scientific breakthroughs. However, data might become unavailable for scientists to access when outages or maintenance of the storage system occur, which severely hinders scientific discovery. To improve the data availability, data duplication and erasure coding (EC) are often used. But as the scientific data gets larger, using these two methods can cause considerable storage and network overhead.In this paper, we propose RAPIDS, a hybrid approach that combines the multigrid-based error-bounded lossy compression with erasure coding, to significantly reduce the storage and network overhead required for maintaining high data availability. Our experiments show that RAPIDS reduces the storage overhead by up to 7.5x and network overhead by up to 3x to achieve the same level of availability compared to the regular EC method. We improve RAPIDS by building two models to optimize the fault tolerance configurations and data gathering strategy. We demonstrate that RAPIDS significantly improves performance when running on many CPU cores in parallel or on GPUs.

Wan, Lipeng↗