Search NASA⌕ Search

SEARCH · Search NASA

Results for “Big Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Reducing Data Center Peak Cooling Demand and Energy Costs with Underground Thermal Energy Storage (UTES)

By recent estimates, data center energy demands are projected to consume between 6.7% and 12% of U.S. annual electricity generation by the year 2028, driven primarily by expanded demands from cloud services, big data analytics, and Artificial Intelligence (AI) (Shehabi et al., 2024). As much as 40% of data center total energy consumption are loads associated with the site infrastructure cooling systems, and these are often highly water consumptive (Aljbour et al., 2024). For energy system planners, this presents significant challenges to meeting and managing the anticipated loads, and especially the peak loads of projected data center deployments. Geothermal technologies offer two unique solutions to these challenges: 1) by serving loads through the deployment of new conventional and/or next-generation geothermal power technologies such as EGS and 2) through an often-overlooked opportunity to reduce data center peak cooling loads. The latter is the focus of this paper which explores Cold Underground Thermal Energy Storage ("Cold UTES") as an emerging industrial-scale geothermal cooling solution. This cooling solution is energy efficient, non-water-consumptive, and utilizes long duration energy storage (LDES) on both diurnal and seasonal time scales. Cold UTES has the potential to also function as a virtual power plant (VPP). The US Department of Energy's Geothermal Technologies Office is supporting R&D to understand the grid and system-wide value, costs, and impacts of deploying this emergent cooling solution at scale.

AI↗

Battery Health Quantification for TDRS Spacecraft by Using Signature Discriminability Measurement

The NASA/GSFC Space Network Project Office (SN) currently operates a constellation of ten geosynchronous TDRS spacecraft launched over the past 30 years. The SN project collects up to 16.5 Gigabytes of telemetry every month. Generally, the spacecraft health and functionality are obtained by the use of real-time telemetry data for the multiple spacecraft subsystems, which are transmitted to the main ground station at the White Sands Complex in Las Cruces, NM. Recently, the SN has instituted a program of Big Data to analyze the large amounts of data using a variety of tools including Machine Learning, Artificial Intelligence, development of training sets, and a variety of mathematical modeling tools. The goal is to improve spacecraft management and obtain a more accurate prediction of the spacecraft end of life. The combination of these efforts with those of the Aerospace Corporation, which has a contract with the SN to produce yearly reliability estimates for the TDRS fleet, will be performed. This paper presents a new concept called telemetry quality quantification (TQQ) and discusses the progress that has been made in battery performance estimation for the second-generation TDRS spacecraft using a signature discriminability measures (SDM) algorithm combined with the Aerospace Corp. battery life estimation models. This activity is important because many of the TDRS fleet of spacecraft have exceeded their on-orbit design lifetime and, therefore, NASA must carefully manage the spacecraft to continue operations while avoiding an end-of-mission scenario that leaves a non-functioning spacecraft in geosynchronous orbit.

Ma, Kenneth Y.↗

Fusion Approach for Remotely-Sensed Mapping of Agriculture (FARMA): A Scalable Open Source Method for Land Cover Monitoring Using Data Fusion

The increasing availability of very-high resolution (VHR; <2 m) imagery has the potential to enable agricultural monitoring at increased resolution and cadence, particularly when used in combination with widely available moderate-resolution imagery. However, scaling limitations exist at the regional level due to big data volumes and processing constraints. Here, we demonstrate the Fusion Approach for Remotely-Sensed Mapping of Agriculture (FARMA), using a suite of open source software capable of efficiently characterizing time-series field-scale statistics across large geographical areas at VHR resolution. We provide distinct implementation examples in Vietnam and Senegal to demonstrate the approach using WorldView VHR optical, Sentinel-1 Synthetic Aperture Radar, and Sentinel-2 and Sentinel-3 optical imagery. This distributed software is open source and entirely scalable, enabling large area mapping even with modest computing power. FARMA provides the ability to extract and monitor sub-hectare fields with multisensor raster signals, which previously could only be achieved at scale with large computational resources. Implementing FARMA could enhance predictive yield models by delineating boundaries and tracking productivity of smallholder fields, enabling more precise food security observations in low and lower-middle income countries.

fusion↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗

Air Traffic Management TestBed Data Exchange Model

The Air Traffic Management (ATM) TestBed is a Platform as a Service that is being developed by the National Aeronautics and Space Administration (NASA) to help design, configure, integrate, run, and monitor air traffic simulations. The platform is designed to provide cloud services including back-end, big-data analytics tools, on-demand computing resource management, data storage, and communication middleware. The ATM TestBed reduces the time to test concepts and technologies, supports interactions among various methods such as human-in-the-loop and automation-in-the-loop simulations, and enables collaborative simulations by sharing technologies and tools in the ATM community. In order to allow easier access to simulation components, TestBed provides a messaging support layer for connectivity using a consistent set of input/output interfaces. In addition, a standard data format is introduced to facilitate communication between the components. The data exchange model, supported in the messaging support layer, standardizes the format of the information to be exchanged among the components. This document describes the messaging data model currently developed in TestBed and provides data dictionaries for references to component developers as well as simulation engineers.

air traffic simulation↗

Integrating HPC, AI, and Workflows for Scientific Data Analysis: Report from Dagstuhl Seminar 23352

The Dagstuhl Seminar 23352, titled “Integrating HPC, AI, and Workflows for Scientific Data Analysis,” held from August 27 to September 1, 2023, was a significant event focusing on the synergy between High-Performance Computing (HPC), Artificial Intelligence (AI), and scientific workflow technologies. The seminar recognized that modern Big Data analysis in science rests on three pillars: workflow technologies for reproducibility and steering, AI and Machine Learning (ML) for versatile analysis, and HPC for handling large data sets. These elements, while crucial, have traditionally been researched separately, leading to gaps in their integration. The seminar aimed to bridge these gaps, acknowledging the challenges and opportunities at the intersection of these technologies. The event highlighted the complex interplay between HPC, workflows, and ML, noting how ML has increasingly been integrated into scientific workflows, thereby enhancing resource demands and bringing new requirements to HPC architectures, like support for GPUs and iterative computations. The seminar also addressed the challenges in adapting HPC for large-scale ML tasks, including in areas like deep learning, and the need for workflow systems to evolve to leverage ML in data analysis fully. Moreover, the seminar explored how ML could optimize scientific workflow systems and HPC operations, such as through improved scheduling and fault tolerance. A key focus was on identifying prestigious use cases of ML in HPC and understanding their unique, unmet requirements. The stochastic nature of ML and its impact on the reproducibility of data analysis on HPC systems was also a topic of discussion.

97 MATHEMATICS AND COMPUTING↗

Enabling Real-time Multi-messenger Astrophysics Discoveries with Deep Learning

Multi-messenger astrophysics is a fast-growing, interdisciplinary field that combines data, which vary in volume and speed of data processing, from many different instruments that probe the Universe using different cosmic messengers: electromagnetic waves, cosmic rays, gravitational waves and neutrinos. In this Expert Recommendation, we review the key challenges of real-time observations of gravitational wave sources and their electromagnetic and astroparticle counterparts, and make a number of recommendations to maximize their potential for scientific discovery. These recommendations refer to the design of scalable and computationally efficient machine learning algorithms; the cyber-infrastructure to numerically simulate astrophysical sources, and to process and interpret multi-messenger astrophysics data; the management of gravitational wave detections to trigger real-time alerts for electromagnetic and astroparticle follow-ups; a vision to harness future developments of machine learning and cyber-infrastructure resources to cope with the big-data requirements; and the need to build a community of experts to realize the goals of multi-messenger astrophysics.

E A Huerta↗

Integrating Intelligent Hydro-informatics into an effective Early Warning System for risk-informed urban flood management

The urban drainage system constantly facing flooding issues in coastal and urban areas. Robust and accurate urban flood management, particularly considering fast-moving compound floods, is crucial to minimize the impact of flood disasters in coastal cities. Till now, Ho Chi Minh City (HCMC) lacks an effective means of urban flood management because of flood risk communication among residents. Existing flood risk communication tools rely on post-disaster flood model outcomes and data. Therefore, this research proposes a real-time Early Urban Flooding Warning System (EUFWS) integrated with a user-friendly web and app interface. The backbone of this system consists of flood models developed using machine learning (ML) algorithms, combined with big data and Web-GIS visualization, with ML serving as the core for constructing the EUFWS. EUFWS offer several key advantages: they are available at all times, accessible from anywhere, and provide a real-time, multi-user working platform. Additionally, the system is flexible, allowing for the easy addition of components and services and scalable, adjusting to workload demands. EUFWS have been successfully deployed in Thu Duc City, Vietnam, as a case study and are operating effectively. EUFWS have been successfully deployed in Thu Duc City, Vietnam, as a case study and are operating effectively. Research results indicate that EUFWS supported decision-makers to be effectively risk informed and make intelligent decisions during urban flood emergencies. Finally, this underscores the significant potential of integrating ML and information technology to enhance the management of smart urban drainage systems in flood-prone cities worldwide.

54 ENVIRONMENTAL SCIENCES↗

Applications of Anomaly Detection and Precursor Identification in Airspace Operations

As we continue to advance the U.S. National Airspace into the next generation of air traffic, we face challenges in both increase in complexity, as well as, a significant growth in traffic volume. Addressing these challenges, while maintaining the same level of safety is an important application of data mining. Because of these significant shifts in airspace design and usage there is a need to identify current and emergent safety risks along with their potential precursors. In recent years NASA has made advancements in developing scalable methods to address this effort in the Big Data paradigm. Multiple kernel anomaly detection approaches have been employed on both surveillance radar data and flight operational quality assurance data to identify operationally significant safety risks. Additionally, events have been explored with a recently developed precursor identification tool to discover states that reveal an increased probability of a safety event. These tools can be used to discover emerging safety risks that may not be currently monitored, which allows for mitigation tactics to be employed and ultimately make the overall airspace safer. This talk will discuss an overview of these methods and a discussion of the findings.

anomaly detection↗

Investigation and Evaluation of Advanced Spectrum Management Concepts for Aeronautical Communications

With the emergence of new aerial vehicles into the airspace and the continued growth of aviation operations, there will be an increasing demand for voice and data communications within the National Airspace System (NAS). The continued use of existing VHF and UHF frequency allocations is not a sustainable approach, and as a result, a new spectrum management solution is required to support future mission needs. The proposed spectrum management concepts leverage modern advancements such as artificial intelligence (AI) and big data to dynamically optimize the spectrum utilization based on the predicted communications demand throughout the airspace. This technical investigation considers both air-ground and air-air communications networks, and can be applied to both existing applications such as the air traffic control (ATC) system, as well as future applications, such as the emerging Advanced Air Mobility (AAM). To support the evaluation of the proposed concepts and technologies, a modeling and simulation capability is currently under development and will continue to evolve to support new and advanced airspace applications. It is anticipated that this proposed spectrum concept will better serve the spectrum needs of future NAS applications.

Eric J Knoblock↗

Investigation and Evaluation of Advanced Spectrum Management Concepts for Aeronautical Communications

With the emergence of new aerial vehicles into the airspace and the continued growth of aviation operations, there will be an increasing demand for voice and data communications within the National Airspace System (NAS). The continued use of existing VHF and UHF frequency allocations is not a sustainable approach, and as a result, a new spectrum management solution is required to support future mission needs. The proposed spectrum management concepts leverage modern advancements such as artificial intelligence (AI) and big data to dynamically optimize the spectrum utilization based on the predicted communications demand throughout the airspace. This technical investigation considers both air-ground and air-air communications networks, and can be applied to both existing applications such as the air traffic control (ATC) system, as well as future applications, such as the emerging Advanced Air Mobility (AAM). To support the evaluation of the proposed concepts and technologies, a modeling and simulation capability is currently under development and will continue to evolve to support new and advanced airspace applications. It is anticipated that this proposed spectrum concept will better serve the spectrum needs of future NAS applications.

Eric J Knoblock↗

The Importance of Data Visualization: Incorporating Storytelling into the Scientific Presentation

From its inception in 2000, one of the primary tasks of the Biomedical Data Reduction Analysis (BDRA) group has been translation of large amounts of data into information that is relevant to the audience receiving it. BDRA helps translate data into an integrated model that supports both operational and research activities. This data integrated model and subsequent visual data presentations have contributed to BDRA's success in delivering the message (i.e., the story) that its customers have needed to communicate. This success has led to additional collaborations among groups that had previously not felt they had much in common until they worked together to develop solutions in an integrated fashion. As more emphasis is placed on working with "big data" and on showing how NASA's efforts contribute to the greater good of the American people and of the world, it becomes imperative to visualize the story of our data to communicate the greater message we need to share. METHODS To create and expand its data integrated model, BDRA has incorporated data from many different collaborating partner labs and other sources. Data are compiled from the repositories of the Lifetime Surveillance of Astronaut Health and the Life Sciences Data Archive, and from the individual laboratories at Johnson Space Center that support collection of data from medical testing, environmental monitoring, and countermeasures, as designated in the Medical Requirements Integration Documents. Ongoing communication with the participating collaborators is maintained to ensure that the message and story of the data are retained as data are translated into information and visual data presentations are delivered in different venues and to different audiences. RESULTS We will describe the importance of storytelling through an integrated model and of subsequent data visualizations in today's scientific presentations and discuss the collaborative methods used. We will illustrate the discussion with examples of graphs from BDRA's past work supporting operations and/or research efforts.

Babiak-Vazquez, A.↗

A Heuristic Approach to Correlating ERAM Flight Data from Twenty Centers

Among its many other functions, the Federal Aviation Administration’s En Route Automation Modernization (ERAM) provides external systems with real-time air traffic data for flights in enroute airspace in the National Airspace System. It replaced the En Route Host computer and backup system used at 20 FAA Air Route Traffic Control Centers (Centers) nationwide. Among the new features of ERAM, its output data stream of flight plan and track data includes a unique identifier for a flight originating in any one of the 20 ERAM Centers. The unique identifier, called the Global Unique Flight Identifier (GUFI), is persistent across all the Centers that track the flight. However, certain factors make it difficult to correlate data using the GUFI. First, the value of the GUFI is only unique within a time window of seven days. Second, the GUFI is attached only to flight-plan related data messages. Finally, track positions reported by ERAM do not reference the GUFI. In order to correlate historical as well as real time flight-plan and position related ERAM data, an efficient, heuristic approach was developed, and a prototype was developed. The approach showed that the processing speed, through parallel processing, is sufficient to correlate ERAM data in real-time. As described in this paper, when there are multiple track positions reported from multiple Centers within a few seconds, each position is assigned with a weighted score to indicate the quality of the position relative to its last know position. The weighted score can be used to eliminate potentially duplicate track positions. The approach is database-agnostic, and can be implemented in a Big Data system such as an Apache Hadoop system, as well as in traditional database systems.

Correlating ERAM Flight Data↗

EDX ClaiMM

EDX ClaiMM is a centralized data & analytical platform designed to revolutionize U.S. critical minerals and materials (CMM) activities. By providing a robust digital infrastructure, ClaiMM will accelerate the combination, leveraging, and rapid utilization of vital data, advanced tools, and cutting-edge research advancements in CMM. This adaptive digital research hub connects the CMM community to essential knowledge products and offers access to interoperable datasets, databases, models, software, and tools from the National Energy Technology’s (NETL’s) Energy Data eXchange (EDX) and other authoritative sources, serving both public and private sectors. EDX ClaiMM delivers AI-informed solutions to address fundamental knowledge gaps and fosters the innovation of new techniques for enhanced characterization and recovery of CMMs within the U.S. By leveraging cloud-hosted, scalable digital infrastructure, ClaiMM meets public–private applied energy needs. It equips the CMM community with priority digital resources that harness on-site and cloud compute capabilities, enabling big data storage, advanced processing, analytics, and visualization.

Critical Materials; Critical Minerals; Rare Earth ↗

Spatial validation reveals poor predictive performance of large-scale ecological mapping models

Mapping aboveground forest biomass is central for assessing the global carbon balance. However, current large-scale maps show strong disparities, despite good validation statistics of their underlying models. Here, we attribute this contradiction to a flaw in the validation methods, which ignore spatial autocorrelation (SAC) in data, leading to overoptimistic assessment of model predictive power. To illustrate this issue, we reproduce the approach of large-scale mapping studies using a massive forest inventory dataset of 11.8 million trees in central Africa to train and validate a random forest model based on multispectral and environmental variables. A standard nonspatial validation method suggests that the model predicts more than half of the forest biomass variation, while spatial validation methods accounting for SAC reveal quasi-null predictive power. This study underscores how a common practice in big data mapping studies shows an apparent high predictive power, even when predictors have poor relationships with the ecological variable of interest, thus possibly leading to erroneous maps and interpretations.

biomass↗

Leveraging the Usage of GPUs in SAR Processing for the NISAR Mission

The NASA ISRO Synthetic Aperture Radar (NISAR) mission will redefine the future of earth science in terms of both the quality as well as the quantity of data that will be downlinked daily. The current software architecture used to process this data is the InSAR Scientific Computing Environment (ISCE), a powerful and modular platform that applies a combination of novel and legacy processing modules to many sources of SAR data. Until recently, this architecture could process most images in a reasonable amount of time; however in the case of the NISAR mission (where the daily influx as well as the size of the images themselves are significantly larger) the current architecture can take hours to process even a single image. This paper explores new efforts to use a Graphics Processing Unit (GPU) to accelerate one of the processing modules to achieve unprecedented runtimes with no loss in precision, potentially setting a new standard in radar processing in the world of “Big Data”.

Cohen, Joshua↗

Time and Measurement Days

Questions in data analysis involving the concepts of time and measurement are often pushed into the background or reserved for a philosophical discussion. Some examples are: a) Is causality a consequence of the laws of physics, or can the arrow of time be reversed? b) Can we determine the arrow of time of an event? c) Do we need the continuum hypothesis for the underlying function in any measurement process? d) Can we say anything about the analyticity of the underlying process of an event? e) Would it be valid to model a non-analytical process as function of time? f) What are the implications of all these questions for classical Fourier techniques? However, in the age of big data gathered either from space missions supplying ultra-precise long time series, or e.g. LIGO data from the ground, the moment to bring these questions to the foreground seems arrived. The limitations of our understanding of some fundamental processes is emphasized by the lack of solution for problems open for more than 2 decades, such as the non-detection of solar g-modes, or the modal identification of main sequence stellar pulsators like delta Scuti stars. Flicker noise or 1/f noise, for example, attributed in solar-like stars to granulation, is analyzed mostly only to apply noise reduction techniques, neither considering the classical problem of 1/f noise that was introduced a 100 years ago, nor taking into account ergodic or non-ergodic solutions that make inapplicable spectral analysis techniques in practice. This topic was discussed by Nicholas W. Watkins during the ITISE meeting held in Granada in 2016. There he presented preliminary results of his research on Mandelbrot's related work. We reproduce here his quotation of Mandelbrot (1999) "There is a sharp contrast between a highly anomalous ("non-white") noise that proceeds in ordinary clock time and a noise whose principal anomaly is that it is restricted to fractal time", suggesting a connection with the above proposed topics that could be phrased as the following additional questions:a) Is self-organized criticality (SOC) frequent in astrophysical phenomena? b) Could all fractals in nature be considered stochastic? c) Could we establish mathematical/physical relationships between chaotic and fractal behaviors in time series? d) Could the differences between fractals and chaos in terms of analyticity be used to understand the residuals of the fitting of stellar light curves? In this meeting we would like to approximate these problems from a holistic and multidisciplinary perspective, taking into account not only technical issues but also the deeper implications. In particular the concept of connectivity (introduced in Pascual-Granado et al. A&A, 2015) could be used to implement, within the framework of ARMA processes, an "arrow of time" (see attached document), and so studying the possible implications in the concept of time as envisaged by Watkins.

data analysis↗

Few measurement shots challenge generalization in learning to classify entanglement

The ability to extract general laws from a few known examples depends on the complexity of the problem and on the amount of training data. In the quantum setting, the learner's generalization performance is further challenged by the destructive nature of quantum measurements that, together with the no-cloning theorem, limits the amount of information that can be extracted from each training sample. In this paper we focus on hybrid quantum learning techniques where classical machine-learning methods are paired with quantum algorithms and show that, in some settings, the uncertainty coming from a few measurement shots can be the dominant source of errors. We identify an instance of this possibly general issue by focusing on the classification of maximally entangled vs. separable states, showing that this toy problem becomes challenging for learners unaware of entanglement theory. Finally, we introduce an estimator based on classical shadows that performs better in the big data, few copy regime. Our results show that the naive application of classical machine-learning methods to the quantum setting is problematic, and that a better theoretical foundation of quantum learning is required.

97 MATHEMATICS AND COMPUTING↗